system

The system addresses the challenge of meeting individual user needs by converting voice input to text, analyzing it for personalized responses, and providing tailored suggestions through audio and visual output, improving user interaction and satisfaction.

JP2026064812APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing systems struggle to meet individual user needs and preferences in online shopping and remote services, lacking advanced interaction capabilities and personalized product/service suggestions.

Method used

A system that acquires user voice input, converts it to text, analyzes the text using natural language processing, generates appropriate responses, converts responses to audio, extracts user requests and preferences, stores them in a database, and proposes personalized products/services based on historical data, while providing visual information through a device.

Benefits of technology

Enables personalized and efficient product/service suggestions tailored to individual user needs, enhancing user experience by considering past conversation and transaction history.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064812000001_ABST
    Figure 2026064812000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of obtaining user voice input, A means of converting voice input to text, A means of analyzing the converted text and generating an appropriate reply, A method for converting replies into speech and outputting them, A means of extracting user requests and preferences and storing them in a database, A means of proposing products and services based on stored information, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, the demand for online shopping and remote services has been increasing, but it is difficult for existing systems to fully meet the individual needs and desires of users. In addition, there is a lack of advanced systems for improving the interaction with users and proposing products and services to individual users. Therefore, there is a need for a new system aimed at improving the user experience and work efficiency.

Means for Solving the Problems

[0005] This invention provides a system that includes means for acquiring user voice input, means for converting voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, means for extracting user requests and preferences and storing them in a database, and means for proposing products and services based on the stored information. This enables appropriate suggestions tailored to individual needs through natural conversation with the user, thereby improving the user experience. Furthermore, a system that includes means for making suggestions that take into account past conversation history and transaction history in association with user data, and means for providing a user interface through a device with a display and presenting visual information, enables the provision of more personalized and advanced services.

[0006] "Means for acquiring user voice input" refers to a device or software that has the function of receiving words spoken by a user as voice and converting them into a format that can be input into a computer system.

[0007] "Means of converting voice input to text" refers to a technology or process that analyzes acquired voice data and converts its content into text data.

[0008] "Means for analyzing converted text and generating appropriate responses" refers to an algorithm or system that analyzes text data using natural language processing techniques and generates appropriate responses or messages based on its content.

[0009] "Method for converting replies into audio and outputting them" refers to a technology that converts generated text messages into audio data and plays that audio through an output device such as a speaker.

[0010] "A means of extracting user requests and preferences and storing them in a database" refers to a system that identifies requests and preferences from the user's speech, organizes that information, and records it in a database.

[0011] "Means of proposing products and services based on stored information" refers to an algorithm or mechanism that references user information stored in a database and selects and presents the most suitable products and services based on that information.

[0012] "A means of making suggestions that take into account past conversation history and transaction history in relation to user data" refers to a system that analyzes data such as the user's past utterances and purchase history, and makes personalized suggestions based on that data.

[0013] "A means of providing a user interface through a device with a display and presenting visual information" refers to a technology that uses a device equipped with a display to visually show text and images to the user. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined. **Modes for Carrying Out the Invention**

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a system for suggesting appropriate products and services through dialogue with users. This system is realized by combining technologies such as speech recognition, natural language processing, data analysis, and speech synthesis.

[0036] System-wide configuration

[0037] The system consists of the following main components:

[0038] 1. Means for obtaining user voice input

[0039] 2. Means of converting voice input to text

[0040] 3. Means for analyzing the converted text and generating an appropriate response.

[0041] 4. A means of converting replies into speech and outputting them.

[0042] 5. A means of extracting user requests and preferences and storing them in a database.

[0043] 6. Means of proposing products and services based on stored information

[0044] 7. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[0045] 8. Means of presenting visual information using a device equipped with a display.

[0046] System operation details

[0047] 1. Obtain user voice input.

[0048] The device acquires user voice input through its microphone. When the user speaks into the device, the microphone captures the audio data.

[0049] 2. Convert voice input to text

[0050] The acquired audio data is converted into text using speech recognition software. This ensures that what the user says is treated as text data.

[0051] 3. Analyze the converted text and generate a reply.

[0052] The server analyzes text data using natural language processing technology. Based on the analyzed data, an AI model generates appropriate responses. For example, in response to a question like "I want to travel," it generates a response suggesting a suitable travel destination.

[0053] 4. Convert the reply to speech and output it.

[0054] The generated text reply is converted into speech using speech synthesis technology. The converted speech data is then played back to the user through the device's speaker.

[0055] 5. Extract requests and preferences and save them to a database.

[0056] The server analyzes the user's speech to extract their requests and preferences. The extracted information is stored in the database as a user profile.

[0057] 6. Propose products and services

[0058] The server uses stored user profile information to run an algorithm that suggests appropriate products and services. The suggested products and services are presented to the user in text or voice.

[0059] 7. Proposals that take past history into consideration

[0060] The server provides more personalized suggestions based on the user's past conversation and transaction history. This makes it possible to offer suggestions that match the user's preferences.

[0061] 8. Presentation of visual information

[0062] Using devices equipped with displays, visual information is also provided to the user. This includes product images and detailed information.

[0063] Specific example

[0064] Specific examples of users:

[0065] The user speaks to a smart speaker with a display.

[0066] "I want to travel, could you recommend some places?"

[0067] Processing flow:

[0068] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[0069] 2. Speech-to-text conversion: Speech input is converted into text data.

[0070] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[0071] 4. Audio output of replies: The generated text reply is converted to audio and played through the speaker.

[0072] 5. Extraction and saving of preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[0073] 6. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[0074] 7. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[0075] 8. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[0076] The above describes a specific form for carrying out the invention. This system allows users to receive personalized product and service suggestions while engaging in voice-based dialogue.

[0077] The following describes the processing flow.

[0078] Program processing

[0079] Step 1: Obtain user voice input

[0080] Operation details:

[0081] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[0082] Step 2: Convert voice input to text

[0083] Operation details:

[0084] The device converts the acquired audio data into text data using speech recognition software. This ensures that what the user says is processed as text.

[0085] Step 3: Send the converted text to the server.

[0086] Operation details:

[0087] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[0088] Step 4: Analyze the text data and generate an appropriate response.

[0089] Operation details:

[0090] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[0091] Step 5: Send the reply from the server to the terminal as text data.

[0092] Operation details:

[0093] The server sends the generated text reply to the terminal. This text data is converted to speech and played back on the terminal.

[0094] Step 6: Convert the reply to speech and output it.

[0095] Operation details:

[0096] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[0097] Step 7: Extract user requests and preferences and save them to the database.

[0098] Operation details:

[0099] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[0100] Step 8: Propose products and services based on the saved information.

[0101] Operation details:

[0102] The server uses information stored in the database to suggest products and services that meet the user's needs. These suggestions are generated as a reply message and presented to the user.

[0103] Step 9: Consider past conversation and transaction history in relation to user data.

[0104] Operation details:

[0105] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[0106] Step 10: Provide visual information

[0107] Operation details:

[0108] The terminal uses a display-equipped device to show the user visual information. This includes images, details, and prices of the proposed products and services.

[0109] The above outlines the specific processing steps of the program in this system. This allows users to receive personalized product and service suggestions while engaging in voice-based conversations.

[0110] (Example 1)

[0111] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0112] Modern users handle vast amounts of information and demand personalized services and product recommendations, but traditional systems struggle to efficiently achieve this. In particular, few systems can start with voice input and then provide recommendations that take into account the user's past history and preferences. Therefore, there is a need to develop systems that improve user satisfaction.

[0113] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0114] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text, means for analyzing the converted text and generating an appropriate response using a generation AI model, means for converting the response into voice and outputting it, means for extracting the user's requests and preferences and storing them in a database, and means for suggesting products and services based on the stored information. This makes it possible to efficiently provide personalized suggestions that start with voice input and take into account the user's past history and preferences.

[0115] "Means for acquiring user voice input" refers to a function in which the terminal uses a microphone to capture the user's voice in real time and sends it to the server as voice data.

[0116] "Means of converting voice input to text" refers to a function in which a server uses speech recognition technology to convert acquired voice data into text data.

[0117] "Means for analyzing converted text and generating appropriate responses using a generative AI model" refers to a function in which the server uses natural language processing technology and a generative AI model to analyze text data and automatically generate appropriate responses to user questions and requests.

[0118] "Method for converting replies into audio and outputting them" refers to a function in which the server uses speech synthesis technology to convert text replies into audio data and plays it back to the user through the terminal's speaker.

[0119] "A means of extracting user requests and preferences and saving them to a database" refers to a function where the server analyzes the user's speech and saves information about their requests and preferences to a database.

[0120] "A means of suggesting products and services based on stored information" refers to a function in which the server uses an algorithm to suggest appropriate products and services based on user information stored in the database.

[0121] "A means of making suggestions that take into account past conversation and transaction history in relation to user data" refers to a function in which the server refers to the user's past conversation and transaction history and makes personalized suggestions based on that.

[0122] "Means of providing a user interface through a display device and presenting visual information" refers to the function of a terminal that uses its display to provide the user with visual information (for example, detailed information or images of products or services).

[0123] Modes for carrying out the invention

[0124] This invention is a system for proposing appropriate products and services through interaction with the user. This system is realized by combining the following hardware and software technologies.

[0125] System-wide configuration

[0126] The system consists of the following main components:

[0127] 1. Means for obtaining user voice input

[0128] 2. Means of converting voice input to text

[0129] 3. A means of analyzing the converted text and generating an appropriate response using a generative AI model.

[0130] 4. A means of converting replies into speech and outputting them.

[0131] 5. A means of extracting user requests and preferences and storing them in a database.

[0132] 6. Means of proposing products and services based on stored information

[0133] 7. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[0134] 8. Means of presenting visual information through a display device.

[0135] Acquiring voice input

[0136] The device uses a microphone to acquire user voice input. When the user speaks into the device, the microphone captures the voice data and sends it to the server as digital data. This data is processed using speech recognition technology.

[0137] Speech-to-text conversion

[0138] The server uses speech recognition software (e.g., a speech recognition API) to convert the acquired speech data into text. This allows the speech input to be treated as text data.

[0139] Text analysis and response generation

[0140] The converted text data is analyzed by the server using natural language processing techniques and generative AI models (e.g., large-scale language models). Based on the analyzed data, an appropriate response is generated. For example, if a user says "I want to travel," the generative AI model is given the prompt "The user says they want to travel. What travel destination would you suggest?" and the model generates a response such as "Hokkaido is a recommended travel destination."

[0141] Voice output of the reply

[0142] The generated text reply is converted into audio data using speech synthesis technology (e.g., a speech synthesis API). The converted audio data is played back to the user through the device's speaker. The user can then hear the response, such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[0143] Extraction of requests and preferences and storage in a database.

[0144] The server analyzes the user's speech to extract their requests and preferences. Customer insight tools are used for this purpose. The extracted information is stored in a database as a user profile. For example, if a user is identified as being interested in "travel," that information is saved.

[0145] Product and service proposals

[0146] The server uses stored user profile information to execute algorithms (e.g., collaborative filtering) to suggest appropriate products and services. This information is presented to the user in text or audio format.

[0147] Proposals that take past history into consideration

[0148] The server references the user's past conversation and transaction history to provide more personalized suggestions. For example, it can pique the user's interest by prioritizing travel destinations they have never visited before.

[0149] Presentation of visual information

[0150] If the device has a display, it will use the display to provide visual information about the suggested products and services. This includes detailed product information and photos. Users can visually confirm the information, leading to a deeper understanding.

[0151] Specific example

[0152] Specific examples of users:

[0153] "I want to travel, could you recommend some places?"

[0154] Processing flow:

[0155] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[0156] 2. Speech-to-text conversion: Speech input is converted into text data using a speech recognition API.

[0157] 3. Text analysis and reply generation: The server uses a generated AI model to generate the reply "Hokkaido is a recommended travel destination."

[0158] 4. Audio output of replies: The generated text reply is converted into audio data using a speech synthesis API and played back through the device's speaker.

[0159] 5. Extraction and saving of preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[0160] 6. Proposal: The server uses collaborative filtering to suggest services and products related to the travel destination.

[0161] 7. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[0162] 8. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the display device.

[0163] As described above, the system of the present invention allows users to receive personalized product and service suggestions through voice and visual information. This enables users to efficiently acquire information and make choices that meet their needs.

[0164] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0165] Step 1:

[0166] The device acquires the user's voice input. When the user speaks into the device, saying "I want to travel," the device's microphone captures the voice. The input is an analog audio signal, which the microphone sensor converts into digital data. Digital audio data is generated, and this becomes the initial output.

[0167] Step 2:

[0168] The server receives the digital audio data obtained in step 1 and uses speech recognition software (e.g., a speech recognition API) to convert the audio into text. Specifically, it converts the audio data "I want to go on a trip" into text data "Text:I want to go on a trip". This converted text data becomes the output.

[0169] Step 3:

[0170] The server, upon receiving the converted text data, analyzes it using natural language processing techniques. This analysis utilizes a generative AI model (e.g., a large-scale language model). The prompt "The user says they want to travel. What travel destination would you suggest?" is input to the generative AI model, and the model generates an appropriate response, "Hokkaido is a recommended travel destination." The analyzed data is then output.

[0171] Step 4:

[0172] The server converts the text reply generated in step 3 into audio data using speech synthesis technology (e.g., a speech synthesis API). Specifically, it converts "My recommended travel destination is Hokkaido" into audio data and sends that audio data to the terminal. This audio data becomes the output of step 4.

[0173] Step 5:

[0174] The device receives audio data sent from the server and plays the audio to the user through its speaker. The user hears the audio say, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food." Through this action, the user obtains the necessary information.

[0175] Step 6:

[0176] Based on what the user says, the server automatically analyzes the data and extracts the user's requests and preferences. For example, keywords such as "travel" and "nature" may be extracted. The extracted information is stored in the database as a user profile. This stored user profile data is then output.

[0177] Step 7:

[0178] The server executes data analysis algorithms (e.g., collaborative filtering) based on user profile information stored in the database, and suggests appropriate products and services. For example, it might suggest accommodations or tour packages related to a travel destination. The content of these suggestions becomes the output.

[0179] Step 8:

[0180] The server considers the user's past conversation and transaction history to provide more personalized product and service suggestions. It also uses past data to suggest travel destinations the user has never visited before. This personalized suggestion information is then output.

[0181] Step 9:

[0182] If the device has a display, the server also provides visual information. For example, it might display detailed information or photos of a suggested travel destination. The user can visually confirm the displayed information. This visual information becomes the final output.

[0183] Through the steps outlined above, this system comprehensively achieves everything from voice input, text analysis, appropriate response generation using a generative AI model, speech synthesis, extraction and storage of user preferences, suggestions based on data analysis, and even the provision of visual information.

[0184] (Application Example 1)

[0185] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0186] In modern autonomous vehicles, there is a need for systems that suggest appropriate destinations and sightseeing spots in real time to enhance the passenger experience. However, conventional in-car entertainment and navigation systems have the problem of not being able to make suggestions that meet the individual needs and preferences of users. Furthermore, there has been a challenge in providing more personalized suggestions that take past history into account.

[0187] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0188] In this invention, the server includes means for acquiring user voice input, means for converting voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, means for extracting user requests and preferences and storing them in a database, means for suggesting products and services based on the stored information, and means for suggesting appropriate drive destinations and sightseeing spots to passengers in the vehicle. This enables real-time appropriate suggestions that meet the requests and preferences of passengers.

[0189] "Acquiring user voice input" means that microphones installed inside the vehicle capture the voices of passengers.

[0190] "Converting voice input to text" means converting acquired voice data into text data using speech recognition technology.

[0191] "Analyzing the converted text and generating appropriate responses" means using natural language processing technology to analyze text data and generate appropriate information and suggestions for passengers.

[0192] "Converting replies to audio and outputting them" means converting the generated replies into an audio format using speech synthesis technology and playing them through the car's speakers.

[0193] "Extracting user requests and preferences and storing them in a database" means extracting requests and preferences from passengers' statements and storing them in a database.

[0194] "Proposing products and services based on stored information" means selecting and proposing appropriate products and services based on customer needs and preferences stored in a database.

[0195] "Suggesting appropriate driving destinations and sightseeing spots to passengers inside the vehicle" means suggesting appropriate driving destinations and sightseeing spots in real time through dialogue with passengers inside the autonomous vehicle.

[0196] This invention relates to a system for suggesting appropriate driving destinations and sightseeing spots to passengers in an autonomous vehicle. This system is implemented using the following key hardware and software components, combining speech recognition, natural language processing, speech synthesis, and database technologies.

[0197] System Configuration

[0198] 1. Hardware Configuration

[0199] Microphone: A high-sensitivity microphone installed inside the vehicle is used. This microphone is used to capture passenger voice input. For example, the SureBity USB microphone is suitable for this purpose.

[0200] Speakers: Audio is output using the in-vehicle speakers. Voice replies to passengers are also delivered through these speakers.

[0201] Display-equipped devices: Display devices installed inside the vehicle are used to provide visual information.

[0202] 2. Software Configuration

[0203] Speech recognition software: Using the speech_recognition library, audio data acquired from the microphone is converted into text data.

[0204] Natural language processing software: Using the pipeline function of the transformers library, the converted text is analyzed to generate appropriate responses. In particular, the "GPT-3(registered trademark)" is used as the generative AI model.

[0205] Text-to-speech software: Using the gTTS (Google® Text-to-Speech) library, the generated text replies are converted into audio data and played back through the car's speakers.

[0206] Database: A database system used to store user requests and preferences. It stores past conversation and transaction history and is used to provide more personalized suggestions.

[0207] System operation

[0208] The operation of this system is described as follows:

[0209] 1. Speech acquisition and recognition

[0210] The terminal acquires the passenger's voice input through the microphone. Voice recognition software converts this into text data and sends it to the server.

[0211] 2. Natural Language Processing and Proposal Generation

[0212] The server analyzes text data using natural language processing techniques and generates appropriate suggestions. These suggestions are then converted into speech using speech synthesis software.

[0213] 3. Audio output and data storage

[0214] The converted audio data is played back to passengers through the in-car speakers. At the same time, information about the suggestions, as well as passenger requests and preferences, is stored in a database and used to provide more personalized suggestions in the future.

[0215] 4. Visual information provision

[0216] Information on suggested driving destinations and tourist spots is presented visually through the in-car display system. This allows passengers to understand the suggestions more concretely.

[0217] Examples and prompts for generative AI models

[0218] Specific example:

[0219] When a user says, "I'm looking for a good restaurant," the speech recognition system converts this into text and generates a prompt message like the following, which is then input into the AI ​​model.

[0220] Prompts for the generative AI model:

[0221] To make the most of our drive, could you recommend some good restaurants?

[0222] The implementation of this system will provide a comfortable and personalized experience within autonomous vehicles. In the concrete implementation of the invention, the aforementioned hardware and software will be integrated to enable real-time and highly accurate suggestions.

[0223] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0224] Step 1:

[0225] The user speaks into a microphone installed inside the vehicle. The terminal acquires the user's voice input. The input is the passenger's voice data, and the output is a captured raw audio file. This audio data is then passed on to subsequent speech recognition processing.

[0226] Step 2:

[0227] The device converts the acquired audio data into text data using speech recognition software (speech_recognition library). The input is the audio file acquired in step 1, and the output is text data. Specifically, the speech recognition software analyzes the audio waveform and converts it into a string based on a language model.

[0228] Step 3:

[0229] The server analyzes the converted text using natural language processing software (the pipeline function of the transformers library) and generates an appropriate response. The input is the text data obtained in step 2, and the output is a text response containing the suggested content. A generative AI model (e.g., gpt-3) is used to analyze the text data and generate prompt sentences that respond to the user's questions and requests.

[0230] Step 4:

[0231] The server converts the generated text reply into audio data using speech synthesis technology (gTTS library). The input is the text reply generated in step 3, and the output is an audio file. Specifically, the speech synthesis technology converts the text into an audio signal and saves it as an MP3 file.

[0232] Step 5:

[0233] The terminal plays the converted audio data to the user through the in-car speakers. The input is the audio file generated in step 4, and the output is the presentation to the user as audio. The audio is played using software that plays MP3 files (e.g., mpg321).

[0234] Step 6:

[0235] The server extracts user requests and preferences and stores them in a database. The input is the text data obtained in step 3, and the output is a new record in the user profile database. Natural language processing techniques are used to extract information about requests and preferences from the text, organize it, and add it to the database.

[0236] Step 7:

[0237] The server executes an algorithm that suggests products and services based on stored information. The input is user profile information from the database, and the output is text data containing new suggestions. These suggestions are then provided to the user in subsequent speech recognition and natural language processing steps.

[0238] Step 8:

[0239] The server or terminal visually presents information about suggested drive destinations and tourist spots via an in-car display device. The input is the suggested content obtained in step 7, and the output is the visual information displayed on the display. Image data and map data are used to visually represent the suggested content.

[0240] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0241] This invention combines a system that proposes appropriate products and services through dialogue with the user with an emotion engine that recognizes the user's emotions. This system is realized by combining technologies such as speech recognition, natural language processing, emotion analysis, data analysis, and speech synthesis.

[0242] System-wide configuration

[0243] The system consists of the following main components:

[0244] 1. Means for obtaining user voice input

[0245] 2. Means of converting voice input to text

[0246] 3. Means for analyzing the converted text and generating an appropriate response.

[0247] 4. A means of converting replies into speech and outputting them.

[0248] 5. An emotion engine that recognizes user emotions and reflects that emotional information in the analysis results.

[0249] 6. A means of extracting user requests and preferences and storing them in a database.

[0250] 7. Means of proposing products and services based on stored information

[0251] 8. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[0252] 9. Means of presenting visual information using a device equipped with a display.

[0253] System operation details

[0254] 1. Obtain user voice input.

[0255] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[0256] 2. Convert voice input to text

[0257] The acquired audio data is converted into text data using speech recognition software. This ensures that the user's spoken content is treated as text data.

[0258] 3. Send the converted text to the server.

[0259] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[0260] 4. Analyze text data and generate a reply.

[0261] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[0262] 5. Emotional analysis using an emotion engine

[0263] The server uses an emotion engine to recognize the user's emotions from voice and text data. The emotion engine analyzes the user's voice tone and text content to determine the user's emotional state. This emotional information is then reflected in the replies and suggestions.

[0264] 6. Send the reply as text data from the server to the terminal.

[0265] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[0266] 7. Convert the reply to speech and output it.

[0267] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[0268] 8. Extract user requests and preferences and save them to a database.

[0269] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[0270] 9. Propose products and services.

[0271] The server uses stored user profile information to run an algorithm that suggests products and services tailored to the user's needs. The suggestions are generated as a reply message and presented to the user.

[0272] 10. Proposals that take past history into consideration

[0273] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[0274] 11. Presentation of visual information

[0275] The terminal uses a display-equipped device to show the user visual information. This includes product images, detailed information, and prices.

[0276] Specific example

[0277] Specific examples of users:

[0278] The user speaks to a smart speaker with a display.

[0279] "I want to travel, but please tell me a recommended place."

[0280] Process flow:

[0281] 1. Acquisition of voice input: The user's voice is acquired by the microphone of the terminal.

[0282] 2. Conversion from voice to text: The voice input is converted into text data.

[0283] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply saying, "A recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[0284] 4. Sentiment analysis: The server's sentiment engine analyzes the user's voice and text and determines that the user is excited.

[0285] 5. Reflection of sentiment information: Reflect the sentiment information in the generated reply and adjust the reply in an energetic tone.

[0286] 6. Voice output of the reply: The generated text reply is converted into voice and played back from the speaker in an energetic tone.

[0287] 7. Extraction and storage of preferences: The user's preferences are extracted with "travel" as the keyword and stored in the database.

[0288] 8. Recommendation: Based on the information stored in the server, the server recommends services and products related to the travel destination.

[0289] 9. Recommendation considering history: Referring to the past history, the server preferentially recommends places that the user has never visited before.

[0290] 10. Presentation of visual information: Photos and information about the recommended travel destination are displayed on the display.

[0291] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[0292] The following describes the processing flow.

[0293] Program processing

[0294] Step 1: Obtain user voice input

[0295] Operation details:

[0296] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[0297] Step 2: Convert voice input to text

[0298] Operation details:

[0299] The device converts the acquired audio data into text data using speech recognition software. This ensures that what the user says is processed as text.

[0300] Step 3: Send the converted text to the server.

[0301] Operation details:

[0302] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[0303] Step 4: Analyze the text data and generate a reply.

[0304] Operation details:

[0305] The server analyzes the received text data using natural language processing technology. Based on the analysis results, a generative AI model generates an appropriate reply to the user's question or request.

[0306] Step 5: Perform sentiment analysis using the sentiment engine

[0307] Operation details:

[0308] The server uses a sentiment engine to recognize the user's sentiment from voice data or text data. The sentiment engine analyzes the tone of the user's voice and the text content to determine the user's emotional state.

[0309] Step 6: Reflect sentiment information in the reply

[0310] Operation details:

[0311] Based on the sentiment information obtained from the sentiment engine, the server adjusts the content and tone of the generated reply. For example, if the user is excited, the server will reply in an energetic tone.

[0312] Step 7: Transmit the reply from the server to the terminal as text data

[0313] Operation details:

[0314] The server transmits the generated reply to the terminal as text data. This text data is converted to voice and played back on the terminal side.

[0315] Step 8: Convert the reply to voice and output it

[0316] Operation details:

[0317] The terminal converts the received text reply to voice data using voice synthesis technology. The converted voice data is played back to the user through the speaker.

[0318] Step 9: Extract user requests and preferences and save them to the database.

[0319] Operation details:

[0320] The server analyzes the user's speech to extract requests and preferences. The extracted information is stored in a database as a user profile.

[0321] Step 10: Propose products and services based on the saved information.

[0322] Operation details:

[0323] The server uses information stored in the database to suggest products and services that meet the user's needs. These suggestions are generated as a reply message and presented to the user.

[0324] Step 11: Make a proposal that takes past history into consideration.

[0325] Operation details:

[0326] The server provides more personalized suggestions based on the user's past conversation and transaction history. This allows for suggestions tailored to the user's preferences and characteristics.

[0327] Step 12: Provide visual information

[0328] Operation details:

[0329] The terminal uses a display-equipped device to show the user visual information. This includes images, details, and prices of the proposed products and services.

[0330] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[0331] (Example 2)

[0332] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0333] Modern users demand fast and personalized product and service recommendations, requiring more accurate suggestions that take their emotional state into account. However, existing systems struggle to recognize user emotions and adjust their recommendations accordingly. Furthermore, they often fail to fully utilize user conversation and transaction history, resulting in a lack of information that users truly need.

[0334] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0335] In this invention, the server includes means for using an emotion engine to recognize the user's emotions, means for sending the generated response from the server to the terminal as text data, and means for making suggestions that take into account past conversation and transaction history in association with user data. This makes it possible to make more personalized product and service suggestions based on the user's emotions and past history.

[0336] "Means for acquiring user voice input" refers to a function that captures voice data spoken by the user using the device's microphone.

[0337] "Methods for converting voice input to text" refers to technologies that use speech recognition software to convert captured voice data into text format.

[0338] "Means for sending converted text to a server" refers to the function by which the terminal sends the text data generated by speech recognition to a server via the network.

[0339] "Means for analyzing text data and generating appropriate replies" refers to a function in which the server analyzes the text data it receives using natural language processing technology and generates an appropriate reply based on its content.

[0340] "Methods of using an emotion engine that recognizes user emotions from voice and text data" refers to a technology in which a server analyzes voice and text data, recognizes the user's emotional state, and incorporates that information into processing.

[0341] "Means for sending generated replies as text data from the server to the terminal" refers to a function for sending replies generated by the server in text format to the terminal.

[0342] "A means of converting replies into speech and outputting them" refers to a function that converts text data received by the device into speech format using speech synthesis technology and plays it back to the user through the speaker.

[0343] "A means of extracting user requests and preferences and storing them in a database" refers to a technology in which a server analyzes the user's speech to extract requests and preferences and stores that information in a database.

[0344] "A means of suggesting products and services based on stored information" refers to a technology in which a server refers to stored user profile information and automatically suggests products and services that are suitable for the user.

[0345] "A means of making suggestions that take into account past conversation and transaction history in relation to user data" refers to a technology in which a server operates an algorithm based on the user's past conversation and transaction history to generate personalized suggestions.

[0346] "A means of providing a user interface through a device with a display and presenting visual information" refers to the function of a terminal that uses its display to present visual information (images, detailed information, price, etc.) to the user.

[0347] This invention combines a system that proposes appropriate products and services through dialogue with the user with an emotion engine that recognizes the user's emotions. This system is realized by combining technologies such as speech recognition, natural language processing, emotion analysis, data analysis, and speech synthesis.

[0348] System-wide configuration

[0349] The system consists of the following main components:

[0350] 1. Means for obtaining user voice input

[0351] 2. Means of converting voice input to text

[0352] 3. Means for sending the converted text to the server

[0353] 4. Means for analyzing text data and generating replies

[0354] 5. An emotion engine that recognizes user emotions and reflects that emotional information in the analysis results.

[0355] 6. Means for sending the generated reply from the server to the terminal.

[0356] 7. A means of converting replies into speech and outputting them.

[0357] 8. A means of extracting user requests and preferences and storing them in a database.

[0358] 9. Means of proposing products and services based on stored information

[0359] 10. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[0360] 11. Means for presenting visual information using a device equipped with a display.

[0361] System operation details

[0362] Get user voice input

[0363] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[0364] Convert voice input to text

[0365] The acquired audio data is converted into text data using speech recognition software. For example, the Google Cloud Speech-to-Text API can be used to treat what the user says as text data.

[0366] Send the converted text to the server.

[0367] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[0368] Analyze text data and generate a reply.

[0369] The server analyzes the received text data using natural language processing technology. For example, it uses a generative AI model such as OpenAI® GPT-4® to generate appropriate responses to user questions and requests.

[0370] Emotional analysis using an emotion engine

[0371] The server uses an emotion engine to recognize user emotions from voice and text data. For example, it uses "IBM Watson® Tone Analyzer" to analyze the user's voice tone and text content to determine the user's emotional state. This emotion information is then reflected in the content of replies and suggestions.

[0372] Send the generated reply from the server to the terminal.

[0373] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[0374] Convert the reply into speech and output it.

[0375] The device converts received text replies into audio data using speech synthesis technology. For example, the converted audio data, using "Amazon Polly," is played back to the user through the speaker.

[0376] Extract user requests and preferences and save them to a database.

[0377] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[0378] Propose products and services

[0379] The server uses stored user profile information to run an algorithm that suggests products and services tailored to the user's needs. The suggestions are generated as a reply message and presented to the user.

[0380] Proposals that take past history into consideration

[0381] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[0382] Presentation of visual information

[0383] The terminal uses a display-equipped device to show the user visual information. This includes product images, detailed information, and prices.

[0384] Specific example

[0385] Specific examples of users:

[0386] The user speaks to a smart speaker with a display: "I want to travel, could you recommend some places?"

[0387] Processing flow:

[0388] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[0389] 2. Speech-to-text conversion: Speech is converted to text using speech recognition software (e.g., Google Cloud Speech-to-Text).

[0390] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[0391] 4. Sentiment Analysis: An emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's voice and text to determine if the user is excited.

[0392] 5. Reflecting Emotional Information: Reflect emotional information in the generated replies and adjust them to have an energetic tone.

[0393] 6. Voice output of replies: The generated text replies are converted into speech using speech synthesis technology (e.g., Amazon Polly) and played back through the speaker in an energetic tone.

[0394] 7. Extracting and saving preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[0395] 8. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[0396] 9. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[0397] 10. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[0398] Example of a prompt:

[0399] I want to go on a trip, could you recommend some places?

[0400] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[0401] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0402] Step 1:

[0403] The device acquires the user's voice input.

[0404] Input: The microphone captures the user's speech.

[0405] Operation: When the user says, "I want to travel, can you recommend some places?", the device's microphone captures the audio data.

[0406] Output: Captured audio data.

[0407] Step 2:

[0408] The device converts voice input into text.

[0409] Input: Acquired audio data.

[0410] Operation: The device uses the Google Cloud Speech-to-Text API to convert speech data into text data.

[0411] Output: Converted text data (e.g., "I want to travel, could you recommend some places?").

[0412] Step 3:

[0413] The terminal sends the converted text to the server.

[0414] Input: Text data.

[0415] Operation: The terminal sends text data to the server over the network.

[0416] Output: Text data sent to the server.

[0417] Step 4:

[0418] The server analyzes the text data and generates a reply.

[0419] Input: Text data received by the server.

[0420] Operation: The server uses natural language processing techniques (e.g., OpenAI GPT-4) to analyze text data and generate appropriate responses.

[0421] Output: Generated reply text (Example: "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food.").

[0422] Step 5:

[0423] The server uses an emotion engine to perform emotion analysis.

[0424] Input: Audio data or text data.

[0425] Operation: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. For example, it can determine whether the user is excited or not.

[0426] Output: Sentiment analysis results (e.g., user is agitated).

[0427] Step 6:

[0428] The server sends the generated reply as text data to the terminal.

[0429] Input: Generated reply text and sentiment analysis results.

[0430] Operation: The server sends a reply text to the terminal and adjusts it to reflect sentiment information.

[0431] Output: Reply text sent to the terminal.

[0432] Step 7:

[0433] The device converts the reply text into speech and outputs it.

[0434] Input: Reply text data.

[0435] Operation: The device uses speech synthesis technology such as Amazon Polly to convert text into audio data, which is then played through the speaker.

[0436] Output: Audio data that the user hears (e.g., "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food.").

[0437] Step 8:

[0438] The server extracts user requests and preferences and stores them in a database.

[0439] Input: Text data of the user's spoken content.

[0440] Operation: The server analyzes the spoken content, extracts requests and preferences, and stores them in a database. For example, it might extract keywords related to "travel."

[0441] Output: Extracted user profile information.

[0442] Step 9:

[0443] The server suggests products and services based on the information it stores.

[0444] Input: User profile information.

[0445] Operation: The server executes an algorithm that suggests products and services that meet the user's needs.

[0446] Output: A reply message regarding the proposed product or service (e.g., "We have a promotional discount for travel. Are you interested?").

[0447] Step 10:

[0448] The server makes suggestions that take past history into consideration.

[0449] Input: User's past conversation and transaction history.

[0450] Operation: The server generates personalized suggestions based on past conversation and transaction history. For example, it might suggest avoiding places you've visited before.

[0451] Output: Personalized suggestions.

[0452] Step 11:

[0453] The device displays visual information.

[0454] Input: Data about the proposed product or service.

[0455] Operation: The device uses its display to present visual information to the user. This includes displaying product images, detailed information, and prices.

[0456] Output: Visual information displayed on the screen (e.g., photos and prices of suggested travel destinations).

[0457] (Application Example 2)

[0458] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0459] Traditional advertising systems have a problem in that they present uniform ads without considering user emotions, thus failing to maximize advertising effectiveness. Furthermore, they lack sufficient personalized suggestions based on user requests and preferences, and there is a need to improve the user experience. The challenge is to improve advertising effectiveness while increasing user satisfaction by recognizing emotions and presenting timely and appropriate ads.

[0460] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for acquiring the user's voice input, means for converting the voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, an emotion engine means for recognizing the user's emotions and reflecting that emotion information in the response and suggestions, means for extracting the user's requests and preferences and storing them in a database, means for suggesting products and services based on the stored information, and a user interface means for presenting the suggestions in audio and visual ways. This makes it possible to provide personalized advertising suggestions that take the user's emotions into consideration.

[0461] "Acquiring voice input" is the process of capturing the voice spoken by the user through the device's microphone.

[0462] "Converting to text" is the process of converting acquired voice input into text data using a speech recognition algorithm.

[0463] "Text analysis" is the process of analyzing converted text data using natural language processing technology to understand and interpret its content.

[0464] "Generating appropriate responses" is the process of creating the best possible response to a user's question or request based on the results of text analysis.

[0465] "Converting replies to speech" is the process of converting generated text responses into speech data using speech synthesis technology.

[0466] "Audio output" is the process of making audio data audible to the user through an audio device such as a speaker.

[0467] An "emotion engine" is an algorithm or component that analyzes a user's voice or text to determine the user's emotional state (joy, sadness, anxiety, etc.).

[0468] "Extracting requests and preferences" is the process of identifying what an individual wants and likes from their spoken content and recording it in a database.

[0469] "Product and service proposals" refer to the process of selecting the most suitable products and services for a given user based on their saved user data, and then creating a proposal.

[0470] A "user interface" is an interface for a user to interact with a system, and includes means of presenting audio and visual information.

[0471] Overall System Overview

[0472] The system implementing this invention provides personalized advertisements to users via terminals such as smartphones and smart glasses. It acquires the user's voice input and performs a series of processes to present appropriate responses and advertisements that reflect emotions. The system components include voice input means, voice recognition means, text analysis means, emotion analysis means, database storage means, product / service suggestion means, and visual and audio output means. Its specific operation is described below.

[0473] Hardware and software used

[0474] The entire system will be implemented using the following hardware and software:

[0475] A device such as a smartphone or smart glasses (including microphone, speaker, and display)

[0476] Microphone: A device for acquiring voice input.

[0477] Speaker: A device for outputting sound.

[0478] Display: A device for displaying visual information.

[0479] Speech recognition software: Software used to convert acquired speech into text (e.g., Google Speech Recognition API)

[0480] Sentiment analysis model: A model for determining a user's emotions (for example, a sentiment analysis model based on BERT).

[0481] Natural Language Processing Models: Models that analyze user text and generate optimal replies or advertisements (e.g., generative AI models).

[0482] Text-to-speech conversion software (TTS engine)

[0483] Database: Storage for saving user preferences and requests (e.g., SQL database)

[0484] Detailed Operation Description

[0485] 1. Acquisition of voice input:

[0486] The device acquires user voice input through the microphone. When the user speaks, the voice data is acquired and the process proceeds to the next stage.

[0487] 2. Convert voice input to text:

[0488] The acquired audio data is converted into text data using speech recognition software (e.g., Google Speech Recognition API). This ensures that what the user says is treated as text data.

[0489] 3. Analysis of text data:

[0490] The server analyzes text data using natural language processing techniques and generates appropriate responses to user questions and requests. It uses a generative AI model to generate the responses the user expects.

[0491] 4. Emotion analysis:

[0492] The server uses an emotion engine to recognize the user's emotions from text and audio data. It analyzes the tone of voice and the content of speech to determine the user's emotional state. For example, if the user is excited, it will generate an energetic advertisement.

[0493] 5. Replies and ad generation:

[0494] Based on the results of sentiment analysis, ads and replies are generated that match the user's emotional state. For example, if the user is happy, ads with a bright and positive tone are generated.

[0495] 6. Convert replies to speech:

[0496] The server converts the generated text reply into audio data using speech synthesis technology (e.g., a TTS engine). The converted audio data is then played back to the user through the speaker.

[0497] 7. Presentation of visual information:

[0498] The device displays the generated advertisement content on its screen. Visual information such as product images, details, and prices are presented.

[0499] Specific example

[0500] For example, the user might say the following:

[0501] "Please tell me about the latest smartphones."

[0502] The system acquires the audio and generates advertisements through the following process.

[0503] Voice input acquisition

[0504] Text conversion using speech recognition

[0505] Text data analysis and response generation

[0506] Determining a user's happy emotions through emotion analysis.

[0507] Generate an energetic ad ("The latest smartphones have amazing camera features!")

[0508] Convert text replies to speech and output them through the speaker.

[0509] Display visual product information on the screen.

[0510] Example of a prompt

[0511] "The user said, 'Tell me about the latest smartphones.' The user seems very interested and is talking enthusiastically. What kind of ad would be suitable?"

[0512] This prompt allows the generative AI model to generate appropriate advertisements.

[0513] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0514] Step 1: Obtaining voice input

[0515] The user speaks into the device. The device acquires the user's voice input through the microphone. The input is the user's voice data, and the output is this voice data.

[0516] Specific operation: The microphone captures the user's voice and converts it into a digital format.

[0517] Step 2: Convert voice input to text

[0518] The device uses speech recognition software to convert acquired speech data into text data. The input is speech data, and the output is text data.

[0519] Specific operation: Speech recognition software (e.g., Google Speech Recognition API) analyzes the audio data and generates a corresponding string.

[0520] Step 3: Analyzing text data

[0521] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, it generates an appropriate response. The input is text data, and the output is the analysis results and the generated response.

[0522] Specific operation: A natural language processing model (e.g., a generative AI model) analyzes text data and generates a contextually appropriate response.

[0523] Step 4: Emotion Analysis

[0524] The server uses an emotion engine to recognize the user's emotions from text and audio data. The input is text and audio data, and the output is the emotion analysis result.

[0525] Specific operation: The emotion analysis model analyzes the text content and voice tone to determine the user's emotional state (e.g., joy, sadness, excitement, etc.).

[0526] Step 5: Replies and ad generation

[0527] The server considers the results of sentiment analysis to generate advertisements and replies appropriate to the user's emotional state. The input is the analysis results and sentiment analysis results, and the output is the generated replies and advertisements.

[0528] Specific operation: The generative AI model creates advertising content that reflects emotional information. For example, if the user is excited, it will generate an advertisement in an energetic tone.

[0529] Step 6: Convert your reply to speech

[0530] The server converts the generated text replies into speech data using speech synthesis technology. The input is text data, and the output is speech data.

[0531] Specific operation: A speech synthesis engine (e.g., a TTS engine) synthesizes text data and outputs it as natural-sounding speech.

[0532] Step 7: Audio Output

[0533] The device plays the converted audio data to the user through its speaker. The input is audio data, and the output is the audio that the user hears.

[0534] Specific operation: The speaker plays the audio data as a machine-generated voice.

[0535] Step 8: Presenting visual information

[0536] The device uses its display to visually show the generated advertisement content. The input is the advertisement content, and the output is the visual information displayed on the screen.

[0537] Specific operation: The display shows advertising images and detailed information, providing the user with visual feedback.

[0538] Step 9: Extract and save requests and preferences

[0539] The server extracts requests and preferences from the user's speech and stores them in a database. The input is text data, and the output is a user profile stored in the database.

[0540] Specific operation: A natural language processing model extracts important keywords and phrases from text and records them in a database.

[0541] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0542] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0543] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0544] [Second Embodiment]

[0545] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0546] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0547] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0548] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0549] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0550] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0551] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0552] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0553] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0554] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0555] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0556] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0557] This invention is a system for suggesting appropriate products and services through dialogue with users. This system is realized by combining technologies such as speech recognition, natural language processing, data analysis, and speech synthesis.

[0558] System-wide configuration

[0559] The system consists of the following main components:

[0560] 1. Means for obtaining user voice input

[0561] 2. Means of converting voice input to text

[0562] 3. Means for analyzing the converted text and generating an appropriate response.

[0563] 4. A means of converting replies into speech and outputting them.

[0564] 5. A means of extracting user requests and preferences and storing them in a database.

[0565] 6. Means of proposing products and services based on stored information

[0566] 7. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[0567] 8. Means of presenting visual information using a device equipped with a display.

[0568] System operation details

[0569] 1. Obtain user voice input.

[0570] The device acquires user voice input through its microphone. When the user speaks into the device, the microphone captures the audio data.

[0571] 2. Convert voice input to text

[0572] The acquired audio data is converted into text using speech recognition software. This ensures that what the user says is treated as text data.

[0573] 3. Analyze the converted text and generate a reply.

[0574] The server analyzes text data using natural language processing technology. Based on the analyzed data, an AI model generates appropriate responses. For example, in response to a question like "I want to travel," it generates a response suggesting a suitable travel destination.

[0575] 4. Convert the reply to speech and output it.

[0576] The generated text reply is converted into speech using speech synthesis technology. The converted speech data is then played back to the user through the device's speaker.

[0577] 5. Extract requests and preferences and save them to a database.

[0578] The server analyzes the user's speech to extract their requests and preferences. The extracted information is stored in the database as a user profile.

[0579] 6. Propose products and services

[0580] The server uses stored user profile information to run an algorithm that suggests appropriate products and services. The suggested products and services are presented to the user in text or voice.

[0581] 7. Proposals that take past history into consideration

[0582] The server provides more personalized suggestions based on the user's past conversation and transaction history. This makes it possible to offer suggestions that match the user's preferences.

[0583] 8. Presentation of visual information

[0584] Using devices equipped with displays, visual information is also provided to the user. This includes product images and detailed information.

[0585] Specific example

[0586] Specific examples of users:

[0587] The user speaks to a smart speaker with a display.

[0588] "I want to travel, could you recommend some places?"

[0589] Processing flow:

[0590] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[0591] 2. Speech-to-text conversion: Speech input is converted into text data.

[0592] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[0593] 4. Audio output of replies: The generated text reply is converted to audio and played through the speaker.

[0594] 5. Extraction and saving of preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[0595] 6. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[0596] 7. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[0597] 8. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[0598] The above describes a specific form for carrying out the invention. This system allows users to receive personalized product and service suggestions while engaging in voice-based dialogue.

[0599] The following describes the processing flow.

[0600] Program processing

[0601] Step 1: Obtain user voice input

[0602] Operation details:

[0603] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[0604] Step 2: Convert voice input to text

[0605] Operation details:

[0606] The device converts the acquired audio data into text data using speech recognition software. This ensures that what the user says is processed as text.

[0607] Step 3: Send the converted text to the server.

[0608] Operation details:

[0609] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[0610] Step 4: Analyze the text data and generate an appropriate response.

[0611] Operation details:

[0612] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[0613] Step 5: Send the reply from the server to the terminal as text data.

[0614] Operation details:

[0615] The server sends the generated text reply to the terminal. This text data is converted to speech and played back on the terminal.

[0616] Step 6: Convert the reply to speech and output it.

[0617] Operation details:

[0618] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[0619] Step 7: Extract user requests and preferences and save them to the database.

[0620] Operation details:

[0621] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[0622] Step 8: Propose products and services based on the saved information.

[0623] Operation details:

[0624] The server uses information stored in the database to suggest products and services that meet the user's needs. These suggestions are generated as a reply message and presented to the user.

[0625] Step 9: Consider past conversation and transaction history in relation to user data.

[0626] Operation details:

[0627] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[0628] Step 10: Provide visual information

[0629] Operation details:

[0630] The terminal uses a display-equipped device to show the user visual information. This includes images, details, and prices of the proposed products and services.

[0631] The above outlines the specific processing steps of the program in this system. This allows users to receive personalized product and service suggestions while engaging in voice-based conversations.

[0632] (Example 1)

[0633] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0634] Modern users handle vast amounts of information and demand personalized services and product recommendations, but traditional systems struggle to efficiently achieve this. In particular, few systems can start with voice input and then provide recommendations that take into account the user's past history and preferences. Therefore, there is a need to develop systems that improve user satisfaction.

[0635] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0636] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text, means for analyzing the converted text and generating an appropriate response using a generation AI model, means for converting the response into voice and outputting it, means for extracting the user's requests and preferences and storing them in a database, and means for suggesting products and services based on the stored information. This makes it possible to efficiently provide personalized suggestions that start with voice input and take into account the user's past history and preferences.

[0637] "Means for acquiring user voice input" refers to a function in which the terminal uses a microphone to capture the user's voice in real time and sends it to the server as voice data.

[0638] "Means of converting voice input to text" refers to a function in which a server uses speech recognition technology to convert acquired voice data into text data.

[0639] "Means for analyzing converted text and generating appropriate responses using a generative AI model" refers to a function in which the server uses natural language processing technology and a generative AI model to analyze text data and automatically generate appropriate responses to user questions and requests.

[0640] "Method for converting replies into audio and outputting them" refers to a function in which the server uses speech synthesis technology to convert text replies into audio data and plays it back to the user through the terminal's speaker.

[0641] "A means of extracting user requests and preferences and saving them to a database" refers to a function where the server analyzes the user's speech and saves information about their requests and preferences to a database.

[0642] "A means of suggesting products and services based on stored information" refers to a function in which the server uses an algorithm to suggest appropriate products and services based on user information stored in the database.

[0643] "A means of making suggestions that take into account past conversation and transaction history in relation to user data" refers to a function in which the server refers to the user's past conversation and transaction history and makes personalized suggestions based on that.

[0644] "Means of providing a user interface through a display device and presenting visual information" refers to the function of a terminal that uses its display to provide the user with visual information (for example, detailed information or images of products or services).

[0645] Modes for carrying out the invention

[0646] This invention is a system for proposing appropriate products and services through interaction with the user. This system is realized by combining the following hardware and software technologies.

[0647] System-wide configuration

[0648] The system consists of the following main components:

[0649] 1. Means for obtaining user voice input

[0650] 2. Means of converting voice input to text

[0651] 3. A means of analyzing the converted text and generating an appropriate response using a generative AI model.

[0652] 4. A means of converting replies into speech and outputting them.

[0653] 5. A means of extracting user requests and preferences and storing them in a database.

[0654] 6. Means of proposing products and services based on stored information

[0655] 7. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[0656] 8. Means of presenting visual information through a display device.

[0657] Acquiring voice input

[0658] The device uses a microphone to acquire user voice input. When the user speaks into the device, the microphone captures the voice data and sends it to the server as digital data. This data is processed using speech recognition technology.

[0659] Speech-to-text conversion

[0660] The server uses speech recognition software (e.g., a speech recognition API) to convert the acquired speech data into text. This allows the speech input to be treated as text data.

[0661] Text analysis and response generation

[0662] The converted text data is analyzed by the server using natural language processing techniques and generative AI models (e.g., large-scale language models). Based on the analyzed data, an appropriate response is generated. For example, if a user says "I want to travel," the generative AI model is given the prompt "The user says they want to travel. What travel destination would you suggest?" and the model generates a response such as "Hokkaido is a recommended travel destination."

[0663] Voice output of the reply

[0664] The generated text reply is converted into audio data using speech synthesis technology (e.g., a speech synthesis API). The converted audio data is played back to the user through the device's speaker. The user can then hear the response, such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[0665] Extraction of requests and preferences and storage in a database.

[0666] The server analyzes the user's speech to extract their requests and preferences. Customer insight tools are used for this purpose. The extracted information is stored in a database as a user profile. For example, if a user is identified as being interested in "travel," that information is saved.

[0667] Product and service proposals

[0668] The server uses stored user profile information to execute algorithms (e.g., collaborative filtering) to suggest appropriate products and services. This information is presented to the user in text or audio format.

[0669] Proposals that take past history into consideration

[0670] The server references the user's past conversation and transaction history to provide more personalized suggestions. For example, it can pique the user's interest by prioritizing travel destinations they have never visited before.

[0671] Presentation of visual information

[0672] If the device has a display, it will use the display to provide visual information about the suggested products and services. This includes detailed product information and photos. Users can visually confirm the information, leading to a deeper understanding.

[0673] Specific example

[0674] Specific examples of users:

[0675] "I want to travel, could you recommend some places?"

[0676] Processing flow:

[0677] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[0678] 2. Speech-to-text conversion: Speech input is converted into text data using a speech recognition API.

[0679] 3. Text analysis and reply generation: The server uses a generated AI model to generate the reply "Hokkaido is a recommended travel destination."

[0680] 4. Audio output of replies: The generated text reply is converted into audio data using a speech synthesis API and played back through the device's speaker.

[0681] 5. Extraction and saving of preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[0682] 6. Proposal: The server uses collaborative filtering to suggest services and products related to the travel destination.

[0683] 7. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[0684] 8. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the display device.

[0685] As described above, the system of the present invention allows users to receive personalized product and service suggestions through voice and visual information. This enables users to efficiently acquire information and make choices that meet their needs.

[0686] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0687] Step 1:

[0688] The device acquires the user's voice input. When the user speaks into the device, saying "I want to travel," the device's microphone captures the voice. The input is an analog audio signal, which the microphone sensor converts into digital data. Digital audio data is generated, and this becomes the initial output.

[0689] Step 2:

[0690] The server receives the digital audio data obtained in step 1 and uses speech recognition software (e.g., a speech recognition API) to convert the audio into text. Specifically, it converts the audio data "I want to go on a trip" into text data "Text:I want to go on a trip". This converted text data becomes the output.

[0691] Step 3:

[0692] The server, upon receiving the converted text data, analyzes it using natural language processing techniques. This analysis utilizes a generative AI model (e.g., a large-scale language model). The prompt "The user says they want to travel. What travel destination would you suggest?" is input to the generative AI model, and the model generates an appropriate response, "Hokkaido is a recommended travel destination." The analyzed data is then output.

[0693] Step 4:

[0694] The server converts the text reply generated in step 3 into audio data using speech synthesis technology (e.g., a speech synthesis API). Specifically, it converts "My recommended travel destination is Hokkaido" into audio data and sends that audio data to the terminal. This audio data becomes the output of step 4.

[0695] Step 5:

[0696] The device receives audio data sent from the server and plays the audio to the user through its speaker. The user hears the audio say, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food." Through this action, the user obtains the necessary information.

[0697] Step 6:

[0698] Based on what the user says, the server automatically analyzes the data and extracts the user's requests and preferences. For example, keywords such as "travel" and "nature" may be extracted. The extracted information is stored in the database as a user profile. This stored user profile data is then output.

[0699] Step 7:

[0700] The server executes data analysis algorithms (e.g., collaborative filtering) based on user profile information stored in the database, and suggests appropriate products and services. For example, it might suggest accommodations or tour packages related to a travel destination. The content of these suggestions becomes the output.

[0701] Step 8:

[0702] The server considers the user's past conversation and transaction history to provide more personalized product and service suggestions. It also uses past data to suggest travel destinations the user has never visited before. This personalized suggestion information is then output.

[0703] Step 9:

[0704] If the device has a display, the server also provides visual information. For example, it might display detailed information or photos of a suggested travel destination. The user can visually confirm the displayed information. This visual information becomes the final output.

[0705] Through the steps outlined above, this system comprehensively achieves everything from voice input, text analysis, appropriate response generation using a generative AI model, speech synthesis, extraction and storage of user preferences, suggestions based on data analysis, and even the provision of visual information.

[0706] (Application Example 1)

[0707] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0708] In modern autonomous vehicles, there is a need for systems that suggest appropriate destinations and sightseeing spots in real time to enhance the passenger experience. However, conventional in-car entertainment and navigation systems have the problem of not being able to make suggestions that meet the individual needs and preferences of users. Furthermore, there has been a challenge in providing more personalized suggestions that take past history into account.

[0709] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0710] In this invention, the server includes means for acquiring user voice input, means for converting voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, means for extracting user requests and preferences and storing them in a database, means for suggesting products and services based on the stored information, and means for suggesting appropriate drive destinations and sightseeing spots to passengers in the vehicle. This enables real-time appropriate suggestions that meet the requests and preferences of passengers.

[0711] "Acquiring user voice input" means that microphones installed inside the vehicle capture the voices of passengers.

[0712] "Converting voice input to text" means converting acquired voice data into text data using speech recognition technology.

[0713] "Analyzing the converted text and generating appropriate responses" means using natural language processing technology to analyze text data and generate appropriate information and suggestions for passengers.

[0714] "Converting replies to audio and outputting them" means converting the generated replies into an audio format using speech synthesis technology and playing them through the car's speakers.

[0715] "Extracting user requests and preferences and storing them in a database" means extracting requests and preferences from passengers' statements and storing them in a database.

[0716] "Proposing products and services based on stored information" means selecting and proposing appropriate products and services based on customer needs and preferences stored in a database.

[0717] "Suggesting appropriate driving destinations and sightseeing spots to passengers inside the vehicle" means suggesting appropriate driving destinations and sightseeing spots in real time through dialogue with passengers inside the autonomous vehicle.

[0718] This invention relates to a system for suggesting appropriate driving destinations and sightseeing spots to passengers in an autonomous vehicle. This system is implemented using the following key hardware and software components, combining speech recognition, natural language processing, speech synthesis, and database technologies.

[0719] System Configuration

[0720] 1. Hardware Configuration

[0721] Microphone: A high-sensitivity microphone installed inside the vehicle is used. This microphone is used to capture passenger voice input. For example, the SureBity USB microphone is suitable for this purpose.

[0722] Speakers: Audio is output using the in-vehicle speakers. Voice replies to passengers are also delivered through these speakers.

[0723] Display-equipped devices: Display devices installed inside the vehicle are used to provide visual information.

[0724] 2. Software Configuration

[0725] Speech recognition software: Using the speech_recognition library, audio data acquired from the microphone is converted into text data.

[0726] Natural language processing software: Using the pipeline function of the transformers library, the transformed text is analyzed to generate appropriate responses. In particular, "gpt-3" is used as the generative AI model.

[0727] Text-to-speech software: Using the gTTS (Google Text-to-Speech) library, the generated text replies are converted into audio data and played back through the car's speakers.

[0728] Database: A database system used to store user requests and preferences. It stores past conversation and transaction history and is used to provide more personalized suggestions.

[0729] System operation

[0730] The operation of this system is described as follows:

[0731] 1. Speech acquisition and recognition

[0732] The terminal acquires the passenger's voice input through the microphone. Voice recognition software converts this into text data and sends it to the server.

[0733] 2. Natural Language Processing and Proposal Generation

[0734] The server analyzes text data using natural language processing techniques and generates appropriate suggestions. These suggestions are then converted into speech using speech synthesis software.

[0735] 3. Audio output and data storage

[0736] The converted audio data is played back to passengers through the in-car speakers. At the same time, information about the suggestions, as well as passenger requests and preferences, is stored in a database and used to provide more personalized suggestions in the future.

[0737] 4. Visual information provision

[0738] Information on suggested driving destinations and tourist spots is presented visually through the in-car display system. This allows passengers to understand the suggestions more concretely.

[0739] Examples and prompts for generative AI models

[0740] Specific example:

[0741] When a user says, "I'm looking for a good restaurant," the speech recognition system converts this into text and generates a prompt message like the following, which is then input into the AI ​​model.

[0742] Prompts for the generative AI model:

[0743] To make the most of our drive, could you recommend some good restaurants?

[0744] The implementation of this system will provide a comfortable and personalized experience within autonomous vehicles. In the concrete implementation of the invention, the aforementioned hardware and software will be integrated to enable real-time and highly accurate suggestions.

[0745] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0746] Step 1:

[0747] The user speaks into a microphone installed inside the vehicle. The terminal acquires the user's voice input. The input is the passenger's voice data, and the output is a captured raw audio file. This audio data is then passed on to subsequent speech recognition processing.

[0748] Step 2:

[0749] The device converts the acquired audio data into text data using speech recognition software (speech_recognition library). The input is the audio file acquired in step 1, and the output is text data. Specifically, the speech recognition software analyzes the audio waveform and converts it into a string based on a language model.

[0750] Step 3:

[0751] The server analyzes the converted text using natural language processing software (the pipeline function of the transformers library) and generates an appropriate response. The input is the text data obtained in step 2, and the output is a text response containing the suggested content. A generative AI model (e.g., gpt-3) is used to analyze the text data and generate prompt sentences that respond to the user's questions and requests.

[0752] Step 4:

[0753] The server converts the generated text reply into audio data using speech synthesis technology (gTTS library). The input is the text reply generated in step 3, and the output is an audio file. Specifically, the speech synthesis technology converts the text into an audio signal and saves it as an MP3 file.

[0754] Step 5:

[0755] The terminal plays the converted audio data to the user through the in-car speakers. The input is the audio file generated in step 4, and the output is the presentation to the user as audio. The audio is played using software that plays MP3 files (e.g., mpg321).

[0756] Step 6:

[0757] The server extracts user requests and preferences and stores them in a database. The input is the text data obtained in step 3, and the output is a new record in the user profile database. Natural language processing techniques are used to extract information about requests and preferences from the text, organize it, and add it to the database.

[0758] Step 7:

[0759] The server executes an algorithm that suggests products and services based on stored information. The input is user profile information from the database, and the output is text data containing new suggestions. These suggestions are then provided to the user in subsequent speech recognition and natural language processing steps.

[0760] Step 8:

[0761] The server or terminal visually presents information about suggested drive destinations and tourist spots via an in-car display device. The input is the suggested content obtained in step 7, and the output is the visual information displayed on the display. Image data and map data are used to visually represent the suggested content.

[0762] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0763] This invention combines a system that proposes appropriate products and services through dialogue with the user with an emotion engine that recognizes the user's emotions. This system is realized by combining technologies such as speech recognition, natural language processing, emotion analysis, data analysis, and speech synthesis.

[0764] System-wide configuration

[0765] The system consists of the following main components:

[0766] 1. Means for obtaining user voice input

[0767] 2. Means of converting voice input to text

[0768] 3. Means for analyzing the converted text and generating an appropriate response.

[0769] 4. A means of converting replies into speech and outputting them.

[0770] 5. An emotion engine that recognizes user emotions and reflects that emotional information in the analysis results.

[0771] 6. A means of extracting user requests and preferences and storing them in a database.

[0772] 7. Means of proposing products and services based on stored information

[0773] 8. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[0774] 9. Means of presenting visual information using a device equipped with a display.

[0775] System operation details

[0776] 1. Obtain user voice input.

[0777] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[0778] 2. Convert voice input to text

[0779] The acquired audio data is converted into text data using speech recognition software. This ensures that the user's spoken content is treated as text data.

[0780] 3. Send the converted text to the server.

[0781] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[0782] 4. Analyze text data and generate a reply.

[0783] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[0784] 5. Emotional analysis using an emotion engine

[0785] The server uses an emotion engine to recognize the user's emotions from voice and text data. The emotion engine analyzes the user's voice tone and text content to determine the user's emotional state. This emotional information is then reflected in the replies and suggestions.

[0786] 6. Send the reply as text data from the server to the terminal.

[0787] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[0788] 7. Convert the reply to speech and output it.

[0789] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[0790] 8. Extract user requests and preferences and save them to a database.

[0791] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[0792] 9. Propose products and services.

[0793] The server uses stored user profile information to run an algorithm that suggests products and services tailored to the user's needs. The suggestions are generated as a reply message and presented to the user.

[0794] 10. Proposals that take past history into consideration

[0795] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[0796] 11. Presentation of visual information

[0797] The terminal uses a display-equipped device to show the user visual information. This includes product images, detailed information, and prices.

[0798] Specific example

[0799] Specific examples of users:

[0800] The user speaks to a smart speaker with a display.

[0801] "I want to travel, could you recommend some places?"

[0802] Processing flow:

[0803] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[0804] 2. Speech-to-text conversion: Speech input is converted into text data.

[0805] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[0806] 4. Sentiment Analysis: The server's emotion engine analyzes the user's voice and text to determine if the user is excited.

[0807] 5. Reflecting Emotional Information: Reflect emotional information in the generated replies and adjust them to have an energetic tone.

[0808] 6. Voice output of replies: The generated text reply is converted to voice and played back from the speaker in an energetic tone.

[0809] 7. Extracting and saving preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[0810] 8. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[0811] 9. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[0812] 10. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[0813] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[0814] The following describes the processing flow.

[0815] Program processing

[0816] Step 1: Obtain user voice input

[0817] Operation details:

[0818] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[0819] Step 2: Convert voice input to text

[0820] Operation details:

[0821] The device converts the acquired audio data into text data using speech recognition software. This ensures that what the user says is processed as text.

[0822] Step 3: Send the converted text to the server.

[0823] Operation details:

[0824] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[0825] Step 4: Analyze the text data and generate a reply.

[0826] Operation details:

[0827] The server analyzes the received text data using natural language processing technology. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[0828] Step 5: Perform emotion analysis using the emotion engine.

[0829] Operation details:

[0830] The server uses an emotion engine to recognize the user's emotions from voice and text data. The emotion engine analyzes the user's voice tone and text content to determine the user's emotional state.

[0831] Step 6: Reflect emotional information in your replies.

[0832] Operation details:

[0833] The server adjusts the content and tone of the generated response based on emotional information obtained from the emotion engine. For example, if the user is excited, the response will be in an energetic tone.

[0834] Step 7: Send the reply from the server to the terminal as text data.

[0835] Operation details:

[0836] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[0837] Step 8: Convert the reply to speech and output it.

[0838] Operation details:

[0839] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[0840] Step 9: Extract user requests and preferences and save them to the database.

[0841] Operation details:

[0842] The server analyzes the user's speech to extract requests and preferences. The extracted information is stored in a database as a user profile.

[0843] Step 10: Propose products and services based on the saved information.

[0844] Operation details:

[0845] The server uses information stored in the database to suggest products and services that meet the user's needs. These suggestions are generated as a reply message and presented to the user.

[0846] Step 11: Make a proposal that takes past history into consideration.

[0847] Operation details:

[0848] The server provides more personalized suggestions based on the user's past conversation and transaction history. This allows for suggestions tailored to the user's preferences and characteristics.

[0849] Step 12: Provide visual information

[0850] Operation details:

[0851] The terminal uses a display-equipped device to show the user visual information. This includes images, details, and prices of the proposed products and services.

[0852] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[0853] (Example 2)

[0854] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0855] Modern users demand fast and personalized product and service recommendations, requiring more accurate suggestions that take their emotional state into account. However, existing systems struggle to recognize user emotions and adjust their recommendations accordingly. Furthermore, they often fail to fully utilize user conversation and transaction history, resulting in a lack of information that users truly need.

[0856] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0857] In this invention, the server includes means for using an emotion engine to recognize the user's emotions, means for sending the generated response from the server to the terminal as text data, and means for making suggestions that take into account past conversation and transaction history in association with user data. This makes it possible to make more personalized product and service suggestions based on the user's emotions and past history.

[0858] "Means for acquiring user voice input" refers to a function that captures voice data spoken by the user using the device's microphone.

[0859] "Methods for converting voice input to text" refers to technologies that use speech recognition software to convert captured voice data into text format.

[0860] "Means for sending converted text to a server" refers to the function by which the terminal sends the text data generated by speech recognition to a server via the network.

[0861] "Means for analyzing text data and generating appropriate replies" refers to a function in which the server analyzes the text data it receives using natural language processing technology and generates an appropriate reply based on its content.

[0862] "Methods of using an emotion engine that recognizes user emotions from voice and text data" refers to a technology in which a server analyzes voice and text data, recognizes the user's emotional state, and incorporates that information into processing.

[0863] "Means for sending generated replies as text data from the server to the terminal" refers to a function for sending replies generated by the server in text format to the terminal.

[0864] "A means of converting replies into speech and outputting them" refers to a function that converts text data received by the device into speech format using speech synthesis technology and plays it back to the user through the speaker.

[0865] "A means of extracting user requests and preferences and storing them in a database" refers to a technology in which a server analyzes the user's speech to extract requests and preferences and stores that information in a database.

[0866] "A means of suggesting products and services based on stored information" refers to a technology in which a server refers to stored user profile information and automatically suggests products and services that are suitable for the user.

[0867] "A means of making suggestions that take into account past conversation and transaction history in relation to user data" refers to a technology in which a server operates an algorithm based on the user's past conversation and transaction history to generate personalized suggestions.

[0868] "A means of providing a user interface through a device with a display and presenting visual information" refers to the function of a terminal that uses its display to present visual information (images, detailed information, price, etc.) to the user.

[0869] This invention combines a system that proposes appropriate products and services through dialogue with the user with an emotion engine that recognizes the user's emotions. This system is realized by combining technologies such as speech recognition, natural language processing, emotion analysis, data analysis, and speech synthesis.

[0870] System-wide configuration

[0871] The system consists of the following main components:

[0872] 1. Means for obtaining user voice input

[0873] 2. Means of converting voice input to text

[0874] 3. Means for sending the converted text to the server

[0875] 4. Means for analyzing text data and generating replies

[0876] 5. An emotion engine that recognizes user emotions and reflects that emotional information in the analysis results.

[0877] 6. Means for sending the generated reply from the server to the terminal.

[0878] 7. A means of converting replies into speech and outputting them.

[0879] 8. A means of extracting user requests and preferences and storing them in a database.

[0880] 9. Means of proposing products and services based on stored information

[0881] 10. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[0882] 11. Means for presenting visual information using a device equipped with a display.

[0883] System operation details

[0884] Get user voice input

[0885] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[0886] Convert voice input to text

[0887] The acquired audio data is converted into text data using speech recognition software. For example, the Google Cloud Speech-to-Text API can be used to treat what the user says as text data.

[0888] Send the converted text to the server.

[0889] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[0890] Analyze text data and generate a reply.

[0891] The server analyzes the received text data using natural language processing techniques. For example, it uses a generative AI model such as OpenAI GPT-4 to generate appropriate responses to user questions and requests.

[0892] Emotional analysis using an emotion engine

[0893] The server uses an emotion engine to recognize user emotions from voice and text data. For example, it uses "IBM Watson Tone Analyzer" to analyze the user's voice tone and text content to determine their emotional state. This emotional information is then reflected in the replies and suggestions.

[0894] Send the generated reply from the server to the terminal.

[0895] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[0896] Convert the reply into speech and output it.

[0897] The device converts received text replies into audio data using speech synthesis technology. For example, the converted audio data, using "Amazon Polly," is played back to the user through the speaker.

[0898] Extract user requests and preferences and save them to a database.

[0899] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[0900] Propose products and services

[0901] The server uses stored user profile information to run an algorithm that suggests products and services tailored to the user's needs. The suggestions are generated as a reply message and presented to the user.

[0902] Proposals that take past history into consideration

[0903] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[0904] Presentation of visual information

[0905] The terminal uses a display-equipped device to show the user visual information. This includes product images, detailed information, and prices.

[0906] Specific example

[0907] Specific examples of users:

[0908] The user speaks to a smart speaker with a display: "I want to travel, could you recommend some places?"

[0909] Processing flow:

[0910] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[0911] 2. Speech-to-text conversion: Speech is converted to text using speech recognition software (e.g., Google Cloud Speech-to-Text).

[0912] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[0913] 4. Sentiment Analysis: An emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's voice and text to determine if the user is excited.

[0914] 5. Reflecting Emotional Information: Reflect emotional information in the generated replies and adjust them to have an energetic tone.

[0915] 6. Voice output of replies: The generated text replies are converted into speech using speech synthesis technology (e.g., Amazon Polly) and played back through the speaker in an energetic tone.

[0916] 7. Extracting and saving preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[0917] 8. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[0918] 9. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[0919] 10. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[0920] Example of a prompt:

[0921] I want to go on a trip, could you recommend some places?

[0922] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[0923] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0924] Step 1:

[0925] The device acquires the user's voice input.

[0926] Input: The microphone captures the user's speech.

[0927] Operation: When the user says, "I want to travel, can you recommend some places?", the device's microphone captures the audio data.

[0928] Output: Captured audio data.

[0929] Step 2:

[0930] The device converts voice input into text.

[0931] Input: Acquired audio data.

[0932] Operation: The device uses the Google Cloud Speech-to-Text API to convert speech data into text data.

[0933] Output: Converted text data (e.g., "I want to travel, could you recommend some places?").

[0934] Step 3:

[0935] The terminal sends the converted text to the server.

[0936] Input: Text data.

[0937] Operation: The terminal sends text data to the server over the network.

[0938] Output: Text data sent to the server.

[0939] Step 4:

[0940] The server analyzes the text data and generates a reply.

[0941] Input: Text data received by the server.

[0942] Operation: The server uses natural language processing techniques (e.g., OpenAI GPT-4) to analyze text data and generate appropriate responses.

[0943] Output: Generated reply text (Example: "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food.").

[0944] Step 5:

[0945] The server uses an emotion engine to perform emotion analysis.

[0946] Input: Audio data or text data.

[0947] Operation: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. For example, it can determine whether the user is excited or not.

[0948] Output: Sentiment analysis results (e.g., user is agitated).

[0949] Step 6:

[0950] The server sends the generated reply as text data to the terminal.

[0951] Input: Generated reply text and sentiment analysis results.

[0952] Operation: The server sends a reply text to the terminal and adjusts it to reflect sentiment information.

[0953] Output: Reply text sent to the terminal.

[0954] Step 7:

[0955] The device converts the reply text into speech and outputs it.

[0956] Input: Reply text data.

[0957] Operation: The device uses speech synthesis technology such as Amazon Polly to convert text into audio data, which is then played through the speaker.

[0958] Output: Audio data that the user hears (e.g., "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food.").

[0959] Step 8:

[0960] The server extracts user requests and preferences and stores them in a database.

[0961] Input: Text data of the user's spoken content.

[0962] Operation: The server analyzes the spoken content, extracts requests and preferences, and stores them in a database. For example, it might extract keywords related to "travel."

[0963] Output: Extracted user profile information.

[0964] Step 9:

[0965] The server suggests products and services based on the information it stores.

[0966] Input: User profile information.

[0967] Operation: The server executes an algorithm that suggests products and services that meet the user's needs.

[0968] Output: A reply message regarding the proposed product or service (e.g., "We have a promotional discount for travel. Are you interested?").

[0969] Step 10:

[0970] The server makes suggestions that take past history into consideration.

[0971] Input: User's past conversation and transaction history.

[0972] Operation: The server generates personalized suggestions based on past conversation and transaction history. For example, it might suggest avoiding places you've visited before.

[0973] Output: Personalized suggestions.

[0974] Step 11:

[0975] The device displays visual information.

[0976] Input: Data about the proposed product or service.

[0977] Operation: The device uses its display to present visual information to the user. This includes displaying product images, detailed information, and prices.

[0978] Output: Visual information displayed on the screen (e.g., photos and prices of suggested travel destinations).

[0979] (Application Example 2)

[0980] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0981] Traditional advertising systems have a problem in that they present uniform ads without considering user emotions, thus failing to maximize advertising effectiveness. Furthermore, they lack sufficient personalized suggestions based on user requests and preferences, and there is a need to improve the user experience. The challenge is to improve advertising effectiveness while increasing user satisfaction by recognizing emotions and presenting timely and appropriate ads.

[0982] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for acquiring the user's voice input, means for converting the voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, an emotion engine means for recognizing the user's emotions and reflecting that emotion information in the response and suggestions, means for extracting the user's requests and preferences and storing them in a database, means for suggesting products and services based on the stored information, and a user interface means for presenting the suggestions in audio and visual ways. This makes it possible to provide personalized advertising suggestions that take the user's emotions into consideration.

[0983] "Acquiring voice input" is the process of capturing the voice spoken by the user through the device's microphone.

[0984] "Converting to text" is the process of converting acquired voice input into text data using a speech recognition algorithm.

[0985] "Text analysis" is the process of analyzing converted text data using natural language processing technology to understand and interpret its content.

[0986] "Generating appropriate responses" is the process of creating the best possible response to a user's question or request based on the results of text analysis.

[0987] "Converting replies to speech" is the process of converting generated text responses into speech data using speech synthesis technology.

[0988] "Audio output" is the process of making audio data audible to the user through an audio device such as a speaker.

[0989] An "emotion engine" is an algorithm or component that analyzes a user's voice or text to determine the user's emotional state (joy, sadness, anxiety, etc.).

[0990] "Extracting requests and preferences" is the process of identifying what an individual wants and likes from their spoken content and recording it in a database.

[0991] "Product and service proposals" refer to the process of selecting the most suitable products and services for a given user based on their saved user data, and then creating a proposal.

[0992] A "user interface" is an interface for a user to interact with a system, and includes means of presenting audio and visual information.

[0993] Overall System Overview

[0994] The system implementing this invention provides personalized advertisements to users via terminals such as smartphones and smart glasses. It acquires the user's voice input and performs a series of processes to present appropriate responses and advertisements that reflect emotions. The system components include voice input means, voice recognition means, text analysis means, emotion analysis means, database storage means, product / service suggestion means, and visual and audio output means. Its specific operation is described below.

[0995] Hardware and software used

[0996] The entire system will be implemented using the following hardware and software:

[0997] A device such as a smartphone or smart glasses (including microphone, speaker, and display)

[0998] Microphone: A device for acquiring voice input.

[0999] Speaker: A device for outputting sound.

[1000] Display: A device for displaying visual information.

[1001] Speech recognition software: Software used to convert acquired speech into text (e.g., Google Speech Recognition API)

[1002] Sentiment analysis model: A model for determining a user's emotions (for example, a sentiment analysis model based on BERT).

[1003] Natural Language Processing Models: Models that analyze user text and generate optimal replies or advertisements (e.g., generative AI models).

[1004] Text-to-speech conversion software (TTS engine)

[1005] Database: Storage for saving user preferences and requests (e.g., SQL database)

[1006] Detailed Operation Description

[1007] 1. Acquisition of voice input:

[1008] The device acquires user voice input through the microphone. When the user speaks, the voice data is acquired and the process proceeds to the next stage.

[1009] 2. Convert voice input to text:

[1010] The acquired audio data is converted into text data using speech recognition software (e.g., Google Speech Recognition API). This ensures that what the user says is treated as text data.

[1011] 3. Analysis of text data:

[1012] The server analyzes text data using natural language processing techniques and generates appropriate responses to user questions and requests. It uses a generative AI model to generate the responses the user expects.

[1013] 4. Emotion analysis:

[1014] The server uses an emotion engine to recognize the user's emotions from text and audio data. It analyzes the tone of voice and the content of speech to determine the user's emotional state. For example, if the user is excited, it will generate an energetic advertisement.

[1015] 5. Replies and ad generation:

[1016] Based on the results of sentiment analysis, ads and replies are generated that match the user's emotional state. For example, if the user is happy, ads with a bright and positive tone are generated.

[1017] 6. Convert replies to speech:

[1018] The server converts the generated text reply into audio data using speech synthesis technology (e.g., a TTS engine). The converted audio data is then played back to the user through the speaker.

[1019] 7. Presentation of visual information:

[1020] The device displays the generated advertisement content on its screen. Visual information such as product images, details, and prices are presented.

[1021] Specific example

[1022] For example, the user might say the following:

[1023] "Please tell me about the latest smartphones."

[1024] The system acquires the audio and generates advertisements through the following process.

[1025] Voice input acquisition

[1026] Text conversion using speech recognition

[1027] Text data analysis and response generation

[1028] Determining a user's happy emotions through emotion analysis.

[1029] Generate an energetic ad ("The latest smartphones have amazing camera features!")

[1030] Convert text replies to speech and output them through the speaker.

[1031] Display visual product information on the screen.

[1032] Example of a prompt

[1033] "The user said, 'Tell me about the latest smartphones.' The user seems very interested and is talking enthusiastically. What kind of ad would be suitable?"

[1034] This prompt allows the generative AI model to generate appropriate advertisements.

[1035] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1036] Step 1: Obtaining voice input

[1037] The user speaks into the device. The device acquires the user's voice input through the microphone. The input is the user's voice data, and the output is this voice data.

[1038] Specific operation: The microphone captures the user's voice and converts it into a digital format.

[1039] Step 2: Convert voice input to text

[1040] The device uses speech recognition software to convert acquired speech data into text data. The input is speech data, and the output is text data.

[1041] Specific operation: Speech recognition software (e.g., Google Speech Recognition API) analyzes the audio data and generates a corresponding string.

[1042] Step 3: Analyzing text data

[1043] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, it generates an appropriate response. The input is text data, and the output is the analysis results and the generated response.

[1044] Specific operation: A natural language processing model (e.g., a generative AI model) analyzes text data and generates a contextually appropriate response.

[1045] Step 4: Emotion Analysis

[1046] The server uses an emotion engine to recognize the user's emotions from text and audio data. The input is text and audio data, and the output is the emotion analysis result.

[1047] Specific operation: The emotion analysis model analyzes the text content and voice tone to determine the user's emotional state (e.g., joy, sadness, excitement, etc.).

[1048] Step 5: Replies and ad generation

[1049] The server considers the results of sentiment analysis to generate advertisements and replies appropriate to the user's emotional state. The input is the analysis results and sentiment analysis results, and the output is the generated replies and advertisements.

[1050] Specific operation: The generative AI model creates advertising content that reflects emotional information. For example, if the user is excited, it will generate an advertisement in an energetic tone.

[1051] Step 6: Convert your reply to speech

[1052] The server converts the generated text replies into speech data using speech synthesis technology. The input is text data, and the output is speech data.

[1053] Specific operation: A speech synthesis engine (e.g., a TTS engine) synthesizes text data and outputs it as natural-sounding speech.

[1054] Step 7: Audio Output

[1055] The device plays the converted audio data to the user through its speaker. The input is audio data, and the output is the audio that the user hears.

[1056] Specific operation: The speaker plays the audio data as a machine-generated voice.

[1057] Step 8: Presenting visual information

[1058] The device uses its display to visually show the generated advertisement content. The input is the advertisement content, and the output is the visual information displayed on the screen.

[1059] Specific operation: The display shows advertising images and detailed information, providing the user with visual feedback.

[1060] Step 9: Extract and save requests and preferences

[1061] The server extracts requests and preferences from the user's speech and stores them in a database. The input is text data, and the output is a user profile stored in the database.

[1062] Specific operation: A natural language processing model extracts important keywords and phrases from text and records them in a database.

[1063] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1064] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1065] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1066] [Third Embodiment]

[1067] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1068] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1069] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1070] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1071] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1072] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1073] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1074] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1075] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1076] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1077] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1078] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1079] This invention is a system for suggesting appropriate products and services through dialogue with users. This system is realized by combining technologies such as speech recognition, natural language processing, data analysis, and speech synthesis.

[1080] System-wide configuration

[1081] The system consists of the following main components:

[1082] 1. Means for obtaining user voice input

[1083] 2. Means of converting voice input to text

[1084] 3. Means for analyzing the converted text and generating an appropriate response.

[1085] 4. A means of converting replies into speech and outputting them.

[1086] 5. A means of extracting user requests and preferences and storing them in a database.

[1087] 6. Means of proposing products and services based on stored information

[1088] 7. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[1089] 8. Means of presenting visual information using a device equipped with a display.

[1090] System operation details

[1091] 1. Obtain user voice input.

[1092] The device acquires user voice input through its microphone. When the user speaks into the device, the microphone captures the audio data.

[1093] 2. Convert voice input to text

[1094] The acquired audio data is converted into text using speech recognition software. This ensures that what the user says is treated as text data.

[1095] 3. Analyze the converted text and generate a reply.

[1096] The server analyzes text data using natural language processing technology. Based on the analyzed data, an AI model generates appropriate responses. For example, in response to a question like "I want to travel," it generates a response suggesting a suitable travel destination.

[1097] 4. Convert the reply to speech and output it.

[1098] The generated text reply is converted into speech using speech synthesis technology. The converted speech data is then played back to the user through the device's speaker.

[1099] 5. Extract requests and preferences and save them to a database.

[1100] The server analyzes the user's speech to extract their requests and preferences. The extracted information is stored in the database as a user profile.

[1101] 6. Propose products and services

[1102] The server uses stored user profile information to run an algorithm that suggests appropriate products and services. The suggested products and services are presented to the user in text or voice.

[1103] 7. Proposals that take past history into consideration

[1104] The server provides more personalized suggestions based on the user's past conversation and transaction history. This makes it possible to offer suggestions that match the user's preferences.

[1105] 8. Presentation of visual information

[1106] Using devices equipped with displays, visual information is also provided to the user. This includes product images and detailed information.

[1107] Specific example

[1108] Specific examples of users:

[1109] The user speaks to a smart speaker with a display.

[1110] "I want to travel, could you recommend some places?"

[1111] Processing flow:

[1112] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[1113] 2. Speech-to-text conversion: Speech input is converted into text data.

[1114] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[1115] 4. Audio output of replies: The generated text reply is converted to audio and played through the speaker.

[1116] 5. Extraction and saving of preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[1117] 6. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[1118] 7. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[1119] 8. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[1120] The above describes a specific form for carrying out the invention. This system allows users to receive personalized product and service suggestions while engaging in voice-based dialogue.

[1121] The following describes the processing flow.

[1122] Program processing

[1123] Step 1: Obtain user voice input

[1124] Operation details:

[1125] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[1126] Step 2: Convert voice input to text

[1127] Operation details:

[1128] The device converts the acquired audio data into text data using speech recognition software. This ensures that what the user says is processed as text.

[1129] Step 3: Send the converted text to the server.

[1130] Operation details:

[1131] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[1132] Step 4: Analyze the text data and generate an appropriate response.

[1133] Operation details:

[1134] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[1135] Step 5: Send the reply from the server to the terminal as text data.

[1136] Operation details:

[1137] The server sends the generated text reply to the terminal. This text data is converted to speech and played back on the terminal.

[1138] Step 6: Convert the reply to speech and output it.

[1139] Operation details:

[1140] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[1141] Step 7: Extract user requests and preferences and save them to the database.

[1142] Operation details:

[1143] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[1144] Step 8: Propose products and services based on the saved information.

[1145] Operation details:

[1146] The server uses information stored in the database to suggest products and services that meet the user's needs. These suggestions are generated as a reply message and presented to the user.

[1147] Step 9: Consider past conversation and transaction history in relation to user data.

[1148] Operation details:

[1149] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[1150] Step 10: Provide visual information

[1151] Operation details:

[1152] The terminal uses a display-equipped device to show the user visual information. This includes images, details, and prices of the proposed products and services.

[1153] The above outlines the specific processing steps of the program in this system. This allows users to receive personalized product and service suggestions while engaging in voice-based conversations.

[1154] (Example 1)

[1155] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1156] Modern users handle vast amounts of information and demand personalized services and product recommendations, but traditional systems struggle to efficiently achieve this. In particular, few systems can start with voice input and then provide recommendations that take into account the user's past history and preferences. Therefore, there is a need to develop systems that improve user satisfaction.

[1157] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1158] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text, means for analyzing the converted text and generating an appropriate response using a generation AI model, means for converting the response into voice and outputting it, means for extracting the user's requests and preferences and storing them in a database, and means for suggesting products and services based on the stored information. This makes it possible to efficiently provide personalized suggestions that start with voice input and take into account the user's past history and preferences.

[1159] "Means for acquiring user voice input" refers to a function in which the terminal uses a microphone to capture the user's voice in real time and sends it to the server as voice data.

[1160] "Means of converting voice input to text" refers to a function in which a server uses speech recognition technology to convert acquired voice data into text data.

[1161] "Means for analyzing converted text and generating appropriate responses using a generative AI model" refers to a function in which the server uses natural language processing technology and a generative AI model to analyze text data and automatically generate appropriate responses to user questions and requests.

[1162] "Method for converting replies into audio and outputting them" refers to a function in which the server uses speech synthesis technology to convert text replies into audio data and plays it back to the user through the terminal's speaker.

[1163] "A means of extracting user requests and preferences and saving them to a database" refers to a function where the server analyzes the user's speech and saves information about their requests and preferences to a database.

[1164] "A means of suggesting products and services based on stored information" refers to a function in which the server uses an algorithm to suggest appropriate products and services based on user information stored in the database.

[1165] "A means of making suggestions that take into account past conversation and transaction history in relation to user data" refers to a function in which the server refers to the user's past conversation and transaction history and makes personalized suggestions based on that.

[1166] "Means of providing a user interface through a display device and presenting visual information" refers to the function of a terminal that uses its display to provide the user with visual information (for example, detailed information or images of products or services).

[1167] Modes for carrying out the invention

[1168] This invention is a system for proposing appropriate products and services through interaction with the user. This system is realized by combining the following hardware and software technologies.

[1169] System-wide configuration

[1170] The system consists of the following main components:

[1171] 1. Means for obtaining user voice input

[1172] 2. Means of converting voice input to text

[1173] 3. A means of analyzing the converted text and generating an appropriate response using a generative AI model.

[1174] 4. A means of converting replies into speech and outputting them.

[1175] 5. A means of extracting user requests and preferences and storing them in a database.

[1176] 6. Means of proposing products and services based on stored information

[1177] 7. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[1178] 8. Means of presenting visual information through a display device.

[1179] Acquiring voice input

[1180] The device uses a microphone to acquire user voice input. When the user speaks into the device, the microphone captures the voice data and sends it to the server as digital data. This data is processed using speech recognition technology.

[1181] Speech-to-text conversion

[1182] The server uses speech recognition software (e.g., a speech recognition API) to convert the acquired speech data into text. This allows the speech input to be treated as text data.

[1183] Text analysis and response generation

[1184] The converted text data is analyzed by the server using natural language processing techniques and generative AI models (e.g., large-scale language models). Based on the analyzed data, an appropriate response is generated. For example, if a user says "I want to travel," the generative AI model is given the prompt "The user says they want to travel. What travel destination would you suggest?" and the model generates a response such as "Hokkaido is a recommended travel destination."

[1185] Voice output of the reply

[1186] The generated text reply is converted into audio data using speech synthesis technology (e.g., a speech synthesis API). The converted audio data is played back to the user through the device's speaker. The user can then hear the response, such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[1187] Extraction of requests and preferences and storage in a database.

[1188] The server analyzes the user's speech to extract their requests and preferences. Customer insight tools are used for this purpose. The extracted information is stored in a database as a user profile. For example, if a user is identified as being interested in "travel," that information is saved.

[1189] Product and service proposals

[1190] The server uses stored user profile information to execute algorithms (e.g., collaborative filtering) to suggest appropriate products and services. This information is presented to the user in text or audio format.

[1191] Proposals that take past history into consideration

[1192] The server references the user's past conversation and transaction history to provide more personalized suggestions. For example, it can pique the user's interest by prioritizing travel destinations they have never visited before.

[1193] Presentation of visual information

[1194] If the device has a display, it will use the display to provide visual information about the suggested products and services. This includes detailed product information and photos. Users can visually confirm the information, leading to a deeper understanding.

[1195] Specific example

[1196] Specific examples of users:

[1197] "I want to travel, could you recommend some places?"

[1198] Processing flow:

[1199] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[1200] 2. Speech-to-text conversion: Speech input is converted into text data using a speech recognition API.

[1201] 3. Text analysis and reply generation: The server uses a generated AI model to generate the reply "Hokkaido is a recommended travel destination."

[1202] 4. Audio output of replies: The generated text reply is converted into audio data using a speech synthesis API and played back through the device's speaker.

[1203] 5. Extraction and saving of preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[1204] 6. Proposal: The server uses collaborative filtering to suggest services and products related to the travel destination.

[1205] 7. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[1206] 8. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the display device.

[1207] As described above, the system of the present invention allows users to receive personalized product and service suggestions through voice and visual information. This enables users to efficiently acquire information and make choices that meet their needs.

[1208] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1209] Step 1:

[1210] The device acquires the user's voice input. When the user speaks into the device, saying "I want to travel," the device's microphone captures the voice. The input is an analog audio signal, which the microphone sensor converts into digital data. Digital audio data is generated, and this becomes the initial output.

[1211] Step 2:

[1212] The server receives the digital audio data obtained in step 1 and uses speech recognition software (e.g., a speech recognition API) to convert the audio into text. Specifically, it converts the audio data "I want to go on a trip" into text data "Text:I want to go on a trip". This converted text data becomes the output.

[1213] Step 3:

[1214] The server, upon receiving the converted text data, analyzes it using natural language processing techniques. This analysis utilizes a generative AI model (e.g., a large-scale language model). The prompt "The user says they want to travel. What travel destination would you suggest?" is input to the generative AI model, and the model generates an appropriate response, "Hokkaido is a recommended travel destination." The analyzed data is then output.

[1215] Step 4:

[1216] The server converts the text reply generated in step 3 into audio data using speech synthesis technology (e.g., a speech synthesis API). Specifically, it converts "My recommended travel destination is Hokkaido" into audio data and sends that audio data to the terminal. This audio data becomes the output of step 4.

[1217] Step 5:

[1218] The device receives audio data sent from the server and plays the audio to the user through its speaker. The user hears the audio say, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food." Through this action, the user obtains the necessary information.

[1219] Step 6:

[1220] Based on what the user says, the server automatically analyzes the data and extracts the user's requests and preferences. For example, keywords such as "travel" and "nature" may be extracted. The extracted information is stored in the database as a user profile. This stored user profile data is then output.

[1221] Step 7:

[1222] The server executes data analysis algorithms (e.g., collaborative filtering) based on user profile information stored in the database, and suggests appropriate products and services. For example, it might suggest accommodations or tour packages related to a travel destination. The content of these suggestions becomes the output.

[1223] Step 8:

[1224] The server considers the user's past conversation and transaction history to provide more personalized product and service suggestions. It also uses past data to suggest travel destinations the user has never visited before. This personalized suggestion information is then output.

[1225] Step 9:

[1226] If the device has a display, the server also provides visual information. For example, it might display detailed information or photos of a suggested travel destination. The user can visually confirm the displayed information. This visual information becomes the final output.

[1227] Through the steps outlined above, this system comprehensively achieves everything from voice input, text analysis, appropriate response generation using a generative AI model, speech synthesis, extraction and storage of user preferences, suggestions based on data analysis, and even the provision of visual information.

[1228] (Application Example 1)

[1229] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1230] In modern autonomous vehicles, there is a need for systems that suggest appropriate destinations and sightseeing spots in real time to enhance the passenger experience. However, conventional in-car entertainment and navigation systems have the problem of not being able to make suggestions that meet the individual needs and preferences of users. Furthermore, there has been a challenge in providing more personalized suggestions that take past history into account.

[1231] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1232] In this invention, the server includes means for acquiring user voice input, means for converting voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, means for extracting user requests and preferences and storing them in a database, means for suggesting products and services based on the stored information, and means for suggesting appropriate drive destinations and sightseeing spots to passengers in the vehicle. This enables real-time appropriate suggestions that meet the requests and preferences of passengers.

[1233] "Acquiring user voice input" means that microphones installed inside the vehicle capture the voices of passengers.

[1234] "Converting voice input to text" means converting acquired voice data into text data using speech recognition technology.

[1235] "Analyzing the converted text and generating appropriate responses" means using natural language processing technology to analyze text data and generate appropriate information and suggestions for passengers.

[1236] "Converting replies to audio and outputting them" means converting the generated replies into an audio format using speech synthesis technology and playing them through the car's speakers.

[1237] "Extracting user requests and preferences and storing them in a database" means extracting requests and preferences from passengers' statements and storing them in a database.

[1238] "Proposing products and services based on stored information" means selecting and proposing appropriate products and services based on customer needs and preferences stored in a database.

[1239] "Suggesting appropriate driving destinations and sightseeing spots to passengers inside the vehicle" means suggesting appropriate driving destinations and sightseeing spots in real time through dialogue with passengers inside the autonomous vehicle.

[1240] This invention relates to a system for suggesting appropriate driving destinations and sightseeing spots to passengers in an autonomous vehicle. This system is implemented using the following key hardware and software components, combining speech recognition, natural language processing, speech synthesis, and database technologies.

[1241] System Configuration

[1242] 1. Hardware Configuration

[1243] Microphone: A high-sensitivity microphone installed inside the vehicle is used. This microphone is used to capture passenger voice input. For example, the SureBity USB microphone is suitable for this purpose.

[1244] Speakers: Audio is output using the in-vehicle speakers. Voice replies to passengers are also delivered through these speakers.

[1245] Display-equipped devices: Display devices installed inside the vehicle are used to provide visual information.

[1246] 2. Software Configuration

[1247] Speech recognition software: Using the speech_recognition library, audio data acquired from the microphone is converted into text data.

[1248] Natural language processing software: Using the pipeline function of the transformers library, the transformed text is analyzed to generate appropriate responses. In particular, "gpt-3" is used as the generative AI model.

[1249] Text-to-speech software: Using the gTTS (Google Text-to-Speech) library, the generated text replies are converted into audio data and played back through the car's speakers.

[1250] Database: A database system used to store user requests and preferences. It stores past conversation and transaction history and is used to provide more personalized suggestions.

[1251] System operation

[1252] The operation of this system is described as follows:

[1253] 1. Speech acquisition and recognition

[1254] The terminal acquires the passenger's voice input through the microphone. Voice recognition software converts this into text data and sends it to the server.

[1255] 2. Natural Language Processing and Proposal Generation

[1256] The server analyzes text data using natural language processing techniques and generates appropriate suggestions. These suggestions are then converted into speech using speech synthesis software.

[1257] 3. Audio output and data storage

[1258] The converted audio data is played back to passengers through the in-car speakers. At the same time, information about the suggestions, as well as passenger requests and preferences, is stored in a database and used to provide more personalized suggestions in the future.

[1259] 4. Visual information provision

[1260] Information on suggested driving destinations and tourist spots is presented visually through the in-car display system. This allows passengers to understand the suggestions more concretely.

[1261] Examples and prompts for generative AI models

[1262] Specific example:

[1263] When a user says, "I'm looking for a good restaurant," the speech recognition system converts this into text and generates a prompt message like the following, which is then input into the AI ​​model.

[1264] Prompts for the generative AI model:

[1265] To make the most of our drive, could you recommend some good restaurants?

[1266] The implementation of this system will provide a comfortable and personalized experience within autonomous vehicles. In the concrete implementation of the invention, the aforementioned hardware and software will be integrated to enable real-time and highly accurate suggestions.

[1267] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1268] Step 1:

[1269] The user speaks into a microphone installed inside the vehicle. The terminal acquires the user's voice input. The input is the passenger's voice data, and the output is a captured raw audio file. This audio data is then passed on to subsequent speech recognition processing.

[1270] Step 2:

[1271] The device converts the acquired audio data into text data using speech recognition software (speech_recognition library). The input is the audio file acquired in step 1, and the output is text data. Specifically, the speech recognition software analyzes the audio waveform and converts it into a string based on a language model.

[1272] Step 3:

[1273] The server analyzes the converted text using natural language processing software (the pipeline function of the transformers library) and generates an appropriate response. The input is the text data obtained in step 2, and the output is a text response containing the suggested content. A generative AI model (e.g., gpt-3) is used to analyze the text data and generate prompt sentences that respond to the user's questions and requests.

[1274] Step 4:

[1275] The server converts the generated text reply into audio data using speech synthesis technology (gTTS library). The input is the text reply generated in step 3, and the output is an audio file. Specifically, the speech synthesis technology converts the text into an audio signal and saves it as an MP3 file.

[1276] Step 5:

[1277] The terminal plays the converted audio data to the user through the in-car speakers. The input is the audio file generated in step 4, and the output is the presentation to the user as audio. The audio is played using software that plays MP3 files (e.g., mpg321).

[1278] Step 6:

[1279] The server extracts user requests and preferences and stores them in a database. The input is the text data obtained in step 3, and the output is a new record in the user profile database. Natural language processing techniques are used to extract information about requests and preferences from the text, organize it, and add it to the database.

[1280] Step 7:

[1281] The server executes an algorithm that suggests products and services based on stored information. The input is user profile information from the database, and the output is text data containing new suggestions. These suggestions are then provided to the user in subsequent speech recognition and natural language processing steps.

[1282] Step 8:

[1283] The server or terminal visually presents information about suggested drive destinations and tourist spots via an in-car display device. The input is the suggested content obtained in step 7, and the output is the visual information displayed on the display. Image data and map data are used to visually represent the suggested content.

[1284] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1285] This invention combines a system that proposes appropriate products and services through dialogue with the user with an emotion engine that recognizes the user's emotions. This system is realized by combining technologies such as speech recognition, natural language processing, emotion analysis, data analysis, and speech synthesis.

[1286] System-wide configuration

[1287] The system consists of the following main components:

[1288] 1. Means for obtaining user voice input

[1289] 2. Means of converting voice input to text

[1290] 3. Means for analyzing the converted text and generating an appropriate response.

[1291] 4. A means of converting replies into speech and outputting them.

[1292] 5. An emotion engine that recognizes user emotions and reflects that emotional information in the analysis results.

[1293] 6. A means of extracting user requests and preferences and storing them in a database.

[1294] 7. Means of proposing products and services based on stored information

[1295] 8. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[1296] 9. Means of presenting visual information using a device equipped with a display.

[1297] System operation details

[1298] 1. Obtain user voice input.

[1299] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[1300] 2. Convert voice input to text

[1301] The acquired audio data is converted into text data using speech recognition software. This ensures that the user's spoken content is treated as text data.

[1302] 3. Send the converted text to the server.

[1303] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[1304] 4. Analyze text data and generate a reply.

[1305] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[1306] 5. Emotional analysis using an emotion engine

[1307] The server uses an emotion engine to recognize the user's emotions from voice and text data. The emotion engine analyzes the user's voice tone and text content to determine the user's emotional state. This emotional information is then reflected in the replies and suggestions.

[1308] 6. Send the reply as text data from the server to the terminal.

[1309] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[1310] 7. Convert the reply to speech and output it.

[1311] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[1312] 8. Extract user requests and preferences and save them to a database.

[1313] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[1314] 9. Propose products and services.

[1315] The server uses stored user profile information to run an algorithm that suggests products and services tailored to the user's needs. The suggestions are generated as a reply message and presented to the user.

[1316] 10. Proposals that take past history into consideration

[1317] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[1318] 11. Presentation of visual information

[1319] The terminal uses a display-equipped device to show the user visual information. This includes product images, detailed information, and prices.

[1320] Specific example

[1321] Specific examples of users:

[1322] The user speaks to a smart speaker with a display.

[1323] "I want to travel, could you recommend some places?"

[1324] Processing flow:

[1325] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[1326] 2. Speech-to-text conversion: Speech input is converted into text data.

[1327] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[1328] 4. Sentiment Analysis: The server's emotion engine analyzes the user's voice and text to determine if the user is excited.

[1329] 5. Reflecting Emotional Information: Reflect emotional information in the generated replies and adjust them to have an energetic tone.

[1330] 6. Voice output of replies: The generated text reply is converted to voice and played back from the speaker in an energetic tone.

[1331] 7. Extracting and saving preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[1332] 8. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[1333] 9. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[1334] 10. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[1335] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[1336] The following describes the processing flow.

[1337] Program processing

[1338] Step 1: Obtain user voice input

[1339] Operation details:

[1340] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[1341] Step 2: Convert voice input to text

[1342] Operation details:

[1343] The device converts the acquired audio data into text data using speech recognition software. This ensures that what the user says is processed as text.

[1344] Step 3: Send the converted text to the server.

[1345] Operation details:

[1346] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[1347] Step 4: Analyze the text data and generate a reply.

[1348] Operation details:

[1349] The server analyzes the received text data using natural language processing technology. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[1350] Step 5: Perform emotion analysis using the emotion engine.

[1351] Operation details:

[1352] The server uses an emotion engine to recognize the user's emotions from voice and text data. The emotion engine analyzes the user's voice tone and text content to determine the user's emotional state.

[1353] Step 6: Reflect emotional information in your replies.

[1354] Operation details:

[1355] The server adjusts the content and tone of the generated response based on emotional information obtained from the emotion engine. For example, if the user is excited, the response will be in an energetic tone.

[1356] Step 7: Send the reply from the server to the terminal as text data.

[1357] Operation details:

[1358] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[1359] Step 8: Convert the reply to speech and output it.

[1360] Operation details:

[1361] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[1362] Step 9: Extract user requests and preferences and save them to the database.

[1363] Operation details:

[1364] The server analyzes the user's speech to extract requests and preferences. The extracted information is stored in a database as a user profile.

[1365] Step 10: Propose products and services based on the saved information.

[1366] Operation details:

[1367] The server uses information stored in the database to suggest products and services that meet the user's needs. These suggestions are generated as a reply message and presented to the user.

[1368] Step 11: Make a proposal that takes past history into consideration.

[1369] Operation details:

[1370] The server provides more personalized suggestions based on the user's past conversation and transaction history. This allows for suggestions tailored to the user's preferences and characteristics.

[1371] Step 12: Provide visual information

[1372] Operation details:

[1373] The terminal uses a display-equipped device to show the user visual information. This includes images, details, and prices of the proposed products and services.

[1374] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[1375] (Example 2)

[1376] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1377] Modern users demand fast and personalized product and service recommendations, requiring more accurate suggestions that take their emotional state into account. However, existing systems struggle to recognize user emotions and adjust their recommendations accordingly. Furthermore, they often fail to fully utilize user conversation and transaction history, resulting in a lack of information that users truly need.

[1378] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1379] In this invention, the server includes means for using an emotion engine to recognize the user's emotions, means for sending the generated response from the server to the terminal as text data, and means for making suggestions that take into account past conversation and transaction history in association with user data. This makes it possible to make more personalized product and service suggestions based on the user's emotions and past history.

[1380] "Means for acquiring user voice input" refers to a function that captures voice data spoken by the user using the device's microphone.

[1381] "Methods for converting voice input to text" refers to technologies that use speech recognition software to convert captured voice data into text format.

[1382] "Means for sending converted text to a server" refers to the function by which the terminal sends the text data generated by speech recognition to a server via the network.

[1383] "Means for analyzing text data and generating appropriate replies" refers to a function in which the server analyzes the text data it receives using natural language processing technology and generates an appropriate reply based on its content.

[1384] "Methods of using an emotion engine that recognizes user emotions from voice and text data" refers to a technology in which a server analyzes voice and text data, recognizes the user's emotional state, and incorporates that information into processing.

[1385] "Means for sending generated replies as text data from the server to the terminal" refers to a function for sending replies generated by the server in text format to the terminal.

[1386] "A means of converting replies into speech and outputting them" refers to a function that converts text data received by the device into speech format using speech synthesis technology and plays it back to the user through the speaker.

[1387] "A means of extracting user requests and preferences and storing them in a database" refers to a technology in which a server analyzes the user's speech to extract requests and preferences and stores that information in a database.

[1388] "A means of suggesting products and services based on stored information" refers to a technology in which a server refers to stored user profile information and automatically suggests products and services that are suitable for the user.

[1389] "A means of making suggestions that take into account past conversation and transaction history in relation to user data" refers to a technology in which a server operates an algorithm based on the user's past conversation and transaction history to generate personalized suggestions.

[1390] "A means of providing a user interface through a device with a display and presenting visual information" refers to the function of a terminal that uses its display to present visual information (images, detailed information, price, etc.) to the user.

[1391] This invention combines a system that proposes appropriate products and services through dialogue with the user with an emotion engine that recognizes the user's emotions. This system is realized by combining technologies such as speech recognition, natural language processing, emotion analysis, data analysis, and speech synthesis.

[1392] System-wide configuration

[1393] The system consists of the following main components:

[1394] 1. Means for obtaining user voice input

[1395] 2. Means of converting voice input to text

[1396] 3. Means for sending the converted text to the server

[1397] 4. Means for analyzing text data and generating replies

[1398] 5. An emotion engine that recognizes user emotions and reflects that emotional information in the analysis results.

[1399] 6. Means for sending the generated reply from the server to the terminal.

[1400] 7. A means of converting replies into speech and outputting them.

[1401] 8. A means of extracting user requests and preferences and storing them in a database.

[1402] 9. Means of proposing products and services based on stored information

[1403] 10. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[1404] 11. Means for presenting visual information using a device equipped with a display.

[1405] System operation details

[1406] Get user voice input

[1407] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[1408] Convert voice input to text

[1409] The acquired audio data is converted into text data using speech recognition software. For example, the Google Cloud Speech-to-Text API can be used to treat what the user says as text data.

[1410] Send the converted text to the server.

[1411] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[1412] Analyze text data and generate a reply.

[1413] The server analyzes the received text data using natural language processing techniques. For example, it uses a generative AI model such as OpenAI GPT-4 to generate appropriate responses to user questions and requests.

[1414] Emotional analysis using an emotion engine

[1415] The server uses an emotion engine to recognize user emotions from voice and text data. For example, it uses "IBM Watson Tone Analyzer" to analyze the user's voice tone and text content to determine their emotional state. This emotional information is then reflected in the replies and suggestions.

[1416] Send the generated reply from the server to the terminal.

[1417] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[1418] Convert the reply into speech and output it.

[1419] The device converts received text replies into audio data using speech synthesis technology. For example, the converted audio data, using "Amazon Polly," is played back to the user through the speaker.

[1420] Extract user requests and preferences and save them to a database.

[1421] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[1422] Propose products and services

[1423] The server uses stored user profile information to run an algorithm that suggests products and services tailored to the user's needs. The suggestions are generated as a reply message and presented to the user.

[1424] Proposals that take past history into consideration

[1425] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[1426] Presentation of visual information

[1427] The terminal uses a display-equipped device to show the user visual information. This includes product images, detailed information, and prices.

[1428] Specific example

[1429] Specific examples of users:

[1430] The user speaks to a smart speaker with a display: "I want to travel, could you recommend some places?"

[1431] Processing flow:

[1432] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[1433] 2. Speech-to-text conversion: Speech is converted to text using speech recognition software (e.g., Google Cloud Speech-to-Text).

[1434] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[1435] 4. Sentiment Analysis: An emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's voice and text to determine if the user is excited.

[1436] 5. Reflecting Emotional Information: Reflect emotional information in the generated replies and adjust them to have an energetic tone.

[1437] 6. Voice output of replies: The generated text replies are converted into speech using speech synthesis technology (e.g., Amazon Polly) and played back through the speaker in an energetic tone.

[1438] 7. Extracting and saving preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[1439] 8. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[1440] 9. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[1441] 10. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[1442] Example of a prompt:

[1443] I want to go on a trip, could you recommend some places?

[1444] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[1445] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1446] Step 1:

[1447] The device acquires the user's voice input.

[1448] Input: The microphone captures the user's speech.

[1449] Operation: When the user says, "I want to travel, can you recommend some places?", the device's microphone captures the audio data.

[1450] Output: Captured audio data.

[1451] Step 2:

[1452] The device converts voice input into text.

[1453] Input: Acquired audio data.

[1454] Operation: The device uses the Google Cloud Speech-to-Text API to convert speech data into text data.

[1455] Output: Converted text data (e.g., "I want to travel, could you recommend some places?").

[1456] Step 3:

[1457] The terminal sends the converted text to the server.

[1458] Input: Text data.

[1459] Operation: The terminal sends text data to the server over the network.

[1460] Output: Text data sent to the server.

[1461] Step 4:

[1462] The server analyzes the text data and generates a reply.

[1463] Input: Text data received by the server.

[1464] Operation: The server uses natural language processing techniques (e.g., OpenAI GPT-4) to analyze text data and generate appropriate responses.

[1465] Output: Generated reply text (Example: "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food.").

[1466] Step 5:

[1467] The server uses an emotion engine to perform emotion analysis.

[1468] Input: Audio data or text data.

[1469] Operation: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. For example, it can determine whether the user is excited or not.

[1470] Output: Sentiment analysis results (e.g., user is agitated).

[1471] Step 6:

[1472] The server sends the generated reply as text data to the terminal.

[1473] Input: Generated reply text and sentiment analysis results.

[1474] Operation: The server sends a reply text to the terminal and adjusts it to reflect sentiment information.

[1475] Output: Reply text sent to the terminal.

[1476] Step 7:

[1477] The device converts the reply text into speech and outputs it.

[1478] Input: Reply text data.

[1479] Operation: The device uses speech synthesis technology such as Amazon Polly to convert text into audio data, which is then played through the speaker.

[1480] Output: Audio data that the user hears (e.g., "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food.").

[1481] Step 8:

[1482] The server extracts user requests and preferences and stores them in a database.

[1483] Input: Text data of the user's spoken content.

[1484] Operation: The server analyzes the spoken content, extracts requests and preferences, and stores them in a database. For example, it might extract keywords related to "travel."

[1485] Output: Extracted user profile information.

[1486] Step 9:

[1487] The server suggests products and services based on the information it stores.

[1488] Input: User profile information.

[1489] Operation: The server executes an algorithm that suggests products and services that meet the user's needs.

[1490] Output: A reply message regarding the proposed product or service (e.g., "We have a promotional discount for travel. Are you interested?").

[1491] Step 10:

[1492] The server makes suggestions that take past history into consideration.

[1493] Input: User's past conversation and transaction history.

[1494] Operation: The server generates personalized suggestions based on past conversation and transaction history. For example, it might suggest avoiding places you've visited before.

[1495] Output: Personalized suggestions.

[1496] Step 11:

[1497] The device displays visual information.

[1498] Input: Data about the proposed product or service.

[1499] Operation: The device uses its display to present visual information to the user. This includes displaying product images, detailed information, and prices.

[1500] Output: Visual information displayed on the screen (e.g., photos and prices of suggested travel destinations).

[1501] (Application Example 2)

[1502] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1503] Traditional advertising systems have a problem in that they present uniform ads without considering user emotions, thus failing to maximize advertising effectiveness. Furthermore, they lack sufficient personalized suggestions based on user requests and preferences, and there is a need to improve the user experience. The challenge is to improve advertising effectiveness while increasing user satisfaction by recognizing emotions and presenting timely and appropriate ads.

[1504] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for acquiring the user's voice input, means for converting the voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, an emotion engine means for recognizing the user's emotions and reflecting that emotion information in the response and suggestions, means for extracting the user's requests and preferences and storing them in a database, means for suggesting products and services based on the stored information, and a user interface means for presenting the suggestions in audio and visual ways. This makes it possible to provide personalized advertising suggestions that take the user's emotions into consideration.

[1505] "Acquiring voice input" is the process of capturing the voice spoken by the user through the device's microphone.

[1506] "Converting to text" is the process of converting acquired voice input into text data using a speech recognition algorithm.

[1507] "Text analysis" is the process of analyzing converted text data using natural language processing technology to understand and interpret its content.

[1508] "Generating appropriate responses" is the process of creating the best possible response to a user's question or request based on the results of text analysis.

[1509] "Converting replies to speech" is the process of converting generated text responses into speech data using speech synthesis technology.

[1510] "Audio output" is the process of making audio data audible to the user through an audio device such as a speaker.

[1511] An "emotion engine" is an algorithm or component that analyzes a user's voice or text to determine the user's emotional state (joy, sadness, anxiety, etc.).

[1512] "Extracting requests and preferences" is the process of identifying what an individual wants and likes from their spoken content and recording it in a database.

[1513] "Product and service proposals" refer to the process of selecting the most suitable products and services for a given user based on their saved user data, and then creating a proposal.

[1514] A "user interface" is an interface for a user to interact with a system, and includes means of presenting audio and visual information.

[1515] Overall System Overview

[1516] The system implementing this invention provides personalized advertisements to users via terminals such as smartphones and smart glasses. It acquires the user's voice input and performs a series of processes to present appropriate responses and advertisements that reflect emotions. The system components include voice input means, voice recognition means, text analysis means, emotion analysis means, database storage means, product / service suggestion means, and visual and audio output means. Its specific operation is described below.

[1517] Hardware and software used

[1518] The entire system will be implemented using the following hardware and software:

[1519] A device such as a smartphone or smart glasses (including microphone, speaker, and display)

[1520] Microphone: A device for acquiring voice input.

[1521] Speaker: A device for outputting sound.

[1522] Display: A device for displaying visual information.

[1523] Speech recognition software: Software used to convert acquired speech into text (e.g., Google Speech Recognition API)

[1524] Sentiment analysis model: A model for determining a user's emotions (for example, a sentiment analysis model based on BERT).

[1525] Natural Language Processing Models: Models that analyze user text and generate optimal replies or advertisements (e.g., generative AI models).

[1526] Text-to-speech conversion software (TTS engine)

[1527] Database: Storage for saving user preferences and requests (e.g., SQL database)

[1528] Detailed Operation Description

[1529] 1. Acquisition of voice input:

[1530] The device acquires user voice input through the microphone. When the user speaks, the voice data is acquired and the process proceeds to the next stage.

[1531] 2. Convert voice input to text:

[1532] The acquired audio data is converted into text data using speech recognition software (e.g., Google Speech Recognition API). This ensures that what the user says is treated as text data.

[1533] 3. Analysis of text data:

[1534] The server analyzes text data using natural language processing techniques and generates appropriate responses to user questions and requests. It uses a generative AI model to generate the responses the user expects.

[1535] 4. Emotion analysis:

[1536] The server uses an emotion engine to recognize the user's emotions from text and audio data. It analyzes the tone of voice and the content of speech to determine the user's emotional state. For example, if the user is excited, it will generate an energetic advertisement.

[1537] 5. Replies and ad generation:

[1538] Based on the results of sentiment analysis, ads and replies are generated that match the user's emotional state. For example, if the user is happy, ads with a bright and positive tone are generated.

[1539] 6. Convert replies to speech:

[1540] The server converts the generated text reply into audio data using speech synthesis technology (e.g., a TTS engine). The converted audio data is then played back to the user through the speaker.

[1541] 7. Presentation of visual information:

[1542] The device displays the generated advertisement content on its screen. Visual information such as product images, details, and prices are presented.

[1543] Specific example

[1544] For example, the user might say the following:

[1545] "Please tell me about the latest smartphones."

[1546] The system acquires the audio and generates advertisements through the following process.

[1547] Voice input acquisition

[1548] Text conversion using speech recognition

[1549] Text data analysis and response generation

[1550] Determining a user's happy emotions through emotion analysis.

[1551] Generate an energetic ad ("The latest smartphones have amazing camera features!")

[1552] Convert text replies to speech and output them through the speaker.

[1553] Display visual product information on the screen.

[1554] Example of a prompt

[1555] "The user said, 'Tell me about the latest smartphones.' The user seems very interested and is talking enthusiastically. What kind of ad would be suitable?"

[1556] This prompt allows the generative AI model to generate appropriate advertisements.

[1557] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1558] Step 1: Obtaining voice input

[1559] The user speaks into the device. The device acquires the user's voice input through the microphone. The input is the user's voice data, and the output is this voice data.

[1560] Specific operation: The microphone captures the user's voice and converts it into a digital format.

[1561] Step 2: Convert voice input to text

[1562] The device uses speech recognition software to convert acquired speech data into text data. The input is speech data, and the output is text data.

[1563] Specific operation: Speech recognition software (e.g., Google Speech Recognition API) analyzes the audio data and generates a corresponding string.

[1564] Step 3: Analyzing text data

[1565] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, it generates an appropriate response. The input is text data, and the output is the analysis results and the generated response.

[1566] Specific operation: A natural language processing model (e.g., a generative AI model) analyzes text data and generates a contextually appropriate response.

[1567] Step 4: Emotion Analysis

[1568] The server uses an emotion engine to recognize the user's emotions from text and audio data. The input is text and audio data, and the output is the emotion analysis result.

[1569] Specific operation: The emotion analysis model analyzes the text content and voice tone to determine the user's emotional state (e.g., joy, sadness, excitement, etc.).

[1570] Step 5: Replies and ad generation

[1571] The server considers the results of sentiment analysis to generate advertisements and replies appropriate to the user's emotional state. The input is the analysis results and sentiment analysis results, and the output is the generated replies and advertisements.

[1572] Specific operation: The generative AI model creates advertising content that reflects emotional information. For example, if the user is excited, it will generate an advertisement in an energetic tone.

[1573] Step 6: Convert your reply to speech

[1574] The server converts the generated text replies into speech data using speech synthesis technology. The input is text data, and the output is speech data.

[1575] Specific operation: A speech synthesis engine (e.g., a TTS engine) synthesizes text data and outputs it as natural-sounding speech.

[1576] Step 7: Audio Output

[1577] The device plays the converted audio data to the user through its speaker. The input is audio data, and the output is the audio that the user hears.

[1578] Specific operation: The speaker plays the audio data as a machine-generated voice.

[1579] Step 8: Presenting visual information

[1580] The device uses its display to visually show the generated advertisement content. The input is the advertisement content, and the output is the visual information displayed on the screen.

[1581] Specific operation: The display shows advertising images and detailed information, providing the user with visual feedback.

[1582] Step 9: Extract and save requests and preferences

[1583] The server extracts requests and preferences from the user's speech and stores them in a database. The input is text data, and the output is a user profile stored in the database.

[1584] Specific operation: A natural language processing model extracts important keywords and phrases from text and records them in a database.

[1585] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1586] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1587] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1588] [Fourth Embodiment]

[1589] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1590] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1591] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1592] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1593] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1594] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1595] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1596] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1597] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1598] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1599] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1600] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1601] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1602] This invention is a system for suggesting appropriate products and services through dialogue with users. This system is realized by combining technologies such as speech recognition, natural language processing, data analysis, and speech synthesis.

[1603] System-wide configuration

[1604] The system consists of the following main components:

[1605] 1. Means for obtaining user voice input

[1606] 2. Means of converting voice input to text

[1607] 3. Means for analyzing the converted text and generating an appropriate response.

[1608] 4. A means of converting replies into speech and outputting them.

[1609] 5. A means of extracting user requests and preferences and storing them in a database.

[1610] 6. Means of proposing products and services based on stored information

[1611] 7. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[1612] 8. Means of presenting visual information using a device equipped with a display.

[1613] System operation details

[1614] 1. Obtain user voice input.

[1615] The device acquires user voice input through its microphone. When the user speaks into the device, the microphone captures the audio data.

[1616] 2. Convert voice input to text

[1617] The acquired audio data is converted into text using speech recognition software. This ensures that what the user says is treated as text data.

[1618] 3. Analyze the converted text and generate a reply.

[1619] The server analyzes text data using natural language processing technology. Based on the analyzed data, an AI model generates appropriate responses. For example, in response to a question like "I want to travel," it generates a response suggesting a suitable travel destination.

[1620] 4. Convert the reply to speech and output it.

[1621] The generated text reply is converted into speech using speech synthesis technology. The converted speech data is then played back to the user through the device's speaker.

[1622] 5. Extract requests and preferences and save them to a database.

[1623] The server analyzes the user's speech to extract their requests and preferences. The extracted information is stored in the database as a user profile.

[1624] 6. Propose products and services

[1625] The server uses stored user profile information to run an algorithm that suggests appropriate products and services. The suggested products and services are presented to the user in text or voice.

[1626] 7. Proposals that take past history into consideration

[1627] The server provides more personalized suggestions based on the user's past conversation and transaction history. This makes it possible to offer suggestions that match the user's preferences.

[1628] 8. Presentation of visual information

[1629] Using devices equipped with displays, visual information is also provided to the user. This includes product images and detailed information.

[1630] Specific example

[1631] Specific examples of users:

[1632] The user speaks to a smart speaker with a display.

[1633] "I want to travel, could you recommend some places?"

[1634] Processing flow:

[1635] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[1636] 2. Speech-to-text conversion: Speech input is converted into text data.

[1637] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[1638] 4. Audio output of replies: The generated text reply is converted to audio and played through the speaker.

[1639] 5. Extraction and saving of preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[1640] 6. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[1641] 7. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[1642] 8. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[1643] The above describes a specific form for carrying out the invention. This system allows users to receive personalized product and service suggestions while engaging in voice-based dialogue.

[1644] The following describes the processing flow.

[1645] Program processing

[1646] Step 1: Obtain user voice input

[1647] Operation details:

[1648] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[1649] Step 2: Convert voice input to text

[1650] Operation details:

[1651] The device converts the acquired audio data into text data using speech recognition software. This ensures that what the user says is processed as text.

[1652] Step 3: Send the converted text to the server.

[1653] Operation details:

[1654] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[1655] Step 4: Analyze the text data and generate an appropriate response.

[1656] Operation details:

[1657] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[1658] Step 5: Send the reply from the server to the terminal as text data.

[1659] Operation details:

[1660] The server sends the generated text reply to the terminal. This text data is converted to speech and played back on the terminal.

[1661] Step 6: Convert the reply to speech and output it.

[1662] Operation details:

[1663] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[1664] Step 7: Extract user requests and preferences and save them to the database.

[1665] Operation details:

[1666] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[1667] Step 8: Propose products and services based on the saved information.

[1668] Operation details:

[1669] The server uses information stored in the database to suggest products and services that meet the user's needs. These suggestions are generated as a reply message and presented to the user.

[1670] Step 9: Consider past conversation and transaction history in relation to user data.

[1671] Operation details:

[1672] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[1673] Step 10: Provide visual information

[1674] Operation details:

[1675] The terminal uses a display-equipped device to show the user visual information. This includes images, details, and prices of the proposed products and services.

[1676] The above outlines the specific processing steps of the program in this system. This allows users to receive personalized product and service suggestions while engaging in voice-based conversations.

[1677] (Example 1)

[1678] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1679] Modern users handle vast amounts of information and demand personalized services and product recommendations, but traditional systems struggle to efficiently achieve this. In particular, few systems can start with voice input and then provide recommendations that take into account the user's past history and preferences. Therefore, there is a need to develop systems that improve user satisfaction.

[1680] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1681] In this invention, the server includes means for acquiring user voice input, means for converting the voice input into text, means for analyzing the converted text and generating an appropriate response using a generation AI model, means for converting the response into voice and outputting it, means for extracting the user's requests and preferences and storing them in a database, and means for suggesting products and services based on the stored information. This makes it possible to efficiently provide personalized suggestions that start with voice input and take into account the user's past history and preferences.

[1682] "Means for acquiring user voice input" refers to a function in which the terminal uses a microphone to capture the user's voice in real time and sends it to the server as voice data.

[1683] "Means of converting voice input to text" refers to a function in which a server uses speech recognition technology to convert acquired voice data into text data.

[1684] "Means for analyzing converted text and generating appropriate responses using a generative AI model" refers to a function in which the server uses natural language processing technology and a generative AI model to analyze text data and automatically generate appropriate responses to user questions and requests.

[1685] "Method for converting replies into audio and outputting them" refers to a function in which the server uses speech synthesis technology to convert text replies into audio data and plays it back to the user through the terminal's speaker.

[1686] "A means of extracting user requests and preferences and saving them to a database" refers to a function where the server analyzes the user's speech and saves information about their requests and preferences to a database.

[1687] "A means of suggesting products and services based on stored information" refers to a function in which the server uses an algorithm to suggest appropriate products and services based on user information stored in the database.

[1688] "A means of making suggestions that take into account past conversation and transaction history in relation to user data" refers to a function in which the server refers to the user's past conversation and transaction history and makes personalized suggestions based on that.

[1689] "Means of providing a user interface through a display device and presenting visual information" refers to the function of a terminal that uses its display to provide the user with visual information (for example, detailed information or images of products or services).

[1690] Modes for carrying out the invention

[1691] This invention is a system for proposing appropriate products and services through interaction with the user. This system is realized by combining the following hardware and software technologies.

[1692] System-wide configuration

[1693] The system consists of the following main components:

[1694] 1. Means for obtaining user voice input

[1695] 2. Means of converting voice input to text

[1696] 3. A means of analyzing the converted text and generating an appropriate response using a generative AI model.

[1697] 4. A means of converting replies into speech and outputting them.

[1698] 5. A means of extracting user requests and preferences and storing them in a database.

[1699] 6. Means of proposing products and services based on stored information

[1700] 7. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[1701] 8. Means of presenting visual information through a display device.

[1702] Acquiring voice input

[1703] The device uses a microphone to acquire user voice input. When the user speaks into the device, the microphone captures the voice data and sends it to the server as digital data. This data is processed using speech recognition technology.

[1704] Speech-to-text conversion

[1705] The server uses speech recognition software (e.g., a speech recognition API) to convert the acquired speech data into text. This allows the speech input to be treated as text data.

[1706] Text analysis and response generation

[1707] The converted text data is analyzed by the server using natural language processing techniques and generative AI models (e.g., large-scale language models). Based on the analyzed data, an appropriate response is generated. For example, if a user says "I want to travel," the generative AI model is given the prompt "The user says they want to travel. What travel destination would you suggest?" and the model generates a response such as "Hokkaido is a recommended travel destination."

[1708] Voice output of the reply

[1709] The generated text reply is converted into audio data using speech synthesis technology (e.g., a speech synthesis API). The converted audio data is played back to the user through the device's speaker. The user can then hear the response, such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[1710] Extraction of requests and preferences and storage in a database.

[1711] The server analyzes the user's speech to extract their requests and preferences. Customer insight tools are used for this purpose. The extracted information is stored in a database as a user profile. For example, if a user is identified as being interested in "travel," that information is saved.

[1712] Product and service proposals

[1713] The server uses stored user profile information to execute algorithms (e.g., collaborative filtering) to suggest appropriate products and services. This information is presented to the user in text or audio format.

[1714] Proposals that take past history into consideration

[1715] The server references the user's past conversation and transaction history to provide more personalized suggestions. For example, it can pique the user's interest by prioritizing travel destinations they have never visited before.

[1716] Presentation of visual information

[1717] If the device has a display, it will use the display to provide visual information about the suggested products and services. This includes detailed product information and photos. Users can visually confirm the information, leading to a deeper understanding.

[1718] Specific example

[1719] Specific examples of users:

[1720] "I want to travel, could you recommend some places?"

[1721] Processing flow:

[1722] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[1723] 2. Speech-to-text conversion: Speech input is converted into text data using a speech recognition API.

[1724] 3. Text analysis and reply generation: The server uses a generated AI model to generate the reply "Hokkaido is a recommended travel destination."

[1725] 4. Audio output of replies: The generated text reply is converted into audio data using a speech synthesis API and played back through the device's speaker.

[1726] 5. Extraction and saving of preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[1727] 6. Proposal: The server uses collaborative filtering to suggest services and products related to the travel destination.

[1728] 7. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[1729] 8. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the display device.

[1730] As described above, the system of the present invention allows users to receive personalized product and service suggestions through voice and visual information. This enables users to efficiently acquire information and make choices that meet their needs.

[1731] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1732] Step 1:

[1733] The device acquires the user's voice input. When the user speaks into the device, saying "I want to travel," the device's microphone captures the voice. The input is an analog audio signal, which the microphone sensor converts into digital data. Digital audio data is generated, and this becomes the initial output.

[1734] Step 2:

[1735] The server receives the digital audio data obtained in step 1 and uses speech recognition software (e.g., a speech recognition API) to convert the audio into text. Specifically, it converts the audio data "I want to go on a trip" into text data "Text:I want to go on a trip". This converted text data becomes the output.

[1736] Step 3:

[1737] The server, upon receiving the converted text data, analyzes it using natural language processing techniques. This analysis utilizes a generative AI model (e.g., a large-scale language model). The prompt "The user says they want to travel. What travel destination would you suggest?" is input to the generative AI model, and the model generates an appropriate response, "Hokkaido is a recommended travel destination." The analyzed data is then output.

[1738] Step 4:

[1739] The server converts the text reply generated in step 3 into audio data using speech synthesis technology (e.g., a speech synthesis API). Specifically, it converts "My recommended travel destination is Hokkaido" into audio data and sends that audio data to the terminal. This audio data becomes the output of step 4.

[1740] Step 5:

[1741] The device receives audio data sent from the server and plays the audio to the user through its speaker. The user hears the audio say, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food." Through this action, the user obtains the necessary information.

[1742] Step 6:

[1743] Based on what the user says, the server automatically analyzes the data and extracts the user's requests and preferences. For example, keywords such as "travel" and "nature" may be extracted. The extracted information is stored in the database as a user profile. This stored user profile data is then output.

[1744] Step 7:

[1745] The server executes data analysis algorithms (e.g., collaborative filtering) based on user profile information stored in the database, and suggests appropriate products and services. For example, it might suggest accommodations or tour packages related to a travel destination. The content of these suggestions becomes the output.

[1746] Step 8:

[1747] The server considers the user's past conversation and transaction history to provide more personalized product and service suggestions. It also uses past data to suggest travel destinations the user has never visited before. This personalized suggestion information is then output.

[1748] Step 9:

[1749] If the device has a display, the server also provides visual information. For example, it might display detailed information or photos of a suggested travel destination. The user can visually confirm the displayed information. This visual information becomes the final output.

[1750] Through the steps outlined above, this system comprehensively achieves everything from voice input, text analysis, appropriate response generation using a generative AI model, speech synthesis, extraction and storage of user preferences, suggestions based on data analysis, and even the provision of visual information.

[1751] (Application Example 1)

[1752] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1753] In modern autonomous vehicles, there is a need for systems that suggest appropriate destinations and sightseeing spots in real time to enhance the passenger experience. However, conventional in-car entertainment and navigation systems have the problem of not being able to make suggestions that meet the individual needs and preferences of users. Furthermore, there has been a challenge in providing more personalized suggestions that take past history into account.

[1754] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1755] In this invention, the server includes means for acquiring user voice input, means for converting voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, means for extracting user requests and preferences and storing them in a database, means for suggesting products and services based on the stored information, and means for suggesting appropriate drive destinations and sightseeing spots to passengers in the vehicle. This enables real-time appropriate suggestions that meet the requests and preferences of passengers.

[1756] "Acquiring user voice input" means that microphones installed inside the vehicle capture the voices of passengers.

[1757] "Converting voice input to text" means converting acquired voice data into text data using speech recognition technology.

[1758] "Analyzing the converted text and generating appropriate responses" means using natural language processing technology to analyze text data and generate appropriate information and suggestions for passengers.

[1759] "Converting replies to audio and outputting them" means converting the generated replies into an audio format using speech synthesis technology and playing them through the car's speakers.

[1760] "Extracting user requests and preferences and storing them in a database" means extracting requests and preferences from passengers' statements and storing them in a database.

[1761] "Proposing products and services based on stored information" means selecting and proposing appropriate products and services based on customer needs and preferences stored in a database.

[1762] "Suggesting appropriate driving destinations and sightseeing spots to passengers inside the vehicle" means suggesting appropriate driving destinations and sightseeing spots in real time through dialogue with passengers inside the autonomous vehicle.

[1763] This invention relates to a system for suggesting appropriate driving destinations and sightseeing spots to passengers in an autonomous vehicle. This system is implemented using the following key hardware and software components, combining speech recognition, natural language processing, speech synthesis, and database technologies.

[1764] System Configuration

[1765] 1. Hardware Configuration

[1766] Microphone: A high-sensitivity microphone installed inside the vehicle is used. This microphone is used to capture passenger voice input. For example, the SureBity USB microphone is suitable for this purpose.

[1767] Speakers: Audio is output using the in-vehicle speakers. Voice replies to passengers are also delivered through these speakers.

[1768] Display-equipped devices: Display devices installed inside the vehicle are used to provide visual information.

[1769] 2. Software Configuration

[1770] Speech recognition software: Using the speech_recognition library, audio data acquired from the microphone is converted into text data.

[1771] Natural language processing software: Using the pipeline function of the transformers library, the transformed text is analyzed to generate appropriate responses. In particular, "gpt-3" is used as the generative AI model.

[1772] Text-to-speech software: Using the gTTS (Google Text-to-Speech) library, the generated text replies are converted into audio data and played back through the car's speakers.

[1773] Database: A database system used to store user requests and preferences. It stores past conversation and transaction history and is used to provide more personalized suggestions.

[1774] System operation

[1775] The operation of this system is described as follows:

[1776] 1. Speech acquisition and recognition

[1777] The terminal acquires the passenger's voice input through the microphone. Voice recognition software converts this into text data and sends it to the server.

[1778] 2. Natural Language Processing and Proposal Generation

[1779] The server analyzes text data using natural language processing techniques and generates appropriate suggestions. These suggestions are then converted into speech using speech synthesis software.

[1780] 3. Audio output and data storage

[1781] The converted audio data is played back to passengers through the in-car speakers. At the same time, information about the suggestions, as well as passenger requests and preferences, is stored in a database and used to provide more personalized suggestions in the future.

[1782] 4. Visual information provision

[1783] Information on suggested driving destinations and tourist spots is presented visually through the in-car display system. This allows passengers to understand the suggestions more concretely.

[1784] Examples and prompts for generative AI models

[1785] Specific example:

[1786] When a user says, "I'm looking for a good restaurant," the speech recognition system converts this into text and generates a prompt message like the following, which is then input into the AI ​​model.

[1787] Prompts for the generative AI model:

[1788] To make the most of our drive, could you recommend some good restaurants?

[1789] The implementation of this system will provide a comfortable and personalized experience within autonomous vehicles. In the concrete implementation of the invention, the aforementioned hardware and software will be integrated to enable real-time and highly accurate suggestions.

[1790] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1791] Step 1:

[1792] The user speaks into a microphone installed inside the vehicle. The terminal acquires the user's voice input. The input is the passenger's voice data, and the output is a captured raw audio file. This audio data is then passed on to subsequent speech recognition processing.

[1793] Step 2:

[1794] The device converts the acquired audio data into text data using speech recognition software (speech_recognition library). The input is the audio file acquired in step 1, and the output is text data. Specifically, the speech recognition software analyzes the audio waveform and converts it into a string based on a language model.

[1795] Step 3:

[1796] The server analyzes the converted text using natural language processing software (the pipeline function of the transformers library) and generates an appropriate response. The input is the text data obtained in step 2, and the output is a text response containing the suggested content. A generative AI model (e.g., gpt-3) is used to analyze the text data and generate prompt sentences that respond to the user's questions and requests.

[1797] Step 4:

[1798] The server converts the generated text reply into audio data using speech synthesis technology (gTTS library). The input is the text reply generated in step 3, and the output is an audio file. Specifically, the speech synthesis technology converts the text into an audio signal and saves it as an MP3 file.

[1799] Step 5:

[1800] The terminal plays the converted audio data to the user through the in-car speakers. The input is the audio file generated in step 4, and the output is the presentation to the user as audio. The audio is played using software that plays MP3 files (e.g., mpg321).

[1801] Step 6:

[1802] The server extracts user requests and preferences and stores them in a database. The input is the text data obtained in step 3, and the output is a new record in the user profile database. Natural language processing techniques are used to extract information about requests and preferences from the text, organize it, and add it to the database.

[1803] Step 7:

[1804] The server executes an algorithm that suggests products and services based on stored information. The input is user profile information from the database, and the output is text data containing new suggestions. These suggestions are then provided to the user in subsequent speech recognition and natural language processing steps.

[1805] Step 8:

[1806] The server or terminal visually presents information about suggested drive destinations and tourist spots via an in-car display device. The input is the suggested content obtained in step 7, and the output is the visual information displayed on the display. Image data and map data are used to visually represent the suggested content.

[1807] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1808] This invention combines a system that proposes appropriate products and services through dialogue with the user with an emotion engine that recognizes the user's emotions. This system is realized by combining technologies such as speech recognition, natural language processing, emotion analysis, data analysis, and speech synthesis.

[1809] System-wide configuration

[1810] The system consists of the following main components:

[1811] 1. Means for obtaining user voice input

[1812] 2. Means of converting voice input to text

[1813] 3. Means for analyzing the converted text and generating an appropriate response.

[1814] 4. A means of converting replies into speech and outputting them.

[1815] 5. An emotion engine that recognizes user emotions and reflects that emotional information in the analysis results.

[1816] 6. A means of extracting user requests and preferences and storing them in a database.

[1817] 7. Means of proposing products and services based on stored information

[1818] 8. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[1819] 9. Means of presenting visual information using a device equipped with a display.

[1820] System operation details

[1821] 1. Obtain user voice input.

[1822] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[1823] 2. Convert voice input to text

[1824] The acquired audio data is converted into text data using speech recognition software. This ensures that the user's spoken content is treated as text data.

[1825] 3. Send the converted text to the server.

[1826] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[1827] 4. Analyze text data and generate a reply.

[1828] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[1829] 5. Emotional analysis using an emotion engine

[1830] The server uses an emotion engine to recognize the user's emotions from voice and text data. The emotion engine analyzes the user's voice tone and text content to determine the user's emotional state. This emotional information is then reflected in the replies and suggestions.

[1831] 6. Send the reply as text data from the server to the terminal.

[1832] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[1833] 7. Convert the reply to speech and output it.

[1834] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[1835] 8. Extract user requests and preferences and save them to a database.

[1836] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[1837] 9. Propose products and services.

[1838] The server uses stored user profile information to run an algorithm that suggests products and services tailored to the user's needs. The suggestions are generated as a reply message and presented to the user.

[1839] 10. Proposals that take past history into consideration

[1840] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[1841] 11. Presentation of visual information

[1842] The terminal uses a display-equipped device to show the user visual information. This includes product images, detailed information, and prices.

[1843] Specific example

[1844] Specific examples of users:

[1845] The user speaks to a smart speaker with a display.

[1846] "I want to travel, could you recommend some places?"

[1847] Processing flow:

[1848] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[1849] 2. Speech-to-text conversion: Speech input is converted into text data.

[1850] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[1851] 4. Sentiment Analysis: The server's emotion engine analyzes the user's voice and text to determine if the user is excited.

[1852] 5. Reflecting Emotional Information: Reflect emotional information in the generated replies and adjust them to have an energetic tone.

[1853] 6. Voice output of replies: The generated text reply is converted to voice and played back from the speaker in an energetic tone.

[1854] 7. Extracting and saving preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[1855] 8. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[1856] 9. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[1857] 10. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[1858] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[1859] The following describes the processing flow.

[1860] Program processing

[1861] Step 1: Obtain user voice input

[1862] Operation details:

[1863] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[1864] Step 2: Convert voice input to text

[1865] Operation details:

[1866] The device converts the acquired audio data into text data using speech recognition software. This ensures that what the user says is processed as text.

[1867] Step 3: Send the converted text to the server.

[1868] Operation details:

[1869] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[1870] Step 4: Analyze the text data and generate a reply.

[1871] Operation details:

[1872] The server analyzes the received text data using natural language processing technology. Based on the analysis results, a generative AI model generates appropriate responses to the user's questions and requests.

[1873] Step 5: Perform emotion analysis using the emotion engine.

[1874] Operation details:

[1875] The server uses an emotion engine to recognize the user's emotions from voice and text data. The emotion engine analyzes the user's voice tone and text content to determine the user's emotional state.

[1876] Step 6: Reflect emotional information in your replies.

[1877] Operation details:

[1878] The server adjusts the content and tone of the generated response based on emotional information obtained from the emotion engine. For example, if the user is excited, the response will be in an energetic tone.

[1879] Step 7: Send the reply from the server to the terminal as text data.

[1880] Operation details:

[1881] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[1882] Step 8: Convert the reply to speech and output it.

[1883] Operation details:

[1884] The device converts received text replies into audio data using speech synthesis technology. The converted audio data is then played back to the user through the speaker.

[1885] Step 9: Extract user requests and preferences and save them to the database.

[1886] Operation details:

[1887] The server analyzes the user's speech to extract requests and preferences. The extracted information is stored in a database as a user profile.

[1888] Step 10: Propose products and services based on the saved information.

[1889] Operation details:

[1890] The server uses information stored in the database to suggest products and services that meet the user's needs. These suggestions are generated as a reply message and presented to the user.

[1891] Step 11: Make a proposal that takes past history into consideration.

[1892] Operation details:

[1893] The server provides more personalized suggestions based on the user's past conversation and transaction history. This allows for suggestions tailored to the user's preferences and characteristics.

[1894] Step 12: Provide visual information

[1895] Operation details:

[1896] The terminal uses a display-equipped device to show the user visual information. This includes images, details, and prices of the proposed products and services.

[1897] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[1898] (Example 2)

[1899] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1900] Modern users demand fast and personalized product and service recommendations, requiring more accurate suggestions that take their emotional state into account. However, existing systems struggle to recognize user emotions and adjust their recommendations accordingly. Furthermore, they often fail to fully utilize user conversation and transaction history, resulting in a lack of information that users truly need.

[1901] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1902] In this invention, the server includes means for using an emotion engine to recognize the user's emotions, means for sending the generated response from the server to the terminal as text data, and means for making suggestions that take into account past conversation and transaction history in association with user data. This makes it possible to make more personalized product and service suggestions based on the user's emotions and past history.

[1903] "Means for acquiring user voice input" refers to a function that captures voice data spoken by the user using the device's microphone.

[1904] "Methods for converting voice input to text" refers to technologies that use speech recognition software to convert captured voice data into text format.

[1905] "Means for sending converted text to a server" refers to the function by which the terminal sends the text data generated by speech recognition to a server via the network.

[1906] "Means for analyzing text data and generating appropriate replies" refers to a function in which the server analyzes the text data it receives using natural language processing technology and generates an appropriate reply based on its content.

[1907] "Methods of using an emotion engine that recognizes user emotions from voice and text data" refers to a technology in which a server analyzes voice and text data, recognizes the user's emotional state, and incorporates that information into processing.

[1908] "Means for sending generated replies as text data from the server to the terminal" refers to a function for sending replies generated by the server in text format to the terminal.

[1909] "A means of converting replies into speech and outputting them" refers to a function that converts text data received by the device into speech format using speech synthesis technology and plays it back to the user through the speaker.

[1910] "A means of extracting user requests and preferences and storing them in a database" refers to a technology in which a server analyzes the user's speech to extract requests and preferences and stores that information in a database.

[1911] "A means of suggesting products and services based on stored information" refers to a technology in which a server refers to stored user profile information and automatically suggests products and services that are suitable for the user.

[1912] "A means of making suggestions that take into account past conversation and transaction history in relation to user data" refers to a technology in which a server operates an algorithm based on the user's past conversation and transaction history to generate personalized suggestions.

[1913] "A means of providing a user interface through a device with a display and presenting visual information" refers to the function of a terminal that uses its display to present visual information (images, detailed information, price, etc.) to the user.

[1914] This invention combines a system that proposes appropriate products and services through dialogue with the user with an emotion engine that recognizes the user's emotions. This system is realized by combining technologies such as speech recognition, natural language processing, emotion analysis, data analysis, and speech synthesis.

[1915] System-wide configuration

[1916] The system consists of the following main components:

[1917] 1. Means for obtaining user voice input

[1918] 2. Means of converting voice input to text

[1919] 3. Means for sending the converted text to the server

[1920] 4. Means for analyzing text data and generating replies

[1921] 5. An emotion engine that recognizes user emotions and reflects that emotional information in the analysis results.

[1922] 6. Means for sending the generated reply from the server to the terminal.

[1923] 7. A means of converting replies into speech and outputting them.

[1924] 8. A means of extracting user requests and preferences and storing them in a database.

[1925] 9. Means of proposing products and services based on stored information

[1926] 10. A means of making suggestions that take into account past conversation and transaction history in relation to user data.

[1927] 11. Means for presenting visual information using a device equipped with a display.

[1928] System operation details

[1929] Get user voice input

[1930] The device acquires user voice input through its microphone. When the user speaks into the device, the voice data is captured by the microphone.

[1931] Convert voice input to text

[1932] The acquired audio data is converted into text data using speech recognition software. For example, the Google Cloud Speech-to-Text API can be used to treat what the user says as text data.

[1933] Send the converted text to the server.

[1934] The device sends the text data converted by speech recognition to the server. This data is used on the server side for further analysis.

[1935] Analyze text data and generate a reply.

[1936] The server analyzes the received text data using natural language processing techniques. For example, it uses a generative AI model such as OpenAI GPT-4 to generate appropriate responses to user questions and requests.

[1937] Emotional analysis using an emotion engine

[1938] The server uses an emotion engine to recognize user emotions from voice and text data. For example, it uses "IBM Watson Tone Analyzer" to analyze the user's voice tone and text content to determine their emotional state. This emotional information is then reflected in the replies and suggestions.

[1939] Send the generated reply from the server to the terminal.

[1940] The server sends the generated reply as text data to the terminal. This text data is then converted to speech and played back on the terminal.

[1941] Convert the reply into speech and output it.

[1942] The device converts received text replies into audio data using speech synthesis technology. For example, the converted audio data, using "Amazon Polly," is played back to the user through the speaker.

[1943] Extract user requests and preferences and save them to a database.

[1944] The server performs additional analysis to extract requests and preferences from the user's speech. The extracted information is stored in the database as a user profile.

[1945] Propose products and services

[1946] The server uses stored user profile information to run an algorithm that suggests products and services tailored to the user's needs. The suggestions are generated as a reply message and presented to the user.

[1947] Proposals that take past history into consideration

[1948] The server provides more personalized suggestions based on the user's past conversation and transaction history. This historical information is retrieved from a database and incorporated into the suggestion algorithm.

[1949] Presentation of visual information

[1950] The terminal uses a display-equipped device to show the user visual information. This includes product images, detailed information, and prices.

[1951] Specific example

[1952] Specific examples of users:

[1953] The user speaks to a smart speaker with a display: "I want to travel, could you recommend some places?"

[1954] Processing flow:

[1955] 1. Voice input acquisition: The user's voice is acquired by the device's microphone.

[1956] 2. Speech-to-text conversion: Speech is converted to text using speech recognition software (e.g., Google Cloud Speech-to-Text).

[1957] 3. Text analysis and reply generation: The server analyzes the text data and generates a reply such as, "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food."

[1958] 4. Sentiment Analysis: An emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's voice and text to determine if the user is excited.

[1959] 5. Reflecting Emotional Information: Reflect emotional information in the generated replies and adjust them to have an energetic tone.

[1960] 6. Voice output of replies: The generated text replies are converted into speech using speech synthesis technology (e.g., Amazon Polly) and played back through the speaker in an energetic tone.

[1961] 7. Extracting and saving preferences: User preferences are extracted using "travel" as a keyword and saved in the database.

[1962] 8. Suggestion: Based on the stored information, the server will suggest services and products related to the travel destination.

[1963] 9. History-Based Suggestions: Based on past history, prioritize suggesting locations the user has not visited before.

[1964] 10. Presentation of visual information: Photos and information about the suggested travel destination are displayed on the screen.

[1965] Example of a prompt:

[1966] I want to go on a trip, could you recommend some places?

[1967] The above describes the specific form of the invention. This system allows users to receive personalized product and service suggestions while interacting with the system via voice. Furthermore, by recognizing the user's emotions and reflecting them in the suggestions, a more personalized experience can be provided.

[1968] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1969] Step 1:

[1970] The device acquires the user's voice input.

[1971] Input: The microphone captures the user's speech.

[1972] Operation: When the user says, "I want to travel, can you recommend some places?", the device's microphone captures the audio data.

[1973] Output: Captured audio data.

[1974] Step 2:

[1975] The device converts voice input into text.

[1976] Input: Acquired audio data.

[1977] Operation: The device uses the Google Cloud Speech-to-Text API to convert speech data into text data.

[1978] Output: Converted text data (e.g., "I want to travel, could you recommend some places?").

[1979] Step 3:

[1980] The terminal sends the converted text to the server.

[1981] Input: Text data.

[1982] Operation: The terminal sends text data to the server over the network.

[1983] Output: Text data sent to the server.

[1984] Step 4:

[1985] The server analyzes the text data and generates a reply.

[1986] Input: Text data received by the server.

[1987] Operation: The server uses natural language processing techniques (e.g., OpenAI GPT-4) to analyze text data and generate appropriate responses.

[1988] Output: Generated reply text (Example: "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food.").

[1989] Step 5:

[1990] The server uses an emotion engine to perform emotion analysis.

[1991] Input: Audio data or text data.

[1992] Operation: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. For example, it can determine whether the user is excited or not.

[1993] Output: Sentiment analysis results (e.g., user is agitated).

[1994] Step 6:

[1995] The server sends the generated reply as text data to the terminal.

[1996] Input: Generated reply text and sentiment analysis results.

[1997] Operation: The server sends a reply text to the terminal and adjusts it to reflect sentiment information.

[1998] Output: Reply text sent to the terminal.

[1999] Step 7:

[2000] The device converts the reply text into speech and outputs it.

[2001] Input: Reply text data.

[2002] Operation: The device uses speech synthesis technology such as Amazon Polly to convert text into audio data, which is then played through the speaker.

[2003] Output: Audio data that the user hears (e.g., "My recommended travel destination is Hokkaido. You can enjoy beautiful nature and delicious food.").

[2004] Step 8:

[2005] The server extracts user requests and preferences and stores them in a database.

[2006] Input: Text data of the user's spoken content.

[2007] Operation: The server analyzes the spoken content, extracts requests and preferences, and stores them in a database. For example, it might extract keywords related to "travel."

[2008] Output: Extracted user profile information.

[2009] Step 9:

[2010] The server suggests products and services based on the information it stores.

[2011] Input: User profile information.

[2012] Operation: The server executes an algorithm that suggests products and services that meet the user's needs.

[2013] Output: A reply message regarding the proposed product or service (e.g., "We have a promotional discount for travel. Are you interested?").

[2014] Step 10:

[2015] The server makes suggestions that take past history into consideration.

[2016] Input: User's past conversation and transaction history.

[2017] Operation: The server generates personalized suggestions based on past conversation and transaction history. For example, it might suggest avoiding places you've visited before.

[2018] Output: Personalized suggestions.

[2019] Step 11:

[2020] The device displays visual information.

[2021] Input: Data about the proposed product or service.

[2022] Operation: The device uses its display to present visual information to the user. This includes displaying product images, detailed information, and prices.

[2023] Output: Visual information displayed on the screen (e.g., photos and prices of suggested travel destinations).

[2024] (Application Example 2)

[2025] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2026] Traditional advertising systems have a problem in that they present uniform ads without considering user emotions, thus failing to maximize advertising effectiveness. Furthermore, they lack sufficient personalized suggestions based on user requests and preferences, and there is a need to improve the user experience. The challenge is to improve advertising effectiveness while increasing user satisfaction by recognizing emotions and presenting timely and appropriate ads.

[2027] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for acquiring the user's voice input, means for converting the voice input into text, means for analyzing the converted text and generating an appropriate response, means for converting the response into voice and outputting it, an emotion engine means for recognizing the user's emotions and reflecting that emotion information in the response and suggestions, means for extracting the user's requests and preferences and storing them in a database, means for suggesting products and services based on the stored information, and a user interface means for presenting the suggestions in audio and visual ways. This makes it possible to provide personalized advertising suggestions that take the user's emotions into consideration.

[2028] "Acquiring voice input" is the process of capturing the voice spoken by the user through the device's microphone.

[2029] "Converting to text" is the process of converting acquired voice input into text data using a speech recognition algorithm.

[2030] "Text analysis" is the process of analyzing converted text data using natural language processing technology to understand and interpret its content.

[2031] "Generating appropriate responses" is the process of creating the best possible response to a user's question or request based on the results of text analysis.

[2032] "Converting replies to speech" is the process of converting generated text responses into speech data using speech synthesis technology.

[2033] "Audio output" is the process of making audio data audible to the user through an audio device such as a speaker.

[2034] An "emotion engine" is an algorithm or component that analyzes a user's voice or text to determine the user's emotional state (joy, sadness, anxiety, etc.).

[2035] "Extracting requests and preferences" is the process of identifying what an individual wants and likes from their spoken content and recording it in a database.

[2036] "Product and service proposals" refer to the process of selecting the most suitable products and services for a given user based on their saved user data, and then creating a proposal.

[2037] A "user interface" is an interface for a user to interact with a system, and includes means of presenting audio and visual information.

[2038] Overall System Overview

[2039] The system implementing this invention provides personalized advertisements to users via terminals such as smartphones and smart glasses. It acquires the user's voice input and performs a series of processes to present appropriate responses and advertisements that reflect emotions. The system components include voice input means, voice recognition means, text analysis means, emotion analysis means, database storage means, product / service suggestion means, and visual and audio output means. Its specific operation is described below.

[2040] Hardware and software used

[2041] The entire system will be implemented using the following hardware and software:

[2042] A device such as a smartphone or smart glasses (including microphone, speaker, and display)

[2043] Microphone: A device for acquiring voice input.

[2044] Speaker: A device for outputting sound.

[2045] Display: A device for displaying visual information.

[2046] Speech recognition software: Software used to convert acquired speech into text (e.g., Google Speech Recognition API)

[2047] Sentiment analysis model: A model for determining a user's emotions (for example, a sentiment analysis model based on BERT).

[2048] Natural Language Processing Models: Models that analyze user text and generate optimal replies or advertisements (e.g., generative AI models).

[2049] Text-to-speech conversion software (TTS engine)

[2050] Database: Storage for saving user preferences and requests (e.g., SQL database)

[2051] Detailed Operation Description

[2052] 1. Acquisition of voice input:

[2053] The device acquires user voice input through the microphone. When the user speaks, the voice data is acquired and the process proceeds to the next stage.

[2054] 2. Convert voice input to text:

[2055] The acquired audio data is converted into text data using speech recognition software (e.g., Google Speech Recognition API). This ensures that what the user says is treated as text data.

[2056] 3. Analysis of text data:

[2057] The server analyzes text data using natural language processing techniques and generates appropriate responses to user questions and requests. It uses a generative AI model to generate the responses the user expects.

[2058] 4. Emotion analysis:

[2059] The server uses an emotion engine to recognize the user's emotions from text and audio data. It analyzes the tone of voice and the content of speech to determine the user's emotional state. For example, if the user is excited, it will generate an energetic advertisement.

[2060] 5. Replies and ad generation:

[2061] Based on the results of sentiment analysis, ads and replies are generated that match the user's emotional state. For example, if the user is happy, ads with a bright and positive tone are generated.

[2062] 6. Convert replies to speech:

[2063] The server converts the generated text reply into audio data using speech synthesis technology (e.g., a TTS engine). The converted audio data is then played back to the user through the speaker.

[2064] 7. Presentation of visual information:

[2065] The device displays the generated advertisement content on its screen. Visual information such as product images, details, and prices are presented.

[2066] Specific example

[2067] For example, the user might say the following:

[2068] "Please tell me about the latest smartphones."

[2069] The system acquires the audio and generates advertisements through the following process.

[2070] Voice input acquisition

[2071] Text conversion using speech recognition

[2072] Text data analysis and response generation

[2073] Determining a user's happy emotions through emotion analysis.

[2074] Generate an energetic ad ("The latest smartphones have amazing camera features!")

[2075] Convert text replies to speech and output them through the speaker.

[2076] Display visual product information on the screen.

[2077] Example of a prompt

[2078] "The user said, 'Tell me about the latest smartphones.' The user seems very interested and is talking enthusiastically. What kind of ad would be suitable?"

[2079] This prompt allows the generative AI model to generate appropriate advertisements.

[2080] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[2081] Step 1: Obtaining voice input

[2082] The user speaks into the device. The device acquires the user's voice input through the microphone. The input is the user's voice data, and the output is this voice data.

[2083] Specific operation: The microphone captures the user's voice and converts it into a digital format.

[2084] Step 2: Convert voice input to text

[2085] The device uses speech recognition software to convert acquired speech data into text data. The input is speech data, and the output is text data.

[2086] Specific operation: Speech recognition software (e.g., Google Speech Recognition API) analyzes the audio data and generates a corresponding string.

[2087] Step 3: Analyzing text data

[2088] The server analyzes the received text data using natural language processing techniques. Based on the analysis results, it generates an appropriate response. The input is text data, and the output is the analysis results and the generated response.

[2089] Specific operation: A natural language processing model (e.g., a generative AI model) analyzes text data and generates a contextually appropriate response.

[2090] Step 4: Emotion Analysis

[2091] The server uses an emotion engine to recognize the user's emotions from text and audio data. The input is text and audio data, and the output is the emotion analysis result.

[2092] Specific operation: The emotion analysis model analyzes the text content and voice tone to determine the user's emotional state (e.g., joy, sadness, excitement, etc.).

[2093] Step 5: Replies and ad generation

[2094] The server considers the results of sentiment analysis to generate advertisements and replies appropriate to the user's emotional state. The input is the analysis results and sentiment analysis results, and the output is the generated replies and advertisements.

[2095] Specific operation: The generative AI model creates advertising content that reflects emotional information. For example, if the user is excited, it will generate an advertisement in an energetic tone.

[2096] Step 6: Convert your reply to speech

[2097] The server converts the generated text replies into speech data using speech synthesis technology. The input is text data, and the output is speech data.

[2098] Specific operation: A speech synthesis engine (e.g., a TTS engine) synthesizes text data and outputs it as natural-sounding speech.

[2099] Step 7: Audio Output

[2100] The device plays the converted audio data to the user through its speaker. The input is audio data, and the output is the audio that the user hears.

[2101] Specific operation: The speaker plays the audio data as a machine-generated voice.

[2102] Step 8: Presenting visual information

[2103] The device uses its display to visually show the generated advertisement content. The input is the advertisement content, and the output is the visual information displayed on the screen.

[2104] Specific operation: The display shows advertising images and detailed information, providing the user with visual feedback.

[2105] Step 9: Extract and save requests and preferences

[2106] The server extracts requests and preferences from the user's speech and stores them in a database. The input is text data, and the output is a user profile stored in the database.

[2107] Specific operation: A natural language processing model extracts important keywords and phrases from text and records them in a database.

[2108] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[2109] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2110] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[2111] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2112] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated ...

Claims

1. A means of obtaining user voice input, A means of converting voice input to text, A means of analyzing the converted text and generating an appropriate reply, A method for converting replies into speech and outputting them, A means of extracting user requests and preferences and storing them in a database, A means of proposing products and services based on stored information, A system that includes this.

2. The system according to claim 1, including means for making suggestions that take into account past conversation history and transaction history in relation to user data.

3. The system according to claim 1, further comprising means for providing a user interface through a device with a display and for presenting visual information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A