System

The system addresses the lack of personalization in radio by using AI to generate and deliver personalized audio content based on user data, enhancing user engagement and advertising effectiveness.

JP2026014924APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116398
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Traditional radio programs struggle to provide personalized information, leading to decreased user satisfaction, while visually demanding platforms like YouTube and TikTok result in users skipping advertisements, highlighting a need for personalized and engaging audio content.

Method used

A system that collects user profile information and usage history, analyzes interests, generates personalized conversations and advertisements using AI models, and delivers them as audio files, reducing visual strain and enhancing user engagement.

Benefits of technology

Provides personalized radio experiences, improving user satisfaction and advertising effectiveness by delivering content tailored to individual interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014924000001_ABST
    Figure 2026014924000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A means for inputting profile information of a user, a means for collecting and storing the input profile information, a means for recording use history data of the user, a means for periodically transmitting the recorded use history data to a server, a means for analyzing an interest and concern of the user based on the transmitted use history data and profile information, a means for selecting a topic based on the analysis result and generating conversation content, and a means for receiving a message from the user, this system is provided with a means for generating related information and an answer after analyzing contents, a means for selecting a related advertisement based on the interest / concern of a user, and for generating conversation and the advertisement as a voice file, a means for distributing the generated voice file to a user terminal, and a means for allowing the user terminal to reproduce the voice file.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In today's society, where a large amount of visual information is provided, users tend to fast-forward or skip advertisements and content. This problem is particularly pronounced on platforms like YouTube and TikTok. On the other hand, radio, which is less visually demanding, is often listened to naturally, and radio listenership has increased, especially since the COVID-19 pandemic. However, traditional radio programs have struggled to provide personalized information, making it difficult to increase user satisfaction. The present invention aims to realize a radio system that provides personalized conversations relevant to each individual user and maintains their interest. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system that includes the following means. First, a means is provided for a user to install an application and input profile information when the application is first launched. Next, a means is provided for collecting and saving the input profile information and recording the user's usage history data on a daily basis. The system also includes a means for periodically transmitting the collected data to a server, which then analyzes the user's interests and concerns. The system also includes a means for selecting topics based on the analysis results and generating conversation content using a natural language generation model. The system also provides a means for receiving letters from users, analyzing their content, and generating related information and answers. The system also includes a means for selecting relevant advertisements based on the analysis results and generating and distributing these advertisements together with the conversation as an audio file. Finally, the system provides a means for the user's device to play the distributed audio file and for the user to operate the audio file during playback, thereby reducing the visual burden on the user and providing personalized information provision.

[0006] "User" refers to an individual who uses the application to input information and listen to radio programs.

[0007] "Profile information" refers to personal data such as a user's age, gender, occupation, hobbies, etc.

[0008] "Usage history data" refers to information such as playback time, number of skips, and selected topics when a user uses an application.

[0009] "Server" refers to a central processing unit that receives, analyzes, and stores user profile information and usage history data.

[0010] "Analysis" refers to the process of identifying user interests based on collected profile information and historical usage data.

[0011] A "topic" refers to a topic selected based on a user's interests and concerns.

[0012] A "natural language generation model" refers to an algorithm or model for generating sentences in natural language based on given input data.

[0013] "Message" refers to a message or question that a user sends through the application.

[0014] "Advertisement" refers to a commercial message that is selected based on user interests and presented within a radio program.

[0015] "Audio File" means a digital file that encodes the Generated Speech and Advertisements as audio.

[0016] "Playback means" refers to a function that plays back the audio file received by the user terminal for the user.

[0017] "Terminal" refers to the electronic device on which the user operates the application, communicates with the server, and plays the audio files. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] This invention is an AI radio system that develops conversations based on the user's interests and concerns. The processing of the system's program is explained in natural language with specific examples.

[0040] User Data Collection and Initial Settings

[0041] User

[0042] First, the user installs the application and enters profile information when the application is first launched.

[0043] Terminal

[0044] The entered profile information is transmitted from the user terminal to the server.

[0045] server

[0046] The server stores the received profile information in a database for later analysis.

[0047] Daily data collection

[0048] Terminal

[0049] Record usage history data (playback time, number of skips, selected topics, etc.) recorded when a user uses the application.

[0050] server

[0051] The server receives usage history data periodically sent from the device and stores it in a database. The server then analyzes this data using AI algorithms to identify the user's interests.

[0052] Generate personalized conversations

[0053] server

[0054] The server then selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in technology, the server might select a topic such as "latest smartwatch features." Based on the selected topic, the server generates the conversation using a natural language generation model. At this time, an appropriate speech synthesis model is also applied to add emotion and tone.

[0055] Processing letters and providing information

[0056] User

[0057] Users can submit questions or comments using the application's letter feature.

[0058] Terminal

[0059] A letter is sent from the user terminal to the server.

[0060] server

[0061] The server analyzes the received letter and generates related information and answers based on its content. For example, if a user asks for a "recommended fitness app," the server generates an answer such as "The recommended fitness app is XX." Related advertisements are also selected.

[0062] Delivering personalized ads

[0063] server

[0064] The server selects relevant advertisements based on the user's profile information and current conversation content, and encodes the generated conversation and advertisements as audio files.

[0065] Examples:

[0066] For technology-loving users, the server will select advertisements for the latest smartwatches and seamlessly integrate them into the conversation.

[0067] Content Delivery and Playback

[0068] server

[0069] The generated audio file is delivered to the user terminal.

[0070] Terminal

[0071] The user terminal plays the audio file, allowing the user to listen to it, and allows the user to control the playback (stop, skip, play, etc.).

[0072] User

[0073] Users listen to the AI ​​radio and provide playback control and feedback as needed.

[0074] Explanation based on concrete examples

[0075] For example, suppose a user sends a letter asking, "What fitness app do you recommend?" In this case, the server analyzes the letter, replies, "The fitness app I recommend is ____," and generates a related advertisement (e.g., a sale on a fitness app). The generated audio file is then delivered to the user's device, and when the user listens to it, the answer to the question and the advertisement information are simultaneously provided.

[0076] The system enables a personalized radio experience based on the user's interests, provides useful information while reducing visual strain, and improves advertising effectiveness through personalized advertising.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[0080] Step 2:

[0081] The device sends the entered profile information to the server.

[0082] Step 3:

[0083] The server stores the received profile information in a database.

[0084] Step 4:

[0085] The device records data about the user's use of the application (playback time, number of skips, selected topics, etc.).

[0086] Step 5:

[0087] The terminal periodically transmits the recorded usage history data to the server.

[0088] Step 6:

[0089] The server receives the usage history data and stores it in a database.

[0090] Step 7:

[0091] The server uses AI algorithms to analyze users' interests based on their profile information and usage history data.

[0092] Step 8:

[0093] The server selects topics appropriate for the user based on the analysis results.

[0094] Step 9:

[0095] The server uses a natural language generation model to generate conversational content based on the selected topic.

[0096] Step 10:

[0097] The server applies a speech synthesis model to add emotion and tone to the generated dialogue.

[0098] Step 11:

[0099] A user submits a question or comment using the in-app letter feature.

[0100] Step 12:

[0101] The terminal sends the letter from the user to the server.

[0102] Step 13:

[0103] The server receives the letter and analyzes its contents.

[0104] Step 14:

[0105] The server generates relevant information and answers based on the analysis results.

[0106] Step 15:

[0107] The server selects relevant advertisements based on the analysis results.

[0108] Step 16:

[0109] The server encodes the generated dialogue and advertisements as audio files.

[0110] Step 17:

[0111] The server delivers the encoded audio file to the user terminal.

[0112] Step 18:

[0113] The device plays the received audio file.

[0114] Step 19:

[0115] Allows the user to perform operations such as stop, skip, and play during playback.

[0116] Example 1

[0117] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0118] Conventional radio and audio content have faced challenges in providing content based on user interests, resulting in a lack of personalization for individual users. Furthermore, there were limited ways to respond appropriately to user feedback and questions, resulting in a one-way user experience. Furthermore, there was no established method for maximizing the effectiveness of advertising.

[0119] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0120] In this invention, the server includes means for collecting and saving user profile information, means for periodically transmitting user usage history data to the server, means for analyzing user interests and concerns based on the transmitted usage history data and profile information, means for selecting a topic based on the analysis result and generating conversation content using a generative AI model, means for converting the conversation content into an audio file using an appropriate voice synthesis model, means for receiving letters from users and analyzing the content to generate related information and answers, means for selecting relevant advertisements based on the user's interests and concerns and generating the conversation and advertisements as audio files, and means for delivering the generated audio files to the user terminal. This enables the provision of personalized content based on the user's interests and concerns, improving the user experience and maximizing advertising effectiveness.

[0121] "User Profile Information" means personal and preference information provided by a user, including name, age, hobbies, etc.

[0122] "Usage history data" is data generated when a user uses an application, including play time, skip counts, selected topics, and the like.

[0123] "Server" means a computer system that collects, stores, and analyzes data sent by users and provides necessary information.

[0124] A "generative AI model" is an artificial intelligence model for generating natural language text based on an input prompt.

[0125] A "speech synthesis model" is a system that includes algorithms and techniques for converting text data into speech data.

[0126] A "letter" is a text message such as a question or comment that a user sends through the application.

[0127] "Advertisements" are commercial information selected based on the user's interests and incorporated into content.

[0128] An "audio file" is a file in which audio data is stored in a digital format and is provided in a format that can be played on a user terminal.

[0129] "User terminal" means an electronic device used by a user to access the system, including a smartphone, tablet, or PC.

[0130] "Playback" refers to the act of the user terminal outputting an audio file as sound and the user listening to it.

[0131] The present invention relates to a system that provides personalized audio content based on a user's interests. This system generates and distributes content using a generative AI model and a voice synthesis model based on the user's profile information and usage history data. Specific embodiments of the system are described below.

[0132] User Data Collection and Initial Settings

[0133] User

[0134] First, a user installs the application and enters profile information such as name, age, hobbies, etc. This step lays the foundation for the system to provide a personalized experience tailored to the user.

[0135] Terminal

[0136] The device encrypts the profile information entered by the user and sends it to the server using the HTTPS protocol, ensuring the security of the user information during this process.

[0137] server

[0138] The server stores the received profile information in a database, which allows the integration and management of user-specific data for future analysis.

[0139] Specific examples

[0140] For example, if a user registers as "Yamada Taro" and selects "technology" as a hobby, this information is sent to the server and stored in a database.

[0141] ---

[0142] Daily data collection

[0143] Terminal

[0144] When a user uses an application, the device collects usage history data in real time, such as playback time, number of skips, and selected topics, allowing the application to always reflect the latest user behavior patterns.

[0145] server

[0146] The device sends the collected data to the server in batch processing at regular intervals (e.g., every 5 minutes). The server stores the received data in a database and analyzes it using AI algorithms. This allows the user's interests and concerns to be identified.

[0147] Specific examples

[0148] If a user frequently watches technology-related episodes and skips fitness-related episodes, the server may determine that the user has a strong interest in technology.

[0149] Natural Language Prompt Examples

[0150] "Data shows that users frequently watch technology-related episodes and skip fitness-related episodes."

[0151] ---

[0152] Generate personalized conversations

[0153] server

[0154] The server selects the most appropriate topic for the user based on the analyzed results, generates conversation content using a natural language generation model (e.g., OpenAI's GPT-4), and converts the generated text into speech using a speech synthesis model (e.g., Amazon Polly).

[0155] Specific examples

[0156] If the server determines from the analysis that the user is interested in technology, it selects the topic "Let's talk about the latest smartwatches." It inputs "Tell me about the latest smartwatches" as a prompt to GPT-4, and passes the generated text to Amazon Polly to generate a natural-sounding voice file.

[0157] Natural Language Prompt Examples

[0158] "Tell me about the latest smartwatch."

[0159] ---

[0160] Processing letters and providing information

[0161] User

[0162] Users can use the application's in-app message feature to send questions or comments, such as "What fitness apps do you recommend?"

[0163] Terminal

[0164] The device sends the message in real time to a server, and this information is also encrypted to ensure security.

[0165] server

[0166] The server analyzes the received letter and generates relevant information and answers based on its content. In addition, relevant advertisements are also selected and generated at the same time.

[0167] Specific examples

[0168] When a user sends a letter asking, "Please tell me your recommended fitness app," the server responds, "The recommended fitness app is XX," and also generates advertisements with information about special sales of fitness apps.

[0169] Natural Language Prompt Examples

[0170] What fitness apps do you recommend?

[0171] ---

[0172] Delivering personalized ads

[0173] server

[0174] The server selects relevant advertisements based on the user's profile information and current conversation content and encodes them as audio files.

[0175] Specific examples

[0176] If the server determines that the user is interested in technology, it selects an advertisement for the latest smartwatch, inputs a prompt into the GTP-4 saying, "Generate ad copy for the latest smartwatch," and uses Amazon Polly to convert the generated ad copy into an audio file that is seamlessly integrated into the conversation.

[0177] Natural Language Prompt Examples

[0178] "Generate ad copy for the latest smartwatch."

[0179] ---

[0180] Content Delivery and Playback

[0181] server

[0182] The generated audio file is then delivered to the user's device, taking into account the user's network conditions.

[0183] Terminal

[0184] The device plays the delivered audio file, and the user can perform operations during playback (such as stop, skip, and play).

[0185] User

[0186] Users listen to AI Radio and provide playback control and feedback as needed, allowing the system to continue to learn about the user's emerging interests and preferences.

[0187] Specific examples

[0188] The server delivers the generated audio file to the user's device, which then decodes and plays it. The user can press the "Stop" button or the "Skip" button during playback to move on to the next content.

[0189] ---

[0190] The implementation of this system will enable a personalized radio experience based on the user's interests. By appropriately utilizing generative AI models and prompts, more natural and engaging content will be provided to the user. This is expected to improve the user experience and maximize advertising effectiveness.

[0191] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0192] System program processing flow

[0193] Step 1: Collecting user data and initial setup

[0194] User

[0195] When a user installs the application and launches it for the first time, they enter their profile information, such as their name, age, and hobbies. By entering this information, the system is ready to provide a personalized experience for the user.

[0196] Terminal

[0197] The device encrypts the entered profile information and sends it to the server using the HTTPS protocol, ensuring the security of the user information during this process.

[0198] server

[0199] The server stores the received profile information in a database for later analysis.

[0200] Specific actions

[0201] For example, if a user registers as "Taro Yamada" and selects "technology" as a hobby, this information is encrypted and sent from the device to the server, which then stores the received data in the appropriate tables in the database.

[0202] Input and Output

[0203] Input: Profile information such as name, age, hobbies, etc.

[0204] Output: User profile information stored in a database

[0205] ---

[0206] Step 2: Collecting data on a daily basis

[0207] Terminal

[0208] When a user uses an application, the device collects usage history data such as playback time, number of skips, and selected topics, ensuring that the device always reflects the user's behavioral patterns in an up-to-date manner.

[0209] server

[0210] The device sends the collected usage history data at regular intervals (e.g., every 5 minutes) to the server, which receives it and stores it in a database. The server then uses AI algorithms to analyze the user's interests and concerns.

[0211] Specific actions

[0212] Each time a user listens to a radio episode, the device records data such as the start time, end time, number of skips, and selected topics. Every five minutes, the device sends this data to a server, which stores it in a database. An AI algorithm is used to analyze which topics the user is interested in.

[0213] Input and Output

[0214] Input: Usage history data such as play time, skip counts, and selected topics

[0215] Output: Analyzed user interests

[0216] ---

[0217] Step 3: Generate personalized conversations

[0218] server

[0219] The server selects the most appropriate topics for the user based on the analyzed user's interests, then generates the conversation using a natural language generation model (e.g., OpenAI's GPT-4), and converts the generated text into an audio file using an appropriate speech synthesis model (e.g., Amazon Polly).

[0220] Specific actions

[0221] For example, if the server determines that the user is interested in technology based on the analysis results, it inputs the prompt sentence "Tell me about the latest smartwatch" into GPT-4. The generated text is passed to Amazon Polly to generate an audio file.

[0222] Input and Output

[0223] Input: Topic analysis results based on user interests, prompt text

[0224] Output: Audio file (conversation content)

[0225] ---

[0226] Step 4: Processing your letter and providing information

[0227] User

[0228] Users can use the application's in-app message feature to send questions or comments, such as "What fitness apps do you recommend?"

[0229] Terminal

[0230] The device sends the message in real time to a server, and this information is also encrypted to ensure security.

[0231] server

[0232] The server analyzes the received letter and generates relevant information and answers based on its content. It also simultaneously selects and generates relevant advertisements.

[0233] Specific actions

[0234] For example, if a user sends a message asking, "What fitness apps do you recommend?", the device encrypts the question and sends it to the server, which uses GPT-4 to generate an answer and also selects and generates relevant advertisements.

[0235] Input and Output

[0236] Input: Letter from user

[0237] Output: Answers and related ads

[0238] ---

[0239] Step 5: Delivering personalized ads

[0240] server

[0241] The server selects relevant advertisements based on the user's profile information and current conversation content and encodes them as audio files.

[0242] Specific actions

[0243] For example, if the server determines that the user is interested in technology, it will generate an ad copy by inputting a prompt such as "Generate an ad copy for the latest smartwatch" into GPT-4. The generated ad copy is then converted into speech using Amazon Polly and incorporated into the conversation.

[0244] Input and Output

[0245] Input: User profile information, current conversation

[0246] Output: Audio file (conversation and advertisements)

[0247] ---

[0248] Step 6: Content Delivery and Playback

[0249] server

[0250] The server then delivers the generated audio files to the user's terminal, taking into account the user's network conditions.

[0251] Terminal

[0252] The device decodes and plays the delivered audio file, and the user can perform operations during playback (such as stop, skip, and play).

[0253] User

[0254] Users listen to AI Radio and provide playback control and feedback as needed, allowing the system to continue to learn about the user's emerging interests and preferences.

[0255] Specific actions

[0256] The server delivers the generated audio file to the user's device, which decodes and plays it. The user can press the "Stop" button during playback or the "Skip" button to move on to the next content.

[0257] Input and Output

[0258] Input: Generated audio file

[0259] Output: Played audio content

[0260] ---

[0261] This allows users to enjoy a more fulfilling radio experience through personalized content, while also maximizing advertising effectiveness and improving the user experience.

[0262] (Application example 1)

[0263] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0264] Conventional voice guidance systems often provide only uniform content, without providing sufficient personalized information based on the user's interests. They also lacked the ability to deliver advertisements based on the user's interests, making efficient marketing difficult. Furthermore, the lack of effective means for analyzing usage history and profile information limited the user experience.

[0265] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0266] In this invention, the server includes means for inputting user profile information, means for collecting and saving the input profile information, means for recording user usage history data, means for periodically transmitting the recorded usage history data to the server, means for analyzing user interests and concerns based on the transmitted usage history data and profile information, means for selecting topics and generating conversation content based on the analysis results, means for receiving letters from users and generating related information and responses after analyzing the content, means for selecting relevant advertisements based on the user's interests and concerns and generating the conversations and advertisements as audio files, means for delivering the generated audio files to user terminals, means for the user terminals to play the audio files, and means for generating and providing personalized audio guides for tourist spots and museums, thereby enabling improved user experience and efficient marketing.

[0267] "User Profile Information" means basic personal information and interest data provided by a User.

[0268] "Usage history data" is information related to behavior, such as operation history and playback history when a user uses an application.

[0269] A "server" is a computer system that collects, stores, and analyzes data sent by users.

[0270] "Analysis" refers to the processing of data to identify user interests based on collected profile information and usage history data.

[0271] A "topic" is a specific subject or theme that may be of interest to a user.

[0272] The "conversation content" is a voice message for dialogue with the user that is generated based on the analyzed data.

[0273] A "letter" is a message containing a question or comment from a user.

[0274] "Related information" is answers and information provided based on the user's letter.

[0275] "Advertisement" means commercial information provided based on the user's interests.

[0276] "Audio files" are generated conversations and advertisements encoded as audio data.

[0277] "Personalized audio guides at tourist spots and museums" are audio guides for tourist spots and museums that are customized according to the user's interests and concerns.

[0278] This invention is an AI radio system that develops conversations based on the user's interests and focuses on providing personalized audio guides, particularly at tourist spots and museums. The system is configured as follows:

[0279] User Data Collection and Initial Settings

[0280] The user first installs the application on their smartphone and enters their profile information when they first launch it. This profile information includes the user's name and genres of interest (e.g., history, art, technology, etc.). The entered profile information is sent from the user's device to the server, where it is stored in a database.

[0281] Daily data collection

[0282] When a user uses the application, usage history data (such as play time, number of skips, and selected topics) is recorded. This data is periodically sent from the device to a server, where it is stored and analyzed. The server then uses AI algorithms (e.g., machine learning models) to analyze this data and identify the user's interests.

[0283] Generate personalized conversations and audio guides

[0284] The server selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in history, the server selects a topic such as "historical exhibits in a specific museum." It then uses a natural language generation model (e.g., GPT-3) to generate audio guide content that includes emotion and tone. The generated content is then encoded as an audio file using a speech synthesis model (e.g., a text-to-speech engine).

[0285] Processing letters and providing information

[0286] Users can send questions and comments using the application's message function. For example, if a user asks, "What exhibits do you recommend at this museum?", the server receives the message, analyzes the content, and generates the most appropriate information and answer. At the same time, relevant advertisements are selected and generated as audio files.

[0287] Delivering personalized ads

[0288] The server selects relevant advertisements based on the user's profile information and current conversation content, and the generated conversation and advertisements are encoded as audio files and delivered to the user's device.

[0289] Content Delivery and Playback

[0290] The generated audio file is delivered from the server to the user's terminal, where the user can listen to it. The user can also perform operations (such as stopping, skipping, and playing) during playback.

[0291] Explanation based on concrete examples

[0292] For example, if a user is interested in "Japanese history," the app can provide an audio guide with the latest information on historical museum exhibits and related commentary. It can also provide personalized advertisements for books related to Japanese history.

[0293] An example prompt is:

[0294] "Prompt: User is interested in "Japanese history." Generate a detailed audio guide for a current museum exhibit. Also create related ads."

[0295] The present invention enables improved user experience and efficient marketing, and provides more comprehensive information at tourist spots and museums.

[0296] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0297] Step 1:

[0298] The user installs the application on their smartphone and enters their profile information (such as their name and genres of interest) when they first start it up. The entered profile information is sent from the user's device to the server.

[0299] Input: User profile information

[0300] What happens: A user enters information into an application and presses the "Submit" button.

[0301] Output: Profile information sent to the server

[0302] Step 2:

[0303] The server stores the received profile information in a database, which is used for analytics purposes.

[0304] Input: Profile information sent to the server

[0305] What it does: The server executes an SQL query that stores information in a database.

[0306] Output: Profile information stored in a database

[0307] Step 3:

[0308] When a user uses the application, usage history data (playback time, number of skips, selected topics, etc.) is recorded and periodically sent to the server.

[0309] Input: Application usage history data

[0310] How it works: The application records usage history in a log file or database and periodically sends it to the server.

[0311] Output: Usage history data sent to the server

[0312] Step 4:

[0313] The server uses AI algorithms (e.g., machine learning models) to analyze the user's interests and concerns based on usage history data and profile information.

[0314] Input: Submitted usage history data and profile information

[0315] How it works: The server inputs data into a machine learning model to estimate the user's interests.

[0316] Output: Analyzed user interests

[0317] Step 5:

[0318] Based on the analysis results, the server selects the most appropriate topic for the user and generates conversation content using a natural language generation model (e.g., GPT-3).

[0319] Input: Analyzed user interests

[0320] How it works: The server inputs topics into a natural language generation model and generates conversation content.

[0321] Output: Generated conversation

[0322] Step 6:

[0323] The server generates an audio file using a speech synthesis model (e.g., a text-to-speech engine) to add emotion and tone to the generated dialogue.

[0324] Input: Generated conversation

[0325] How it works: The server inputs text data into a speech synthesis model to generate speech data.

[0326] Output: Generated audio file

[0327] Step 7:

[0328] Users can submit questions and comments using the application's in-app message feature, such as "What exhibits would you recommend at this museum?"

[0329] Input: Letter from user

[0330] What happens: A user enters a question or comment into an application's input field and clicks the submit button.

[0331] Output: Letter sent to the server

[0332] Step 8:

[0333] The server analyzes the received letter, generates relevant information and answers, and selects relevant advertisements if necessary.

[0334] Input: Letter sent to the server

[0335] How it works: The server feeds the letter data into a machine learning model to generate the best answers and relevant ads.

[0336] Output: Generated answers and related ads

[0337] Step 9:

[0338] The server encodes the generated conversations, responses, and advertisements as audio files and delivers them to the user terminal.

[0339] Input: Generated conversations, responses, and advertisements

[0340] Operation: The server encodes the voice data and sends it to the user's terminal.

[0341] Output: Audio file sent to the user's device

[0342] Step 10:

[0343] The user terminal plays the delivered audio file, and the user listens to it. The user can perform operations (stop, skip, play, etc.) during playback.

[0344] Input: Streamed audio file

[0345] Action: The media player on the user's device plays an audio file, and the user operates the playback controls.

[0346] Output: Played audio file and user operation log

[0347] Through the above processing steps, users can always obtain information based on their own interests and concerns, allowing them to enjoy a more fulfilling experience at tourist spots and museums.

[0348] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0349] This invention is an AI radio system that combines an emotion engine that recognizes the user's emotions. Below, we will explain the system's program processing in natural language with specific examples.

[0350] User Data Collection and Initial Settings

[0351] User

[0352] First, the user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching the application for the first time.

[0353] Terminal

[0354] The entered profile information is transmitted from the user terminal to the server.

[0355] server

[0356] The server stores the received profile information in a database for later analysis.

[0357] Daily data collection

[0358] Terminal

[0359] Record usage history data (playback time, number of skips, selected topics, etc.) recorded when a user uses the application.

[0360] server

[0361] The server receives usage history data periodically sent from the device and stores it in a database. The server then analyzes this data using AI algorithms to identify the user's interests.

[0362] Generate personalized conversations

[0363] server

[0364] The server selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in technology, the server will select a topic such as "latest smartwatch features." Based on the selected topic, a natural language generation model is used to generate the conversation. At this time, an appropriate speech synthesis model is also applied to add emotion and tone.

[0365] Emotion recognition and conversation adjustment with emotion engine

[0366] Terminal

[0367] It analyzes the user's voice and input data in real time and recognizes their emotional state through an emotion engine.

[0368] server

[0369] The emotion engine receives emotional data analyzed from the device. The server adjusts the topic and conversation content in real time based on this emotional data. For example, if the user is feeling stressed, the server generates conversation content and a tone that will relax the user.

[0370] Processing letters and providing information

[0371] User

[0372] Users can submit questions or comments using the application's letter feature.

[0373] Terminal

[0374] A letter is sent from the user terminal to the server.

[0375] server

[0376] The server analyzes the received letter and generates related information and answers based on its content. For example, if a user asks for a "recommended fitness app," the server generates an answer such as "The recommended fitness app is XX." Related advertisements are also selected.

[0377] Delivering personalized ads

[0378] server

[0379] The server selects relevant advertisements based on the user's profile information and current conversation content, and encodes the generated conversation and advertisements as audio files.

[0380] Examples:

[0381] For technology-loving users, the server will select advertisements for the latest smartwatches and seamlessly integrate them into the conversation.

[0382] Content Delivery and Playback

[0383] server

[0384] The generated audio file is delivered to the user terminal.

[0385] Terminal

[0386] The user terminal plays the audio file and allows the user to perform operations (stop, skip, play, etc.).

[0387] User

[0388] Users listen to the AI ​​radio and provide playback control and feedback as needed.

[0389] Explanation based on concrete examples

[0390] For example, suppose a user sends a letter asking, "What fitness app do you recommend?" In this case, the server analyzes the letter, replies, "The recommended fitness app is ____," and generates a related advertisement (e.g., a sale on a fitness app). The generated audio file is delivered to the user's device, and when the user listens to it, the answer to the question and the advertisement information are simultaneously provided. In addition, by using an emotion engine, the server can provide a response with a tone and content that matches the emotional state the user was feeling when asking the question.

[0391] In this way, the present invention provides a personalized radio experience based on the user's interests, providing useful information while reducing visual burden. Furthermore, by integrating an emotion engine, the present invention can provide customized conversations and advertisements based on the user's current emotional state, further enhancing satisfaction.

[0392] The processing flow will be explained below.

[0393] Step 1:

[0394] The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[0395] Step 2:

[0396] The device sends the entered profile information to the server.

[0397] Step 3:

[0398] The server stores the received profile information in a database.

[0399] Step 4:

[0400] The device records usage history data such as playback time, number of skips, and selected topics when a user uses an application.

[0401] Step 5:

[0402] The terminal periodically transmits the recorded usage history data to the server.

[0403] Step 6:

[0404] The server receives the usage history data and stores it in a database.

[0405] Step 7:

[0406] The server uses AI algorithms to analyze users' interests based on their profile information and usage history data.

[0407] Step 8:

[0408] The device analyzes the user's voice and input data in real time and recognizes their emotional state through an emotion engine.

[0409] Step 9:

[0410] The terminal transmits the emotion data analyzed by the emotion engine to the server.

[0411] Step 10:

[0412] The server selects the most suitable topic for the user based on the analysis of their emotional data and interests. For example, if the user is feeling stressed, it will select relaxing topics.

[0413] Step 11:

[0414] The server generates conversation content based on the selected topic using a natural language generation model.

[0415] Step 12:

[0416] The server applies a speech synthesis model to add emotion and tone to the generated dialogue.

[0417] Step 13:

[0418] A user submits a question or comment using the in-app letter feature.

[0419] Step 14:

[0420] The terminal sends the letter from the user to the server.

[0421] Step 15:

[0422] The server receives the letter and analyzes its contents.

[0423] Step 16:

[0424] The server generates relevant information and answers based on the analysis results. For example, if a user asks, "What fitness app do you recommend?", the server will respond, "The recommended fitness app is XX."

[0425] Step 17:

[0426] The server selects relevant advertisements based on the analysis results.

[0427] Step 18:

[0428] The server encodes the generated dialogue and advertisements as audio files.

[0429] Step 19:

[0430] The server delivers the encoded audio file to the device.

[0431] Step 20:

[0432] The device plays the received audio file, allowing the user to perform operations such as stop, skip, and play during playback.

[0433] Step 21:

[0434] Users listen to what is being played and find that the content is personalized, for example, responses or advertisements tailored based on the user's current emotional state.

[0435] The above processing steps not only provide a personalized radio experience based on the user's interests, but also utilize an emotion engine to provide content optimized for the user's emotional state. This system reduces visual load while providing useful and satisfying information.

[0436] Example 2

[0437] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0438] Conventional voice information delivery systems have limited functionality for providing personalized content based on a user's profile information and usage history, resulting in issues with not being able to fully respond to user interests. Furthermore, they have not been able to adjust conversation content using real-time emotion recognition or dynamically generate content that reflects user feedback. Furthermore, there have been cases where advertising content and audio content do not always match, hindering the user experience.

[0439] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0440] In this invention, the server includes a means for collecting and saving user profile information, a means for periodically transmitting usage history data, and a means for analyzing the usage history data and profile information. This enables personalized conversation content based on the user's interests and the generation of related advertisements and their delivery as audio files. Furthermore, the user experience can be improved by recognizing the user's emotional state and adjusting the conversation content in real time.

[0441] "User profile information" is information including personal attributes of the user, such as age, sex, occupation, and hobbies.

[0442] "Means of collection and storage" refers to the technical means for receiving data entered by the user and storing it in a database or the like.

[0443] "Usage history data" refers to data such as playback time, number of skips, and selected topics that are recorded when a user uses an application.

[0444] "Means of analysis" refers to technical means using AI algorithms and machine learning models to identify user interests and concerns based on collected data.

[0445] "Means for generating conversational content" means the technical means for generating responses and conversations in natural language using a generative AI model based on the analysis results.

[0446] The "means for receiving letters and analyzing their contents" refers to the technical means for receiving questions and comments sent by users and processing them using a text analysis model.

[0447] The "means for generating relevant information and answers" refers to the technical means for generating appropriate information and answers based on the analyzed content of the letter.

[0448] "Means for selecting relevant advertisements and generating audio files of conversations and advertisements" refers to the technical means for selecting highly relevant advertisements based on the user's interests and generating audio data that integrates the advertisements into the conversation content.

[0449] The "means for playing audio files" refers to the technical means for appropriately playing audio files generated on a user terminal.

[0450] The "means for recognizing emotional states" refers to technical means for analyzing the user's voice and input data in real time and identifying the user's emotions.

[0451] A "conversational content adjustment mechanism" is a technological mechanism for changing the tone or content of a conversation in real time based on a perceived emotional state.

[0452] This invention is an AI radio system that combines an emotion engine that recognizes the user's emotions. The specific processing of the system program is shown below.

[0453] The present invention is a system for providing voice information based on user operations, and has the following main functions.

[0454] User Data Collection and Initial Settings

[0455] User:

[0456] First, the user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[0457] Device:

[0458] The device establishes an Internet connection and sends an HTTP request to the server to transmit the entered profile information to the server.

[0459] server:

[0460] The server receives the HTTP request and stores the profile information in a database, using a database management system such as MySQL.

[0461] Daily data collection

[0462] Device:

[0463] Every time the application is launched, it locally records user actions (playback time, skip count, selected topics, etc.) and saves the data in JSON format.

[0464] server:

[0465] The device sends usage history data to the server at regular intervals or when a specific event occurs. The server stores the received data in a database and processes it on a cloud server for data analysis.

[0466] Generate personalized conversations

[0467] server:

[0468] The server analyzes profile information and usage history data to select the most suitable topics for each user. It uses machine learning libraries such as Python's scikit-learn and TensorFlow to generate prompts for a natural language generation model (e.g., OpenAI GPT-3) and sends them to the model via an API.

[0469] Specific prompt examples:

[0470] "Users are interested in technology, so talk in detail about the latest smartwatch features."

[0471] The server uses the received generated text to generate an audio file using a speech synthesis model (e.g., Google Text-to-Speech) taking into account emotions and tone.

[0472] Emotion recognition and conversation adjustment with emotion engine

[0473] Device:

[0474] The device captures the user's voice input through a microphone and analyzes it in real time using an emotion engine (e.g., Affectiva Emotion AI).

[0475] server:

[0476] The system receives emotional data from the device and stores it in a database. Based on the analysis results, it adjusts the conversation content accordingly. For example, if the user is feeling stressed, it will regenerate the conversation content to provide a calming tone and relaxing topics.

[0477] Processing letters and providing information

[0478] User:

[0479] Users use the letter feature within the app to enter questions or comments and press the send button.

[0480] Device:

[0481] The terminal converts the input content into JSON format and sends it to the server as an HTTP POST request.

[0482] server:

[0483] The server analyzes the received letter using a text analysis model (e.g., spaCy), generates related information and recommendations using a generative AI model (GPT-3), generates the generated answers as audio files, and selects and integrates advertisements into the audio data.

[0484] Delivering personalized ads

[0485] server:

[0486] The server selects relevant ads based on the user's profile information and current conversation content, and encodes the generated conversation content and ads as audio files.

[0487] Specific examples of behavior:

[0488] For technology-loving users, ads about the latest smartwatch features are seamlessly integrated into the conversation.

[0489] Content Delivery and Playback

[0490] server:

[0491] The server sends the generated audio file as an HTTP response to deliver it to the user terminal using an HTTP framework (e.g., Flask).

[0492] Device:

[0493] The device uses the audio playback feature (e.g., AndroidMediaPlayer) to play the audio file, and the user can perform operations such as play, stop, and skip through the app interface.

[0494] User:

[0495] The user listens to the generated audio content and submits feedback or follow-up questions as needed.

[0496] In this way, the system of the present invention delivers personalized audio content based on the user's interests and generates finely tuned responses based on real-time emotion recognition, resulting in a more satisfying user experience.

[0497] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0498] Step 1: Collecting user data and initial setup

[0499] User

[0500] Input: The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[0501] Specific behavior: A user enters profile information into an application's input form and presses the submit button.

[0502] Terminal

[0503] Data processing: The entered profile information is temporarily saved locally.

[0504] Output: This information is converted to JSON format and sent to the server as an HTTP request.

[0505] What happens: The device establishes an Internet connection and generates an HTTP POST request containing the entered profile information.

[0506] server

[0507] Input: The server receives the profile information sent from the device.

[0508] Data processing: Analyze the received information and store it in a database.

[0509] Output: Generates a save confirmation message.

[0510] Specific operation: The server stores the received profile information using a database management system such as MySQL.

[0511] Step 2: Collecting data on a daily basis

[0512] Terminal

[0513] Input: Operational data when a user uses the application (playback time, number of skips, selected topics, etc.).

[0514] Data processing: Save usage history data in JSON format to local storage.

[0515] Output: Set a trigger to periodically send the saved data to the server.

[0516] Specific operation: Detects user operations within the application and stores data locally.

[0517] server

[0518] Input: Usage history data sent periodically from the device.

[0519] Data processing: The received usage history data is stored in a database and sent to an analysis server.

[0520] Output: Generates data analysis results.

[0521] Specific operation: The server analyzes the usage history data stored in the database using an AI algorithm to identify the user's interests and concerns.

[0522] Step 3: Generate personalized conversations

[0523] server

[0524] Input: Analysis of user profile information and usage history data.

[0525] Data processing: Based on the analysis results, prompt sentences are generated for the generative AI model (e.g., GPT-3).

[0526] Output: The generated conversation.

[0527] Specific operation: The server analyzes the data using Python's scikit-learn and TensorFlow and generates a prompt. The generated prompt is sent to the generative AI model via an API and an answer is received. "The user is interested in technology, so please talk in detail about the latest smartwatch features."

[0528] server

[0529] Input: The speech received from the generative AI model.

[0530] Data processing: Applying speech synthesis models to add emotion and tone to the conversation.

[0531] Output: An audio file of the generated conversation.

[0532] What it does: The server uses a speech synthesis model such as Google Text-to-Speech to convert the generated text into an audio file.

[0533] Step 4: Emotion recognition and conversation adjustment using the emotion engine

[0534] Terminal

[0535] Input: User speech and input data.

[0536] Data processing: Analyze emotional states in real time using an emotion engine.

[0537] Output: Emotion data as the analysis result.

[0538] Specific operation: The device captures the user's voice using a microphone device and analyzes it in real time using an emotion engine (e.g., Affectiva Emotion AI).

[0539] server

[0540] Input: Emotion data sent from the device.

[0541] Data processing: Adjusting conversation content in real time based on emotional data.

[0542] Output: Improved conversation.

[0543] Specific operation: Based on the analysis results, the server re-prompts the generative AI model as necessary and corrects the conversation content.

[0544] Step 5: Processing your letter and providing information

[0545] User

[0546] Input: User comments and questions.

[0547] Specific behavior: A user uses the letter feature within the app, enters a question or comment, and presses the send button.

[0548] Terminal

[0549] Input: The letter data entered by the user.

[0550] Data processing: Generate an HTTP POST request in JSON format and send it to the server.

[0551] Output: Send to server.

[0552] Specific operation: The device converts user input into JSON format and sends it to the server via the Internet.

[0553] server

[0554] Input: Letter data sent from the terminal.

[0555] Data processing: Analyze using text analytics models (e.g., spaCy) and generate relevant answers and information using generative AI models.

[0556] Output: The generated answer and related information.

[0557] Specific operation: The server receives the letter data, analyzes it using a natural language processing model, and creates an appropriate response using a generative AI model, which then generates it as an audio file.

[0558] Step 6: Delivering personalized ads

[0559] server

[0560] Input: User profile information and current conversation.

[0561] Data processing: Selecting relevant ads and merging the conversation and ads into an audio file.

[0562] Output: An audio file containing ads.

[0563] How it works: Based on the user information and conversation content, the server selects relevant advertisements and encodes them into audio files.

[0564] Specific examples

[0565] For technology-loving users, ads for the latest smartwatches are selected and seamlessly integrated into conversations.

[0566] Step 7: Content Delivery and Playback

[0567] server

[0568] Input: The generated audio file.

[0569] Data processing: Audio files are delivered to the user's device using the HTTP framework.

[0570] Output: Audio file delivery to user device.

[0571] Specific operation: The server uses an HTTP framework such as Flask to send the audio file to the terminal as an HTTP response.

[0572] Terminal

[0573] Input: The audio file sent from the server.

[0574] Data processing: Play audio files using the audio playback function.

[0575] Output: Playback of audio content.

[0576] Specific behavior: The device plays the audio file using AndroidMediaPlayer or similar, and allows the user to control playback through the interface.

[0577] User

[0578] What it does: The user listens to the generated audio content and provides feedback or asks follow-up questions if necessary. This data is used to generate future content.

[0579] (Application example 2)

[0580] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0581] Conventional radio systems and content distribution services have difficulty providing personalized content based on a user's current emotional state. Furthermore, they are unable to respond to a user's emotional changes in real time, resulting in a poor user experience. Furthermore, in the case of advertising delivery, insufficient personalization based on the user's current situation results in a failure to attract the user's attention and a reduced effectiveness.

[0582] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0583] In this invention, the server includes means for analyzing the user's voice and input data in real time and incorporating an emotion engine that recognizes the user's emotional state, means for adjusting topics and conversation content in real time based on the emotion data, and means for receiving letters from users, analyzing the content, and generating related information and answers, thereby making it possible to provide appropriate content in real time that matches the user's emotional state.

[0584] "User profile information" is basic personal information such as age, sex, occupation, and hobbies that a user enters when using the service for the first time.

[0585] "Usage history data" refers to behavioral data such as playback time, number of skips, and selected topics when a user uses an application.

[0586] The "emotion engine" is an engine that has the function of analyzing the user's voice and input data in real time and recognizing their emotional state.

[0587] The "server" is a computer system that stores and analyzes data collected from users and generates and adjusts topics and conversation content based on their emotional state and interests.

[0588] A "speech synthesis model" is a technology for generating speech based on computer-generated text and reproducing emotion and tone.

[0589] "Emotion data" is data that indicates the user's current emotional state as analyzed by the emotion engine.

[0590] "Advertisements" are commercial information that is selected based on the user's interests and incorporated naturally into conversations.

[0591] "Letters" are messages such as questions, comments, and feedback that users send through the application.

[0592] A "topic" is a subject of conversation selected by the server based on the user's interests and concerns.

[0593] The "audio file" is an audio data file in which the generated conversation content or advertisement content is encoded, and is played back on the user terminal.

[0594] In this invention, a system for exchanging data between a user terminal and a server to provide personalized content to the user will be specifically described. The detailed program processing for realizing this system will be described below.

[0595] User Data Collection and Initial Settings

[0596] After installing the application, the user enters their profile information the first time they start it. This information includes, for example, age, gender, occupation, and hobbies. This information is sent from the user's device to the server, which then stores it in a database. The hardware used at this stage is a smartphone, and the software includes front-end technology (e.g., React Native) that configures the user interface (UI) and a library (e.g., requests) that sends HTTP requests.

[0597] Daily data collection

[0598] When a user uses an application, usage history data such as playback time, number of skips, and selected topics is recorded. This data is periodically sent to a server, which stores the received data in a database. The server uses an AI algorithm to analyze this data and identify the user's interests. The software used for this analysis is a database management system (e.g., MySQL) and an AI analysis algorithm (e.g., TensorFlow).

[0599] Emotion recognition and conversation adjustment with emotion engine

[0600] The user device analyzes the user's voice and input data in real time and recognizes their emotional state using an emotion engine. The hardware used to collect voice data is the smartphone's microphone, and the software is a voice analysis library (e.g., Google Cloud Speech-to-Text API). The server analyzes the user's current emotional state based on the emotion data sent from the emotion engine and adjusts the topic and content of the conversation in real time. For example, if the user is feeling stressed, it generates a relaxing tone of voice.

[0601] Generate personalized conversations

[0602] The server selects the most suitable topic for the user based on the analysis results and generates the conversation using a natural language generation model, while also applying an appropriate speech synthesis model to add emotion and tone. For example, if the user prefers topics related to fitness, the server will generate a conversation about "recommended fitness apps."

[0603] Processing letters and providing information

[0604] Users can use the application's message feature to send questions or comments. For example, they can ask, "What fitness apps do you recommend?" These messages are sent from the user's device to a server, which analyzes the content and generates relevant information and answers. This analysis and generation is performed using a natural language processing model (e.g., GPT-3).

[0605] Delivering personalized ads

[0606] The server selects relevant ads based on the user's profile information and current conversation, and encodes the generated conversation and ads into audio files. For example, a technology-loving user might receive an ad for the latest wearable devices. The selected ads are seamlessly integrated into the conversation.

[0607] Content Delivery and Playback

[0608] The generated audio file is delivered from the server to the user's device, where it is played. When the user listens to AI Radio, they can control playback (pause, skip, play, etc.) as needed. The hardware used in this part is a smartphone, and the software is a media player library with audio playback capabilities.

[0609] Prompt Sentence Examples

[0610] For example, here's a prompt that describes a situation where the user is feeling stressed about fitness:

[0611] Example prompt sentence:

[0612] "A user who has asked a fitness question is stressed. We recommend fitness apps to that user in a relaxing tone."

[0613] In this way, it becomes possible to provide personalized content in real time according to the user's emotional state and interests. This invention not only improves the user experience, but also realizes highly personalized content delivery to increase advertising effectiveness.

[0614] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0615] Step 1:

[0616] The user installs the application and enters profile information when the application is first launched.

[0617] Examples of input: name, age, gender, occupation, hobbies, etc.

[0618] The terminal transmits the input data to the server, which stores the received data in a database.

[0619] Output: User profile information is stored in the server database.

[0620] Step 2:

[0621] When a user uses the application, usage history data such as playback time, number of skips, and selected topics is recorded.

[0622] Example inputs: playback data, skip counts, topic selection.

[0623] The terminal periodically transmits the collected data to the server, and the server stores the received data in a database.

[0624] Output: Usage history data is stored in the server database.

[0625] Step 3:

[0626] The server uses AI algorithms to analyze the collected profile information and usage history data to identify the user's interests.

[0627] Examples of input: profile information, usage history data.

[0628] The server uses AI algorithms (e.g., TensorFlow) to analyze the data.

[0629] Output: Analysis results based on user interests and concerns.

[0630] Step 4:

[0631] The server uses an emotion engine that analyzes the user's voice and input data in real time to recognize their emotional state.

[0632] Examples of input: user voice data, text input data.

[0633] The device sends the audio to an emotion engine (e.g., Google Cloud Speech-to-Text API) for emotion analysis.

[0634] Output: Emotion data indicating the user's current emotional state.

[0635] Step 5:

[0636] The server adjusts topics and conversation content in real time based on emotional data.

[0637] Examples of input: emotional data, historical usage data.

[0638] The server generates the conversation content using a natural language generation model (e.g., GPT-3) and makes any necessary adjustments.

[0639] Output: Conversational content that reflects the user's emotional state.

[0640] Step 6:

[0641] Users can submit questions or comments using the application's letter feature.

[0642] Example input: A user question or comment.

[0643] The terminal sends input from the user to the server, which analyzes the content.

[0644] Output: Analysis results and related information and answers.

[0645] Step 7:

[0646] The server selects relevant advertisements based on the user's profile information and current conversation content.

[0647] Examples of input: profile information, conversation content.

[0648] The server uses AI algorithms to select the most relevant ads for the user and incorporate them into the conversation.

[0649] Output: Advertising information embedded in the conversation.

[0650] Step 8:

[0651] The server encodes the generated conversations and advertisements as audio files and delivers them to the user terminal.

[0652] Examples of input: conversation content, advertising information.

[0653] The server uses a speech synthesis model (e.g., a speech synthesis engine) to generate an audio file and send it to the device.

[0654] Output: The audio file delivered to the user's device.

[0655] Step 9:

[0656] The user terminal plays the audio file delivered from the server and allows the user to operate the file (stop, skip, play, etc.) during playback.

[0657] Example input: An audio file delivered from a server.

[0658] The device plays audio files using a built-in media player and accepts user interaction through the UI.

[0659] Output: Audio playback with user interaction.

[0660] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0661] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0662] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0663] [Second embodiment]

[0664] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0665] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0666] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0667] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0668] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0669] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0670] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0671] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0672] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0673] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0674] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0675] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0676] This invention is an AI radio system that develops conversations based on the user's interests and concerns. The processing of the system's program is explained in natural language with specific examples.

[0677] User Data Collection and Initial Settings

[0678] User

[0679] First, the user installs the application and enters profile information when the application is first launched.

[0680] Terminal

[0681] The entered profile information is transmitted from the user terminal to the server.

[0682] server

[0683] The server stores the received profile information in a database for later analysis.

[0684] Daily data collection

[0685] Terminal

[0686] Record usage history data (playback time, number of skips, selected topics, etc.) recorded when a user uses the application.

[0687] server

[0688] The server receives usage history data periodically sent from the device and stores it in a database. The server then analyzes this data using AI algorithms to identify the user's interests.

[0689] Generate personalized conversations

[0690] server

[0691] The server then selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in technology, the server might select a topic such as "latest smartwatch features." Based on the selected topic, the server generates the conversation using a natural language generation model. At this time, an appropriate speech synthesis model is also applied to add emotion and tone.

[0692] Processing letters and providing information

[0693] User

[0694] Users can submit questions or comments using the application's letter feature.

[0695] Terminal

[0696] A letter is sent from the user terminal to the server.

[0697] server

[0698] The server analyzes the received letter and generates related information and answers based on its content. For example, if a user asks for a "recommended fitness app," the server generates an answer such as "The recommended fitness app is XX." Related advertisements are also selected.

[0699] Delivering personalized ads

[0700] server

[0701] The server selects relevant advertisements based on the user's profile information and current conversation content, and encodes the generated conversation and advertisements as audio files.

[0702] Examples:

[0703] For technology-loving users, the server will select advertisements for the latest smartwatches and seamlessly integrate them into the conversation.

[0704] Content Delivery and Playback

[0705] server

[0706] The generated audio file is delivered to the user terminal.

[0707] Terminal

[0708] The user terminal plays the audio file, allowing the user to listen to it, and allows the user to control the playback (stop, skip, play, etc.).

[0709] User

[0710] Users listen to the AI ​​radio and provide playback control and feedback as needed.

[0711] Explanation based on concrete examples

[0712] For example, suppose a user sends a letter asking, "What fitness app do you recommend?" In this case, the server analyzes the letter, replies, "The fitness app I recommend is ____," and generates a related advertisement (e.g., a sale on a fitness app). The generated audio file is then delivered to the user's device, and when the user listens to it, the answer to the question and the advertisement information are simultaneously provided.

[0713] The system enables a personalized radio experience based on the user's interests, provides useful information while reducing visual strain, and improves advertising effectiveness through personalized advertising.

[0714] The processing flow will be explained below.

[0715] Step 1:

[0716] The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[0717] Step 2:

[0718] The device sends the entered profile information to the server.

[0719] Step 3:

[0720] The server stores the received profile information in a database.

[0721] Step 4:

[0722] The device records data about the user's use of the application (playback time, number of skips, selected topics, etc.).

[0723] Step 5:

[0724] The terminal periodically transmits the recorded usage history data to the server.

[0725] Step 6:

[0726] The server receives the usage history data and stores it in a database.

[0727] Step 7:

[0728] The server uses AI algorithms to analyze users' interests based on their profile information and usage history data.

[0729] Step 8:

[0730] The server selects topics appropriate for the user based on the analysis results.

[0731] Step 9:

[0732] The server uses a natural language generation model to generate conversational content based on the selected topic.

[0733] Step 10:

[0734] The server applies a speech synthesis model to add emotion and tone to the generated dialogue.

[0735] Step 11:

[0736] A user submits a question or comment using the in-app letter feature.

[0737] Step 12:

[0738] The terminal sends the letter from the user to the server.

[0739] Step 13:

[0740] The server receives the letter and analyzes its contents.

[0741] Step 14:

[0742] The server generates relevant information and answers based on the analysis results.

[0743] Step 15:

[0744] The server selects relevant advertisements based on the analysis results.

[0745] Step 16:

[0746] The server encodes the generated dialogue and advertisements as audio files.

[0747] Step 17:

[0748] The server delivers the encoded audio file to the user terminal.

[0749] Step 18:

[0750] The device plays the received audio file.

[0751] Step 19:

[0752] Allows the user to perform operations such as stop, skip, and play during playback.

[0753] Example 1

[0754] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0755] Conventional radio and audio content have faced challenges in providing content based on user interests, resulting in a lack of personalization for individual users. Furthermore, there were limited ways to respond appropriately to user feedback and questions, resulting in a one-way user experience. Furthermore, there was no established method for maximizing the effectiveness of advertising.

[0756] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0757] In this invention, the server includes means for collecting and saving user profile information, means for periodically transmitting user usage history data to the server, means for analyzing user interests and concerns based on the transmitted usage history data and profile information, means for selecting a topic based on the analysis result and generating conversation content using a generative AI model, means for converting the conversation content into an audio file using an appropriate voice synthesis model, means for receiving letters from users and analyzing the content to generate related information and answers, means for selecting relevant advertisements based on the user's interests and concerns and generating the conversation and advertisements as audio files, and means for delivering the generated audio files to the user terminal. This enables the provision of personalized content based on the user's interests and concerns, improving the user experience and maximizing advertising effectiveness.

[0758] "User Profile Information" means personal and preference information provided by a user, including name, age, hobbies, etc.

[0759] "Usage history data" is data generated when a user uses an application, including play time, skip counts, selected topics, and the like.

[0760] "Server" means a computer system that collects, stores, and analyzes data sent by users and provides necessary information.

[0761] A "generative AI model" is an artificial intelligence model for generating natural language text based on an input prompt.

[0762] A "speech synthesis model" is a system that includes algorithms and techniques for converting text data into speech data.

[0763] A "letter" is a text message such as a question or comment that a user sends through the application.

[0764] "Advertisements" are commercial information selected based on the user's interests and incorporated into content.

[0765] An "audio file" is a file in which audio data is stored in a digital format and is provided in a format that can be played on a user terminal.

[0766] "User terminal" means an electronic device used by a user to access the system, including a smartphone, tablet, or PC.

[0767] "Playback" refers to the act of the user terminal outputting an audio file as sound and the user listening to it.

[0768] The present invention relates to a system that provides personalized audio content based on a user's interests. This system generates and distributes content using a generative AI model and a voice synthesis model based on the user's profile information and usage history data. Specific embodiments of the system are described below.

[0769] User Data Collection and Initial Settings

[0770] User

[0771] First, a user installs the application and enters profile information such as name, age, hobbies, etc. This step lays the foundation for the system to provide a personalized experience tailored to the user.

[0772] Terminal

[0773] The device encrypts the profile information entered by the user and sends it to the server using the HTTPS protocol, ensuring the security of the user information during this process.

[0774] server

[0775] The server stores the received profile information in a database, which allows the integration and management of user-specific data for future analysis.

[0776] Specific examples

[0777] For example, if a user registers as "Yamada Taro" and selects "technology" as a hobby, this information is sent to the server and stored in a database.

[0778] ---

[0779] Daily data collection

[0780] Terminal

[0781] When a user uses an application, the device collects usage history data in real time, such as playback time, number of skips, and selected topics, allowing the application to always reflect the latest user behavior patterns.

[0782] server

[0783] The device sends the collected data to the server in batch processing at regular intervals (e.g., every 5 minutes). The server stores the received data in a database and analyzes it using AI algorithms. This allows the user's interests and concerns to be identified.

[0784] Specific examples

[0785] If a user frequently watches technology-related episodes and skips fitness-related episodes, the server may determine that the user has a strong interest in technology.

[0786] Natural Language Prompt Examples

[0787] "Data shows that users frequently watch technology-related episodes and skip fitness-related episodes."

[0788] ---

[0789] Generate personalized conversations

[0790] server

[0791] The server selects the most appropriate topic for the user based on the analyzed results, generates conversation content using a natural language generation model (e.g., OpenAI's GPT-4), and converts the generated text into speech using a speech synthesis model (e.g., Amazon Polly).

[0792] Specific examples

[0793] If the server determines from the analysis that the user is interested in technology, it selects the topic "Let's talk about the latest smartwatches." It inputs "Tell me about the latest smartwatches" as a prompt to GPT-4, and passes the generated text to Amazon Polly to generate a natural-sounding voice file.

[0794] Natural Language Prompt Examples

[0795] "Tell me about the latest smartwatch."

[0796] ---

[0797] Processing letters and providing information

[0798] User

[0799] Users can use the application's in-app message feature to send questions or comments, such as "What fitness apps do you recommend?"

[0800] Terminal

[0801] The device sends the message in real time to a server, and this information is also encrypted to ensure security.

[0802] server

[0803] The server analyzes the received letter and generates relevant information and answers based on its content. In addition, relevant advertisements are also selected and generated at the same time.

[0804] Specific examples

[0805] When a user sends a letter saying, "Please tell me your recommended fitness app," the server responds, "The recommended fitness app is XX," and also generates advertisements with information about special sales of fitness apps.

[0806] Natural Language Prompt Examples

[0807] What fitness apps do you recommend?

[0808] ---

[0809] Delivering personalized ads

[0810] server

[0811] The server selects relevant advertisements based on the user's profile information and current conversation content and encodes them as audio files.

[0812] Specific examples

[0813] If the server determines that the user is interested in technology, it selects an advertisement for the latest smartwatch, inputs a prompt into the GTP-4 saying, "Generate ad copy for the latest smartwatch," and uses Amazon Polly to convert the generated ad copy into an audio file that is seamlessly integrated into the conversation.

[0814] Natural Language Prompt Examples

[0815] "Generate ad copy for the latest smartwatch."

[0816] ---

[0817] Content Delivery and Playback

[0818] server

[0819] The generated audio file is then delivered to the user's device, taking into account the user's network conditions.

[0820] Terminal

[0821] The device plays the delivered audio file, and the user can perform operations during playback (such as stop, skip, and play).

[0822] User

[0823] Users listen to AI Radio and provide playback control and feedback as needed, allowing the system to continue learning about the user's emerging interests and preferences.

[0824] Specific examples

[0825] The server delivers the generated audio file to the user's device, which then decodes and plays it. The user can press the "Stop" button or the "Skip" button during playback to move on to the next content.

[0826] ---

[0827] The implementation of this system will enable a personalized radio experience based on the user's interests. By appropriately utilizing generative AI models and prompts, more natural and engaging content will be provided to the user. This is expected to improve the user experience and maximize advertising effectiveness.

[0828] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0829] System program processing flow

[0830] Step 1: Collecting user data and initial setup

[0831] User

[0832] When a user installs the application and launches it for the first time, they enter their profile information, such as their name, age, and hobbies. By entering this information, the system is ready to provide a personalized experience for the user.

[0833] Terminal

[0834] The device encrypts the entered profile information and sends it to the server using the HTTPS protocol, ensuring the security of the user information during this process.

[0835] server

[0836] The server stores the received profile information in a database for later analysis.

[0837] Specific actions

[0838] For example, if a user registers as "Taro Yamada" and selects "technology" as a hobby, this information is encrypted and sent from the device to the server, which then stores the received data in the appropriate tables in the database.

[0839] Input and Output

[0840] Input: Profile information such as name, age, hobbies, etc.

[0841] Output: User profile information stored in a database

[0842] ---

[0843] Step 2: Collecting data on a daily basis

[0844] Terminal

[0845] When a user uses an application, the device collects usage history data such as playback time, number of skips, and selected topics, ensuring that the device always reflects the user's behavioral patterns in an up-to-date manner.

[0846] server

[0847] The device sends the collected usage history data at regular intervals (e.g., every 5 minutes) to the server, which receives it and stores it in a database. The server then uses AI algorithms to analyze the user's interests and concerns.

[0848] Specific actions

[0849] Each time a user listens to a radio episode, the device records data such as the start time, end time, number of skips, and selected topics. Every five minutes, the device sends this data to a server, which stores it in a database. An AI algorithm is used to analyze which topics the user is interested in.

[0850] Input and Output

[0851] Input: Usage history data such as play time, skip counts, and selected topics

[0852] Output: Analyzed user interests

[0853] ---

[0854] Step 3: Generate personalized conversations

[0855] server

[0856] The server selects the most appropriate topics for the user based on the analyzed user's interests, then generates the conversation using a natural language generation model (e.g., OpenAI's GPT-4), and converts the generated text into an audio file using an appropriate speech synthesis model (e.g., Amazon Polly).

[0857] Specific actions

[0858] For example, if the server determines that the user is interested in technology based on the analysis results, it inputs the prompt sentence "Tell me about the latest smartwatch" into GPT-4. The generated text is passed to Amazon Polly to generate an audio file.

[0859] Input and Output

[0860] Input: Topic analysis results based on user interests, prompt text

[0861] Output: Audio file (conversation content)

[0862] ---

[0863] Step 4: Processing your letter and providing information

[0864] User

[0865] Users can use the application's in-app message feature to send questions or comments, such as "What fitness apps do you recommend?"

[0866] Terminal

[0867] The device sends the message in real time to a server, and this information is also encrypted to ensure security.

[0868] server

[0869] The server analyzes the received letter and generates relevant information and answers based on its content. It also simultaneously selects and generates relevant advertisements.

[0870] Specific actions

[0871] For example, if a user sends a message asking, "What fitness apps do you recommend?", the device encrypts the question and sends it to the server, which uses GPT-4 to generate an answer and also selects and generates relevant advertisements.

[0872] Input and Output

[0873] Input: Letter from user

[0874] Output: Answers and related ads

[0875] ---

[0876] Step 5: Delivering personalized ads

[0877] server

[0878] The server selects relevant advertisements based on the user's profile information and current conversation content and encodes them as audio files.

[0879] Specific actions

[0880] For example, if the server determines that the user is interested in technology, it will generate an ad copy by inputting the prompt "Generate an ad copy for the latest smartwatch" into GPT-4. The generated ad copy is converted into speech by Amazon Polly and incorporated into the conversation.

[0881] Input and Output

[0882] Input: User profile information, current conversation

[0883] Output: Audio file (conversation and advertisements)

[0884] ---

[0885] Step 6: Content Delivery and Playback

[0886] server

[0887] The server then delivers the generated audio files to the user's terminal, taking into account the user's network conditions.

[0888] Terminal

[0889] The device decodes and plays the delivered audio file, and the user can perform operations during playback (such as stop, skip, and play).

[0890] User

[0891] Users listen to AI Radio and provide playback control and feedback as needed, allowing the system to continue learning about the user's emerging interests and preferences.

[0892] Specific actions

[0893] The server delivers the generated audio file to the user's device, which decodes and plays it. The user can press the "Stop" button during playback or the "Skip" button to move on to the next content.

[0894] Input and Output

[0895] Input: Generated audio file

[0896] Output: Played audio content

[0897] ---

[0898] This allows users to enjoy a more fulfilling radio experience through personalized content, while also maximizing advertising effectiveness and improving the user experience.

[0899] (Application example 1)

[0900] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0901] Conventional voice guidance systems often provide only uniform content, without providing sufficient personalized information based on the user's interests. They also lacked the ability to deliver advertisements based on the user's interests, making efficient marketing difficult. Furthermore, the lack of effective means for analyzing usage history and profile information limited the user experience.

[0902] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0903] In this invention, the server includes means for inputting user profile information, means for collecting and saving the input profile information, means for recording user usage history data, means for periodically transmitting the recorded usage history data to the server, means for analyzing user interests and concerns based on the transmitted usage history data and profile information, means for selecting topics and generating conversation content based on the analysis results, means for receiving letters from users and generating related information and responses after analyzing the content, means for selecting relevant advertisements based on the user's interests and concerns and generating the conversations and advertisements as audio files, means for delivering the generated audio files to user terminals, means for the user terminals to play the audio files, and means for generating and providing personalized audio guides for tourist spots and museums, thereby enabling improved user experience and efficient marketing.

[0904] "User Profile Information" means basic personal information and interest data provided by a User.

[0905] "Usage history data" is information related to behavior, such as operation history and playback history when a user uses an application.

[0906] A "server" is a computer system that collects, stores, and analyzes data sent by users.

[0907] "Analysis" refers to the processing of data to identify user interests based on collected profile information and usage history data.

[0908] A "topic" is a specific subject or theme that may be of interest to a user.

[0909] The "conversation content" is a voice message for dialogue with the user that is generated based on the analyzed data.

[0910] A "letter" is a message containing a question or comment from a user.

[0911] "Related information" is answers and information provided based on the user's letter.

[0912] "Advertisement" means commercial information provided based on the user's interests.

[0913] "Audio files" are generated conversations and advertisements encoded as audio data.

[0914] "Personalized audio guides at tourist spots and museums" are audio guides for tourist spots and museums that are customized according to the user's interests.

[0915] This invention is an AI radio system that develops conversations based on the user's interests and focuses on providing personalized audio guides, particularly at tourist spots and museums. The system is configured as follows:

[0916] User Data Collection and Initial Settings

[0917] The user first installs the application on their smartphone and enters their profile information when they first launch it. This profile information includes the user's name and genres of interest (e.g., history, art, technology, etc.). The entered profile information is sent from the user's device to the server, where it is stored in a database.

[0918] Daily data collection

[0919] When a user uses the application, usage history data (such as play time, number of skips, and selected topics) is recorded. This data is periodically sent from the device to a server, where it is stored and analyzed. The server then uses AI algorithms (e.g., machine learning models) to analyze this data and identify the user's interests.

[0920] Generate personalized conversations and audio guides

[0921] The server selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in history, the server selects a topic such as "historical exhibits in a specific museum." It then uses a natural language generation model (e.g., GPT-3) to generate audio guide content that includes emotion and tone. The generated content is then encoded as an audio file using a speech synthesis model (e.g., a text-to-speech engine).

[0922] Processing letters and providing information

[0923] Users can send questions and comments using the application's message function. For example, if a user asks, "What exhibits do you recommend at this museum?", the server receives the message, analyzes the content, and generates the most appropriate information and answer. At the same time, relevant advertisements are selected and generated as audio files.

[0924] Delivering personalized ads

[0925] The server selects relevant advertisements based on the user's profile information and current conversation content, and the generated conversation and advertisements are encoded as audio files and delivered to the user's device.

[0926] Content Delivery and Playback

[0927] The generated audio file is delivered from the server to the user's terminal, where the user can listen to it. The user can also perform operations (such as stopping, skipping, and playing) during playback.

[0928] Explanation based on concrete examples

[0929] For example, if a user is interested in "Japanese history," the app can provide an audio guide with the latest information on historical museum exhibits and related commentary. It can also provide personalized advertisements for books related to Japanese history.

[0930] An example prompt is:

[0931] "Prompt: User is interested in "Japanese history." Generate a detailed audio guide for a current museum exhibit. Also create related ads."

[0932] The present invention enables improved user experience and efficient marketing, and provides more comprehensive information at tourist spots and museums.

[0933] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0934] Step 1:

[0935] The user installs the application on their smartphone and enters their profile information (such as their name and genres of interest) when they first start it up. The entered profile information is sent from the user's device to the server.

[0936] Input: User profile information

[0937] What happens: A user enters information into an application and presses the "Submit" button.

[0938] Output: Profile information sent to the server

[0939] Step 2:

[0940] The server stores the received profile information in a database, which is used for analytics purposes.

[0941] Input: Profile information sent to the server

[0942] What it does: The server executes an SQL query that stores information in a database.

[0943] Output: Profile information stored in a database

[0944] Step 3:

[0945] When a user uses the application, usage history data (playback time, number of skips, selected topics, etc.) is recorded and periodically sent to the server.

[0946] Input: Application usage history data

[0947] How it works: The application records usage history in a log file or database and periodically sends it to the server.

[0948] Output: Usage history data sent to the server

[0949] Step 4:

[0950] The server uses AI algorithms (e.g., machine learning models) to analyze the user's interests and concerns based on usage history data and profile information.

[0951] Input: Submitted usage history data and profile information

[0952] How it works: The server inputs data into a machine learning model to estimate the user's interests.

[0953] Output: Analyzed user interests

[0954] Step 5:

[0955] Based on the analysis results, the server selects the most appropriate topic for the user and generates conversation content using a natural language generation model (e.g., GPT-3).

[0956] Input: Analyzed user interests

[0957] How it works: The server inputs topics into a natural language generation model and generates conversation content.

[0958] Output: Generated conversation

[0959] Step 6:

[0960] The server generates an audio file using a speech synthesis model (e.g., a text-to-speech engine) to add emotion and tone to the generated dialogue.

[0961] Input: Generated conversation

[0962] How it works: The server inputs text data into a speech synthesis model to generate speech data.

[0963] Output: Generated audio file

[0964] Step 7:

[0965] Users can submit questions and comments using the application's in-app message feature, such as "What exhibits would you recommend at this museum?"

[0966] Input: Letter from user

[0967] What happens: A user enters a question or comment into an application's input field and clicks the submit button.

[0968] Output: Letter sent to the server

[0969] Step 8:

[0970] The server analyzes the received letter, generates relevant information and answers, and selects relevant advertisements if necessary.

[0971] Input: Letter sent to the server

[0972] How it works: The server feeds the letter data into a machine learning model to generate the best answers and relevant ads.

[0973] Output: Generated answers and related ads

[0974] Step 9:

[0975] The server encodes the generated conversations, responses, and advertisements as audio files and delivers them to the user terminal.

[0976] Input: Generated conversations, responses, and advertisements

[0977] Operation: The server encodes the voice data and sends it to the user's terminal.

[0978] Output: Audio file sent to the user's device

[0979] Step 10:

[0980] The user terminal plays the delivered audio file, and the user listens to it. The user can perform operations (stop, skip, play, etc.) during playback.

[0981] Input: Streamed audio file

[0982] Action: The media player on the user's device plays an audio file, and the user operates the playback controls.

[0983] Output: Played audio file and user operation log

[0984] Through the above processing steps, users can always obtain information based on their interests and concerns, allowing them to enjoy a more fulfilling experience at tourist spots and museums.

[0985] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0986] This invention is an AI radio system that combines an emotion engine that recognizes the user's emotions. Below, we will explain the system's program processing in natural language with specific examples.

[0987] User Data Collection and Initial Settings

[0988] User

[0989] First, the user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching the application for the first time.

[0990] Terminal

[0991] The entered profile information is transmitted from the user terminal to the server.

[0992] server

[0993] The server stores the received profile information in a database for later analysis.

[0994] Daily data collection

[0995] Terminal

[0996] Record usage history data (playback time, number of skips, selected topics, etc.) recorded when a user uses the application.

[0997] server

[0998] The server receives usage history data periodically sent from the device and stores it in a database. The server then analyzes this data using AI algorithms to identify the user's interests.

[0999] Generate personalized conversations

[1000] server

[1001] The server selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in technology, the server will select a topic such as "latest smartwatch features." Based on the selected topic, a natural language generation model is used to generate the conversation. At this time, an appropriate speech synthesis model is also applied to add emotion and tone.

[1002] Emotion recognition and conversation adjustment with emotion engine

[1003] Terminal

[1004] It analyzes the user's voice and input data in real time and recognizes their emotional state through an emotion engine.

[1005] server

[1006] The emotion engine receives emotional data analyzed from the device. The server adjusts the topic and conversation content in real time based on this emotional data. For example, if the user is feeling stressed, the server generates conversation content and a tone that will relax the user.

[1007] Processing letters and providing information

[1008] User

[1009] Users can submit questions or comments using the application's letter feature.

[1010] Terminal

[1011] A letter is sent from the user terminal to the server.

[1012] server

[1013] The server analyzes the received letter and generates related information and answers based on its content. For example, if a user asks for a "recommended fitness app," the server generates an answer such as "The recommended fitness app is XX." Related advertisements are also selected.

[1014] Delivering personalized ads

[1015] server

[1016] The server selects relevant advertisements based on the user's profile information and current conversation content, and encodes the generated conversation and advertisements as audio files.

[1017] Examples:

[1018] For technology-loving users, the server will select advertisements for the latest smartwatches and seamlessly integrate them into the conversation.

[1019] Content Delivery and Playback

[1020] server

[1021] The generated audio file is delivered to the user terminal.

[1022] Terminal

[1023] The user terminal plays the audio file and allows the user to perform operations (stop, skip, play, etc.).

[1024] User

[1025] Users listen to the AI ​​radio and provide playback control and feedback as needed.

[1026] Explanation based on concrete examples

[1027] For example, suppose a user sends a letter asking, "What fitness app do you recommend?" In this case, the server analyzes the letter, replies, "The recommended fitness app is ____," and generates a related advertisement (e.g., a sale on a fitness app). The generated audio file is delivered to the user's device, and when the user listens to it, the answer to the question and the advertisement information are simultaneously provided. In addition, by using an emotion engine, the server can provide a response with a tone and content that matches the emotional state the user was feeling when asking the question.

[1028] In this way, the present invention provides a personalized radio experience based on the user's interests, providing useful information while reducing visual burden. Furthermore, by integrating an emotion engine, the present invention can provide customized conversations and advertisements based on the user's current emotional state, further enhancing satisfaction.

[1029] The processing flow will be explained below.

[1030] Step 1:

[1031] The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[1032] Step 2:

[1033] The device sends the entered profile information to the server.

[1034] Step 3:

[1035] The server stores the received profile information in a database.

[1036] Step 4:

[1037] The device records usage history data such as playback time, number of skips, and selected topics when a user uses an application.

[1038] Step 5:

[1039] The terminal periodically transmits the recorded usage history data to the server.

[1040] Step 6:

[1041] The server receives the usage history data and stores it in a database.

[1042] Step 7:

[1043] The server uses AI algorithms to analyze users' interests based on their profile information and usage history data.

[1044] Step 8:

[1045] The device analyzes the user's voice and input data in real time and recognizes their emotional state through an emotion engine.

[1046] Step 9:

[1047] The terminal transmits the emotion data analyzed by the emotion engine to the server.

[1048] Step 10:

[1049] The server selects the most suitable topic for the user based on the analysis of their emotional data and interests. For example, if the user is feeling stressed, it will select relaxing topics.

[1050] Step 11:

[1051] The server generates conversation content based on the selected topic using a natural language generation model.

[1052] Step 12:

[1053] The server applies a speech synthesis model to add emotion and tone to the generated dialogue.

[1054] Step 13:

[1055] A user submits a question or comment using the in-app letter feature.

[1056] Step 14:

[1057] The terminal sends the letter from the user to the server.

[1058] Step 15:

[1059] The server receives the letter and analyzes its contents.

[1060] Step 16:

[1061] The server generates relevant information and answers based on the analysis results. For example, if a user asks, "What fitness app do you recommend?", the server will respond, "The recommended fitness app is XX."

[1062] Step 17:

[1063] The server selects relevant advertisements based on the analysis results.

[1064] Step 18:

[1065] The server encodes the generated dialogue and advertisements as audio files.

[1066] Step 19:

[1067] The server delivers the encoded audio file to the device.

[1068] Step 20:

[1069] The device plays the received audio file, allowing the user to perform operations such as stop, skip, and play during playback.

[1070] Step 21:

[1071] Users listen to what is being played and find that the content is personalized, for example, responses or advertisements tailored based on the user's current emotional state.

[1072] The above processing steps not only provide a personalized radio experience based on the user's interests, but also utilize an emotion engine to provide content optimized for the user's emotional state. This system reduces visual load while providing useful and satisfying information.

[1073] Example 2

[1074] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1075] Conventional voice information delivery systems have limited functionality for providing personalized content based on a user's profile information and usage history, resulting in issues with not being able to fully respond to user interests. Furthermore, they have not been able to adjust conversation content using real-time emotion recognition or dynamically generate content that reflects user feedback. Furthermore, there have been cases where advertising content and audio content do not always match, hindering the user experience.

[1076] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1077] In this invention, the server includes a means for collecting and saving user profile information, a means for periodically transmitting usage history data, and a means for analyzing the usage history data and profile information. This enables personalized conversation content based on the user's interests and the generation of related advertisements and their delivery as audio files. Furthermore, the user experience can be improved by recognizing the user's emotional state and adjusting the conversation content in real time.

[1078] "User profile information" is information including personal attributes of the user, such as age, sex, occupation, and hobbies.

[1079] "Means of collection and storage" refers to the technical means for receiving data entered by the user and storing it in a database or the like.

[1080] "Usage history data" refers to data such as playback time, number of skips, and selected topics that are recorded when a user uses an application.

[1081] "Means of analysis" refers to technical means using AI algorithms and machine learning models to identify user interests and concerns based on collected data.

[1082] "Means for generating conversational content" means the technical means for creating natural language responses and conversations using a generative AI model based on the analysis results.

[1083] The "means for receiving letters and analyzing their contents" refers to the technical means for receiving questions and comments sent by users and processing them using a text analysis model.

[1084] The "means for generating relevant information and answers" refers to the technical means for generating appropriate information and answers based on the analyzed content of the letter.

[1085] "Means for selecting relevant advertisements and generating audio files of conversations and advertisements" refers to the technical means for selecting highly relevant advertisements based on the user's interests and generating audio data that integrates the advertisements into the conversation content.

[1086] The "means for playing audio files" refers to the technical means for appropriately playing audio files generated on a user terminal.

[1087] The "means for recognizing emotional states" refers to technical means for analyzing the user's voice and input data in real time and identifying the user's emotions.

[1088] A "conversational content adjustment mechanism" is a technological mechanism for changing the tone or content of a conversation in real time based on a perceived emotional state.

[1089] This invention is an AI radio system that combines an emotion engine that recognizes the user's emotions. The specific processing of the system program is shown below.

[1090] The present invention is a system for providing voice information based on user operations, and has the following main functions.

[1091] User Data Collection and Initial Settings

[1092] User:

[1093] First, the user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[1094] Device:

[1095] The device establishes an Internet connection and sends an HTTP request to the server to transmit the entered profile information to the server.

[1096] server:

[1097] The server receives the HTTP request and stores the profile information in a database, using a database management system such as MySQL.

[1098] Daily data collection

[1099] Device:

[1100] Every time the application is launched, it locally records user actions (playback time, skip count, selected topics, etc.) and saves the data in JSON format.

[1101] server:

[1102] The device sends usage history data to the server at regular intervals or when a specific event occurs. The server stores the received data in a database and processes it on a cloud server for data analysis.

[1103] Generate personalized conversations

[1104] server:

[1105] The server analyzes profile information and usage history data to select the most suitable topics for each user. It uses machine learning libraries such as Python's scikit-learn and TensorFlow to generate prompts for a natural language generation model (e.g., OpenAI GPT-3) and sends them to the model via an API.

[1106] Specific prompt examples:

[1107] "Users are interested in technology, so talk in detail about the latest smartwatch features."

[1108] The server uses the received generated text to generate an audio file using a speech synthesis model (e.g., Google Text-to-Speech) taking into account emotions and tone.

[1109] Emotion recognition and conversation adjustment with emotion engine

[1110] Device:

[1111] The device captures the user's voice input through a microphone and analyzes it in real time using an emotion engine (e.g., Affectiva Emotion AI).

[1112] server:

[1113] The system receives emotional data from the device and stores it in a database. Based on the analysis results, it adjusts the conversation content accordingly. For example, if the user is feeling stressed, it will regenerate the conversation content to provide a calming tone and relaxing topics.

[1114] Processing letters and providing information

[1115] User:

[1116] Users use the letter feature within the app to enter questions or comments and press the send button.

[1117] Device:

[1118] The terminal converts the input content into JSON format and sends it to the server as an HTTP POST request.

[1119] server:

[1120] The server analyzes the received letter using a text analysis model (e.g., spaCy), generates related information and recommendations using a generative AI model (GPT-3), generates the generated answers as audio files, and selects and integrates advertisements into the audio data.

[1121] Delivering personalized ads

[1122] server:

[1123] The server selects relevant ads based on the user's profile information and current conversation content, and encodes the generated conversation content and ads as audio files.

[1124] Specific examples of behavior:

[1125] For technology-loving users, ads about the latest smartwatch features are seamlessly integrated into the conversation.

[1126] Content Delivery and Playback

[1127] server:

[1128] The server sends the generated audio file as an HTTP response to deliver it to the user terminal using an HTTP framework (e.g., Flask).

[1129] Device:

[1130] The device uses the audio playback feature (e.g., AndroidMediaPlayer) to play the audio file, and the user can perform operations such as play, stop, and skip through the app interface.

[1131] User:

[1132] The user listens to the generated audio content and submits feedback or follow-up questions as needed.

[1133] In this way, the system of the present invention delivers personalized audio content based on the user's interests and generates finely tuned responses based on real-time emotion recognition, resulting in a more satisfying user experience.

[1134] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1135] Step 1: Collecting user data and initial setup

[1136] User

[1137] Input: The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[1138] Specific behavior: A user enters profile information into an application's input form and presses the submit button.

[1139] Terminal

[1140] Data processing: The entered profile information is temporarily saved locally.

[1141] Output: This information is converted to JSON format and sent to the server as an HTTP request.

[1142] What happens: The device establishes an Internet connection and generates an HTTP POST request containing the entered profile information.

[1143] server

[1144] Input: The server receives the profile information sent from the device.

[1145] Data processing: Analyze the received information and store it in a database.

[1146] Output: Generates a save confirmation message.

[1147] Specific operation: The server stores the received profile information using a database management system such as MySQL.

[1148] Step 2: Collecting data on a daily basis

[1149] Terminal

[1150] Input: Operational data when a user uses the application (playback time, number of skips, selected topics, etc.).

[1151] Data processing: Save usage history data in JSON format to local storage.

[1152] Output: Set a trigger to periodically send the saved data to the server.

[1153] Specific operation: Detects user operations within the application and stores data locally.

[1154] server

[1155] Input: Usage history data sent periodically from the device.

[1156] Data processing: The received usage history data is stored in a database and sent to an analysis server.

[1157] Output: Generates data analysis results.

[1158] Specific operation: The server analyzes the usage history data stored in the database using an AI algorithm to identify the user's interests and concerns.

[1159] Step 3: Generate personalized conversations

[1160] server

[1161] Input: Analysis of user profile information and usage history data.

[1162] Data processing: Based on the analysis results, prompt sentences are generated for the generative AI model (e.g., GPT-3).

[1163] Output: The generated conversation.

[1164] Specific operation: The server analyzes the data using Python's scikit-learn and TensorFlow and generates a prompt. The generated prompt is sent to the generative AI model via an API and an answer is received. "The user is interested in technology, so please talk in detail about the latest smartwatch features."

[1165] server

[1166] Input: The speech received from the generative AI model.

[1167] Data processing: Applying speech synthesis models to add emotion and tone to the conversation.

[1168] Output: An audio file of the generated conversation.

[1169] What it does: The server uses a speech synthesis model such as Google Text-to-Speech to convert the generated text into an audio file.

[1170] Step 4: Emotion recognition and conversation adjustment using the emotion engine

[1171] Terminal

[1172] Input: User speech and input data.

[1173] Data processing: Analyze emotional states in real time using an emotion engine.

[1174] Output: Emotion data as the analysis result.

[1175] Specific operation: The device captures the user's voice using a microphone device and analyzes it in real time using an emotion engine (e.g., Affectiva Emotion AI).

[1176] server

[1177] Input: Emotion data sent from the device.

[1178] Data processing: Adjusting conversation content in real time based on emotional data.

[1179] Output: Improved conversation.

[1180] Specific operation: Based on the analysis results, the server re-prompts the generative AI model as necessary and corrects the conversation content.

[1181] Step 5: Processing your letter and providing information

[1182] User

[1183] Input: User comments and questions.

[1184] What happens: A user uses the in-app message feature, enters a question or comment, and presses the send button.

[1185] Terminal

[1186] Input: The letter data entered by the user.

[1187] Data processing: Generate an HTTP POST request in JSON format and send it to the server.

[1188] Output: Send to server.

[1189] Specific operation: The device converts user input into JSON format and sends it to the server via the Internet.

[1190] server

[1191] Input: Letter data sent from the terminal.

[1192] Data processing: Analyze using text analytics models (e.g., spaCy) and generate relevant answers and information using generative AI models.

[1193] Output: The generated answer and related information.

[1194] Specific operation: The server receives the letter data, analyzes it using a natural language processing model, and creates an appropriate response using a generative AI model, which then generates it as an audio file.

[1195] Step 6: Delivering personalized ads

[1196] server

[1197] Input: User profile information and current conversation.

[1198] Data processing: Selecting relevant ads and merging the conversation and ads into an audio file.

[1199] Output: An audio file containing ads.

[1200] How it works: Based on the user information and conversation content, the server selects relevant advertisements and encodes them into audio files.

[1201] Specific examples

[1202] For technology-loving users, ads for the latest smartwatches are selected and seamlessly integrated into conversations.

[1203] Step 7: Content Delivery and Playback

[1204] server

[1205] Input: The generated audio file.

[1206] Data processing: Audio files are delivered to the user's device using the HTTP framework.

[1207] Output: Audio file delivery to user device.

[1208] Specific operation: The server uses an HTTP framework such as Flask to send the audio file to the terminal as an HTTP response.

[1209] Terminal

[1210] Input: The audio file sent from the server.

[1211] Data processing: Play audio files using the audio playback function.

[1212] Output: Playback of audio content.

[1213] Specific behavior: The device plays the audio file using AndroidMediaPlayer or similar, and allows the user to control playback through the interface.

[1214] User

[1215] What it does: The user listens to the generated audio content and provides feedback or asks follow-up questions if necessary. This data is used to generate future content.

[1216] (Application example 2)

[1217] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1218] Conventional radio systems and content distribution services have difficulty providing personalized content based on a user's current emotional state. Furthermore, they are unable to respond to a user's emotional changes in real time, resulting in a poor user experience. Furthermore, in the case of advertising delivery, insufficient personalization based on the user's current situation results in a failure to attract the user's attention and a reduced effectiveness.

[1219] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1220] In this invention, the server includes means for analyzing the user's voice and input data in real time and incorporating an emotion engine that recognizes the user's emotional state, means for adjusting topics and conversation content in real time based on the emotion data, and means for receiving letters from users, analyzing the content, and generating related information and answers, thereby making it possible to provide appropriate content in real time that matches the user's emotional state.

[1221] "User profile information" is basic personal information such as age, sex, occupation, and hobbies that a user enters when using the service for the first time.

[1222] "Usage history data" refers to behavioral data such as playback time, number of skips, and selected topics when a user uses an application.

[1223] The "emotion engine" is an engine that has the function of analyzing the user's voice and input data in real time and recognizing their emotional state.

[1224] The "server" is a computer system that stores and analyzes data collected from users and generates and adjusts topics and conversation content based on their emotional state and interests.

[1225] A "speech synthesis model" is a technology for generating speech based on computer-generated text and reproducing emotion and tone.

[1226] "Emotion data" is data that indicates the user's current emotional state as analyzed by the emotion engine.

[1227] "Advertisements" are commercial information that is selected based on the user's interests and incorporated naturally into conversations.

[1228] "Letters" are messages such as questions, comments, and feedback that users send through the application.

[1229] A "topic" is a conversation subject selected by the server based on the user's interests and concerns.

[1230] The "audio file" is an audio data file in which the generated conversation content or advertisement content is encoded, and is played back on the user terminal.

[1231] In this invention, a system for exchanging data between a user terminal and a server to provide personalized content to the user will be specifically described. The detailed program processing for realizing this system will be described below.

[1232] User Data Collection and Initial Settings

[1233] After installing the application, the user enters their profile information the first time they start it. This information includes, for example, age, gender, occupation, and hobbies. This information is sent from the user's device to the server, which then stores it in a database. The hardware used at this stage is a smartphone, and the software includes front-end technology (e.g., React Native) that configures the user interface (UI) and a library (e.g., requests) that sends HTTP requests.

[1234] Daily data collection

[1235] When a user uses an application, usage history data such as playback time, number of skips, and selected topics is recorded. This data is periodically sent to a server, which stores the received data in a database. The server uses an AI algorithm to analyze this data and identify the user's interests. The software used for this analysis is a database management system (e.g., MySQL) and an AI analysis algorithm (e.g., TensorFlow).

[1236] Emotion recognition and conversation adjustment with emotion engine

[1237] The user device analyzes the user's voice and input data in real time and recognizes their emotional state using an emotion engine. The hardware used to collect voice data is the smartphone's microphone, and the software is a voice analysis library (e.g., Google Cloud Speech-to-Text API). The server analyzes the user's current emotional state based on the emotion data sent from the emotion engine and adjusts the topic and content of the conversation in real time. For example, if the user is feeling stressed, it generates a relaxing tone of voice.

[1238] Generate personalized conversations

[1239] The server selects the most suitable topic for the user based on the analysis results and generates the conversation using a natural language generation model, while also applying an appropriate speech synthesis model to add emotion and tone. For example, if the user prefers topics related to fitness, the server will generate a conversation about "recommended fitness apps."

[1240] Processing letters and providing information

[1241] Users can use the application's message feature to send questions or comments. For example, they can ask, "What fitness apps do you recommend?" These messages are sent from the user's device to a server, which analyzes the content and generates relevant information and answers. This analysis and generation is performed using a natural language processing model (e.g., GPT-3).

[1242] Delivering personalized ads

[1243] The server selects relevant ads based on the user's profile information and current conversation, and encodes the generated conversation and ads into audio files. For example, a technology-loving user might receive an ad for the latest wearable devices. The selected ads are seamlessly integrated into the conversation.

[1244] Content Delivery and Playback

[1245] The generated audio file is delivered from the server to the user's device, where it is played. When the user listens to AI Radio, they can control playback (pause, skip, play, etc.) as needed. The hardware used in this part is a smartphone, and the software is a media player library with audio playback capabilities.

[1246] Prompt Sentence Examples

[1247] For example, here is a prompt that describes a situation where the user is feeling stressed about fitness:

[1248] Example prompt sentence:

[1249] "A user who has asked a fitness question is stressed. We recommend fitness apps to that user in a relaxing tone."

[1250] In this way, it becomes possible to provide personalized content in real time according to the user's emotional state and interests. This invention not only improves the user experience, but also realizes highly personalized content delivery to increase advertising effectiveness.

[1251] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1252] Step 1:

[1253] The user installs the application and enters profile information when the application is first launched.

[1254] Examples of input: name, age, gender, occupation, hobbies, etc.

[1255] The terminal transmits the input data to the server, which stores the received data in a database.

[1256] Output: User profile information is stored in the server database.

[1257] Step 2:

[1258] When a user uses the application, usage history data such as playback time, number of skips, and selected topics is recorded.

[1259] Example inputs: playback data, skip counts, topic selection.

[1260] The terminal periodically transmits the collected data to the server, and the server stores the received data in a database.

[1261] Output: Usage history data is stored in the server database.

[1262] Step 3:

[1263] The server uses AI algorithms to analyze the collected profile information and usage history data to identify the user's interests.

[1264] Examples of input: profile information, usage history data.

[1265] The server uses AI algorithms (e.g., TensorFlow) to analyze the data.

[1266] Output: Analysis results based on user interests and concerns.

[1267] Step 4:

[1268] The server uses an emotion engine that analyzes the user's voice and input data in real time to recognize their emotional state.

[1269] Examples of input: user voice data, text input data.

[1270] The device sends the audio to an emotion engine (e.g., Google Cloud Speech-to-Text API) for emotion analysis.

[1271] Output: Emotion data indicating the user's current emotional state.

[1272] Step 5:

[1273] The server adjusts topics and conversation content in real time based on emotional data.

[1274] Examples of input: emotional data, historical usage data.

[1275] The server generates the conversation using a natural language generation model (e.g., GPT-3) and makes any necessary adjustments.

[1276] Output: Conversational content that reflects the user's emotional state.

[1277] Step 6:

[1278] Users can submit questions or comments using the application's letter feature.

[1279] Example input: A user question or comment.

[1280] The terminal sends input from the user to the server, which analyzes the content.

[1281] Output: Analysis results and related information and answers.

[1282] Step 7:

[1283] The server selects relevant advertisements based on the user's profile information and current conversation content.

[1284] Examples of input: profile information, conversation content.

[1285] The server uses AI algorithms to select the most relevant ads for the user and incorporate them into the conversation.

[1286] Output: Advertising information embedded in the conversation.

[1287] Step 8:

[1288] The server encodes the generated conversations and advertisements as audio files and delivers them to the user terminal.

[1289] Examples of input: conversation content, advertising information.

[1290] The server uses a speech synthesis model (e.g., a speech synthesis engine) to generate an audio file and send it to the device.

[1291] Output: The audio file delivered to the user's device.

[1292] Step 9:

[1293] The user terminal plays the audio file delivered from the server and allows the user to operate the file (stop, skip, play, etc.) during playback.

[1294] Example input: An audio file delivered from a server.

[1295] The device plays audio files using a built-in media player and accepts user interaction through the UI.

[1296] Output: Audio playback with user interaction.

[1297] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1298] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1299] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1300] [Third embodiment]

[1301] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1302] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1303] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1304] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1305] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1306] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1307] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1308] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1309] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1310] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1311] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1312] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1313] This invention is an AI radio system that develops conversations based on the user's interests and concerns. The processing of the system's program is explained in natural language with specific examples.

[1314] User Data Collection and Initial Settings

[1315] User

[1316] First, the user installs the application and enters profile information when the application is first launched.

[1317] Terminal

[1318] The entered profile information is transmitted from the user terminal to the server.

[1319] server

[1320] The server stores the received profile information in a database for later analysis.

[1321] Daily data collection

[1322] Terminal

[1323] Record usage history data (playback time, number of skips, selected topics, etc.) recorded when a user uses the application.

[1324] server

[1325] The server receives usage history data periodically sent from the device and stores it in a database. The server then analyzes this data using AI algorithms to identify the user's interests.

[1326] Generate personalized conversations

[1327] server

[1328] The server then selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in technology, the server might select a topic such as "latest smartwatch features." Based on the selected topic, the server generates the conversation using a natural language generation model. At this time, an appropriate speech synthesis model is also applied to add emotion and tone.

[1329] Processing letters and providing information

[1330] User

[1331] Users can submit questions or comments using the application's letter feature.

[1332] Terminal

[1333] A letter is sent from the user terminal to the server.

[1334] server

[1335] The server analyzes the received letter and generates related information and answers based on its content. For example, if a user asks for a "recommended fitness app," the server generates an answer such as "The recommended fitness app is XX." Related advertisements are also selected.

[1336] Delivering personalized ads

[1337] server

[1338] The server selects relevant advertisements based on the user's profile information and current conversation content, and encodes the generated conversation and advertisements as audio files.

[1339] Examples:

[1340] For technology-loving users, the server will select advertisements for the latest smartwatches and seamlessly integrate them into the conversation.

[1341] Content Delivery and Playback

[1342] server

[1343] The generated audio file is delivered to the user terminal.

[1344] Terminal

[1345] The user terminal plays the audio file, allowing the user to listen to it, and allows the user to control the playback (stop, skip, play, etc.).

[1346] User

[1347] Users listen to the AI ​​radio and provide playback control and feedback as needed.

[1348] Explanation based on concrete examples

[1349] For example, suppose a user sends a letter asking, "What fitness app do you recommend?" In this case, the server analyzes the letter, replies, "The fitness app I recommend is ____," and generates a related advertisement (e.g., a sale on a fitness app). The generated audio file is then delivered to the user's device, and when the user listens to it, the answer to the question and the advertisement information are simultaneously provided.

[1350] The system enables a personalized radio experience based on the user's interests, provides useful information while reducing visual strain, and improves advertising effectiveness through personalized advertising.

[1351] The processing flow will be explained below.

[1352] Step 1:

[1353] The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[1354] Step 2:

[1355] The device sends the entered profile information to the server.

[1356] Step 3:

[1357] The server stores the received profile information in a database.

[1358] Step 4:

[1359] The device records data about the user's use of the application (playback time, number of skips, selected topics, etc.).

[1360] Step 5:

[1361] The terminal periodically transmits the recorded usage history data to the server.

[1362] Step 6:

[1363] The server receives the usage history data and stores it in a database.

[1364] Step 7:

[1365] The server uses AI algorithms to analyze users' interests based on their profile information and usage history data.

[1366] Step 8:

[1367] The server selects topics appropriate for the user based on the analysis results.

[1368] Step 9:

[1369] The server uses a natural language generation model to generate conversational content based on the selected topic.

[1370] Step 10:

[1371] The server applies a speech synthesis model to add emotion and tone to the generated dialogue.

[1372] Step 11:

[1373] A user submits a question or comment using the in-app letter feature.

[1374] Step 12:

[1375] The terminal sends the letter from the user to the server.

[1376] Step 13:

[1377] The server receives the letter and analyzes its contents.

[1378] Step 14:

[1379] The server generates relevant information and answers based on the analysis results.

[1380] Step 15:

[1381] The server selects relevant advertisements based on the analysis results.

[1382] Step 16:

[1383] The server encodes the generated dialogue and advertisements as audio files.

[1384] Step 17:

[1385] The server delivers the encoded audio file to the user terminal.

[1386] Step 18:

[1387] The device plays the received audio file.

[1388] Step 19:

[1389] Allows the user to perform operations such as stop, skip, and play during playback.

[1390] Example 1

[1391] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1392] Conventional radio and audio content have faced challenges in providing content based on user interests, resulting in a lack of personalization for individual users. Furthermore, there were limited ways to respond appropriately to user feedback and questions, resulting in a one-way user experience. Furthermore, there was no established method for maximizing the effectiveness of advertising.

[1393] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1394] In this invention, the server includes means for collecting and saving user profile information, means for periodically transmitting user usage history data to the server, means for analyzing user interests and concerns based on the transmitted usage history data and profile information, means for selecting a topic based on the analysis result and generating conversation content using a generative AI model, means for converting the conversation content into an audio file using an appropriate voice synthesis model, means for receiving letters from users and analyzing the content to generate related information and answers, means for selecting relevant advertisements based on the user's interests and concerns and generating the conversation and advertisements as audio files, and means for delivering the generated audio files to the user terminal. This enables the provision of personalized content based on the user's interests and concerns, improving the user experience and maximizing advertising effectiveness.

[1395] "User Profile Information" means personal and preference information provided by a user, including name, age, hobbies, etc.

[1396] "Usage history data" is data generated when a user uses an application, including play time, skip counts, selected topics, and the like.

[1397] "Server" means a computer system that collects, stores, and analyzes data sent by users and provides necessary information.

[1398] A "generative AI model" is an artificial intelligence model for generating natural language text based on an input prompt.

[1399] A "speech synthesis model" is a system that includes algorithms and techniques for converting text data into speech data.

[1400] A "letter" is a text message such as a question or comment that a user sends through the application.

[1401] "Advertisements" are commercial information selected based on the user's interests and incorporated into content.

[1402] An "audio file" is a file in which audio data is stored in a digital format and is provided in a format that can be played on a user terminal.

[1403] "User terminal" means an electronic device used by a user to access the system, including a smartphone, tablet, or PC.

[1404] "Playback" refers to the act of the user terminal outputting an audio file as sound and the user listening to it.

[1405] The present invention relates to a system that provides personalized audio content based on a user's interests. This system generates and distributes content using a generative AI model and a voice synthesis model based on the user's profile information and usage history data. Specific embodiments of the system are described below.

[1406] User Data Collection and Initial Settings

[1407] User

[1408] First, a user installs the application and enters profile information such as name, age, hobbies, etc. This step lays the foundation for the system to provide a personalized experience tailored to the user.

[1409] Terminal

[1410] The device encrypts the profile information entered by the user and sends it to the server using the HTTPS protocol, ensuring the security of the user information during this process.

[1411] server

[1412] The server stores the received profile information in a database, which allows the integration and management of user-specific data for future analysis.

[1413] Specific examples

[1414] For example, if a user registers as "Yamada Taro" and selects "technology" as a hobby, this information is sent to the server and stored in a database.

[1415] ---

[1416] Daily data collection

[1417] Terminal

[1418] When a user uses an application, the device collects usage history data in real time, such as playback time, number of skips, and selected topics, allowing the application to always reflect the latest user behavior patterns.

[1419] server

[1420] The device sends the collected data to the server in batch processing at regular intervals (e.g., every 5 minutes). The server stores the received data in a database and analyzes it using AI algorithms. This allows the user's interests and concerns to be identified.

[1421] Specific examples

[1422] If a user frequently watches technology-related episodes and skips fitness-related episodes, the server may determine that the user has a strong interest in technology.

[1423] Natural Language Prompt Examples

[1424] "Data shows that users frequently watch technology-related episodes and skip fitness-related episodes."

[1425] ---

[1426] Generate personalized conversations

[1427] server

[1428] The server selects the most appropriate topic for the user based on the analyzed results, generates conversation content using a natural language generation model (e.g., OpenAI's GPT-4), and converts the generated text into speech using a speech synthesis model (e.g., Amazon Polly).

[1429] Specific examples

[1430] If the server determines from the analysis that the user is interested in technology, it selects the topic "Let's talk about the latest smartwatches." It inputs "Tell me about the latest smartwatches" as a prompt to GPT-4, and passes the generated text to Amazon Polly to generate a natural-sounding voice file.

[1431] Natural Language Prompt Examples

[1432] "Tell me about the latest smartwatch."

[1433] ---

[1434] Processing letters and providing information

[1435] User

[1436] Users can use the application's in-app message feature to send questions or comments, such as "What fitness apps do you recommend?"

[1437] Terminal

[1438] The device sends the message in real time to a server, and this information is also encrypted to ensure security.

[1439] server

[1440] The server analyzes the received letter and generates relevant information and answers based on its content. In addition, relevant advertisements are also selected and generated at the same time.

[1441] Specific examples

[1442] When a user sends a letter saying, "Please tell me your recommended fitness app," the server responds, "The recommended fitness app is XX," and also generates advertisements with information about special sales of fitness apps.

[1443] Natural Language Prompt Examples

[1444] What fitness apps do you recommend?

[1445] ---

[1446] Delivering personalized ads

[1447] server

[1448] The server selects relevant advertisements based on the user's profile information and current conversation content and encodes them as audio files.

[1449] Specific examples

[1450] If the server determines that the user is interested in technology, it selects an advertisement for the latest smartwatch, inputs a prompt into the GTP-4 saying, "Generate ad copy for the latest smartwatch," and uses Amazon Polly to convert the generated ad copy into an audio file that is seamlessly integrated into the conversation.

[1451] Natural Language Prompt Examples

[1452] "Generate ad copy for the latest smartwatch."

[1453] ---

[1454] Content Delivery and Playback

[1455] server

[1456] The generated audio file is then delivered to the user's device, taking into account the user's network conditions.

[1457] Terminal

[1458] The device plays the delivered audio file, and the user can perform operations during playback (such as stop, skip, and play).

[1459] User

[1460] Users listen to AI Radio and provide playback control and feedback as needed, allowing the system to continue learning about the user's emerging interests and preferences.

[1461] Specific examples

[1462] The server delivers the generated audio file to the user's device, which then decodes and plays it. The user can press the "Stop" button or the "Skip" button during playback to move on to the next content.

[1463] ---

[1464] The implementation of this system will enable a personalized radio experience based on the user's interests. By appropriately utilizing generative AI models and prompts, more natural and engaging content will be provided to the user. This is expected to improve the user experience and maximize advertising effectiveness.

[1465] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1466] System program processing flow

[1467] Step 1: Collecting user data and initial setup

[1468] User

[1469] When a user installs the application and launches it for the first time, they enter their profile information, such as their name, age, and hobbies. By entering this information, the system is ready to provide a personalized experience for the user.

[1470] Terminal

[1471] The device encrypts the entered profile information and sends it to the server using the HTTPS protocol, ensuring the security of the user information during this process.

[1472] server

[1473] The server stores the received profile information in a database for later analysis.

[1474] Specific actions

[1475] For example, if a user registers as "Taro Yamada" and selects "technology" as a hobby, this information is encrypted and sent from the device to the server, which then stores the received data in the appropriate tables in the database.

[1476] Input and Output

[1477] Input: Profile information such as name, age, hobbies, etc.

[1478] Output: User profile information stored in a database

[1479] ---

[1480] Step 2: Collecting data on a daily basis

[1481] Terminal

[1482] When a user uses an application, the device collects usage history data such as playback time, number of skips, and selected topics, ensuring that the device always reflects the user's behavioral patterns in an up-to-date manner.

[1483] server

[1484] The device sends the collected usage history data at regular intervals (e.g., every 5 minutes) to the server, which receives it and stores it in a database. The server then uses AI algorithms to analyze the user's interests and concerns.

[1485] Specific actions

[1486] Each time a user listens to a radio episode, the device records data such as the start time, end time, number of skips, and selected topics. Every five minutes, the device sends this data to a server, which stores it in a database. An AI algorithm is used to analyze which topics the user is interested in.

[1487] Input and Output

[1488] Input: Usage history data such as play time, skip counts, and selected topics

[1489] Output: Analyzed user interests

[1490] ---

[1491] Step 3: Generate personalized conversations

[1492] server

[1493] The server selects the most appropriate topics for the user based on the analyzed user's interests, then generates the conversation using a natural language generation model (e.g., OpenAI's GPT-4), and converts the generated text into an audio file using an appropriate speech synthesis model (e.g., Amazon Polly).

[1494] Specific actions

[1495] For example, if the server determines that the user is interested in technology based on the analysis results, it inputs the prompt sentence "Tell me about the latest smartwatch" into GPT-4. The generated text is passed to Amazon Polly to generate an audio file.

[1496] Input and Output

[1497] Input: Topic analysis results based on user interests, prompt text

[1498] Output: Audio file (conversation content)

[1499] ---

[1500] Step 4: Processing your letter and providing information

[1501] User

[1502] Users can use the application's in-app message feature to send questions or comments, such as "What fitness apps do you recommend?"

[1503] Terminal

[1504] The device sends the message in real time to a server, and this information is also encrypted to ensure security.

[1505] server

[1506] The server analyzes the received letter and generates relevant information and answers based on its content. It also simultaneously selects and generates relevant advertisements.

[1507] Specific actions

[1508] For example, if a user sends a message asking, "What fitness apps do you recommend?", the device encrypts the question and sends it to the server, which uses GPT-4 to generate an answer and also selects and generates relevant advertisements.

[1509] Input and Output

[1510] Input: Letter from user

[1511] Output: Answers and related ads

[1512] ---

[1513] Step 5: Delivering personalized ads

[1514] server

[1515] The server selects relevant advertisements based on the user's profile information and current conversation content and encodes them as audio files.

[1516] Specific actions

[1517] For example, if the server determines that the user is interested in technology, it will generate an ad copy by inputting the prompt "Generate an ad copy for the latest smartwatch" into GPT-4. The generated ad copy is converted into speech by Amazon Polly and incorporated into the conversation.

[1518] Input and Output

[1519] Input: User profile information, current conversation

[1520] Output: Audio file (conversation and advertisements)

[1521] ---

[1522] Step 6: Content Delivery and Playback

[1523] server

[1524] The server then delivers the generated audio files to the user's terminal, taking into account the user's network conditions.

[1525] Terminal

[1526] The device decodes and plays the delivered audio file, and the user can perform operations during playback (such as stop, skip, and play).

[1527] User

[1528] Users listen to AI Radio and provide playback control and feedback as needed, allowing the system to continue learning about the user's emerging interests and preferences.

[1529] Specific actions

[1530] The server delivers the generated audio file to the user's device, which decodes and plays it. The user can press the "Stop" button during playback or the "Skip" button to move on to the next content.

[1531] Input and Output

[1532] Input: Generated audio file

[1533] Output: Played audio content

[1534] ---

[1535] This allows users to enjoy a more fulfilling radio experience through personalized content, while also maximizing advertising effectiveness and improving the user experience.

[1536] (Application example 1)

[1537] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1538] Conventional voice guidance systems often provide only uniform content, without providing sufficient personalized information based on the user's interests. They also lacked the ability to deliver advertisements based on the user's interests, making efficient marketing difficult. Furthermore, the lack of effective means for analyzing usage history and profile information limited the user experience.

[1539] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1540] In this invention, the server includes means for inputting user profile information, means for collecting and saving the input profile information, means for recording user usage history data, means for periodically transmitting the recorded usage history data to the server, means for analyzing user interests and concerns based on the transmitted usage history data and profile information, means for selecting topics and generating conversation content based on the analysis results, means for receiving letters from users and generating related information and responses after analyzing the content, means for selecting relevant advertisements based on the user's interests and concerns and generating the conversations and advertisements as audio files, means for delivering the generated audio files to user terminals, means for the user terminals to play the audio files, and means for generating and providing personalized audio guides for tourist spots and museums, thereby enabling improved user experience and efficient marketing.

[1541] "User Profile Information" means basic personal information and interest data provided by a User.

[1542] "Usage history data" is information related to behavior, such as operation history and playback history when a user uses an application.

[1543] A "server" is a computer system that collects, stores, and analyzes data sent by users.

[1544] "Analysis" refers to the processing of data to identify user interests based on collected profile information and usage history data.

[1545] A "topic" is a specific subject or theme that may be of interest to a user.

[1546] The "conversation content" is a voice message for dialogue with the user that is generated based on the analyzed data.

[1547] A "letter" is a message containing a question or comment from a user.

[1548] "Related information" is answers and information provided based on the user's letter.

[1549] "Advertisement" means commercial information provided based on the user's interests.

[1550] "Audio files" are generated conversations and advertisements encoded as audio data.

[1551] "Personalized audio guides at tourist spots and museums" are audio guides for tourist spots and museums that are customized according to the user's interests.

[1552] This invention is an AI radio system that develops conversations based on the user's interests and focuses on providing personalized audio guides, particularly at tourist spots and museums. The system is configured as follows:

[1553] User Data Collection and Initial Settings

[1554] The user first installs the application on their smartphone and enters their profile information when they first launch it. This profile information includes the user's name and genres of interest (e.g., history, art, technology, etc.). The entered profile information is sent from the user's device to the server, where it is stored in a database.

[1555] Daily data collection

[1556] When a user uses the application, usage history data (such as play time, number of skips, and selected topics) is recorded. This data is periodically sent from the device to a server, where it is stored and analyzed. The server then uses AI algorithms (e.g., machine learning models) to analyze this data and identify the user's interests.

[1557] Generate personalized conversations and audio guides

[1558] The server selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in history, the server selects a topic such as "historical exhibits in a specific museum." It then uses a natural language generation model (e.g., GPT-3) to generate audio guide content that includes emotion and tone. The generated content is then encoded as an audio file using a speech synthesis model (e.g., a text-to-speech engine).

[1559] Processing letters and providing information

[1560] Users can send questions and comments using the application's message function. For example, if a user asks, "What exhibits do you recommend at this museum?", the server receives the message, analyzes the content, and generates the most appropriate information and answer. At the same time, relevant advertisements are selected and generated as audio files.

[1561] Delivering personalized ads

[1562] The server selects relevant advertisements based on the user's profile information and current conversation content, and the generated conversation and advertisements are encoded as audio files and delivered to the user's device.

[1563] Content Delivery and Playback

[1564] The generated audio file is delivered from the server to the user's terminal, where the user can listen to it. The user can also perform operations (such as stopping, skipping, and playing) during playback.

[1565] Explanation based on concrete examples

[1566] For example, if a user is interested in "Japanese history," the app can provide an audio guide with the latest information on historical museum exhibits and related commentary. It can also provide personalized advertisements for books related to Japanese history.

[1567] An example prompt is:

[1568] "Prompt: User is interested in "Japanese history." Generate a detailed audio guide for a current museum exhibit. Also create related ads."

[1569] The present invention enables improved user experience and efficient marketing, and provides more comprehensive information at tourist spots and museums.

[1570] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1571] Step 1:

[1572] The user installs the application on their smartphone and enters their profile information (such as their name and genres of interest) when they first start it up. The entered profile information is sent from the user's device to the server.

[1573] Input: User profile information

[1574] What happens: A user enters information into an application and presses the "Submit" button.

[1575] Output: Profile information sent to the server

[1576] Step 2:

[1577] The server stores the received profile information in a database, which is used for analytics purposes.

[1578] Input: Profile information sent to the server

[1579] What it does: The server executes an SQL query that stores information in a database.

[1580] Output: Profile information stored in a database

[1581] Step 3:

[1582] When a user uses the application, usage history data (playback time, number of skips, selected topics, etc.) is recorded and periodically sent to the server.

[1583] Input: Application usage history data

[1584] How it works: The application records usage history in a log file or database and periodically sends it to the server.

[1585] Output: Usage history data sent to the server

[1586] Step 4:

[1587] The server uses AI algorithms (e.g., machine learning models) to analyze the user's interests and concerns based on usage history data and profile information.

[1588] Input: Submitted usage history data and profile information

[1589] How it works: The server inputs data into a machine learning model to estimate the user's interests.

[1590] Output: Analyzed user interests

[1591] Step 5:

[1592] Based on the analysis results, the server selects the most appropriate topic for the user and generates conversation content using a natural language generation model (e.g., GPT-3).

[1593] Input: Analyzed user interests

[1594] How it works: The server inputs topics into a natural language generation model and generates conversation content.

[1595] Output: Generated conversation

[1596] Step 6:

[1597] The server generates an audio file using a speech synthesis model (e.g., a text-to-speech engine) to add emotion and tone to the generated dialogue.

[1598] Input: Generated conversation

[1599] How it works: The server inputs text data into a speech synthesis model to generate speech data.

[1600] Output: Generated audio file

[1601] Step 7:

[1602] Users can submit questions and comments using the application's in-app message feature, such as "What exhibits would you recommend at this museum?"

[1603] Input: Letter from user

[1604] What happens: A user enters a question or comment into an application's input field and clicks the submit button.

[1605] Output: Letter sent to the server

[1606] Step 8:

[1607] The server analyzes the received letter, generates relevant information and answers, and selects relevant advertisements if necessary.

[1608] Input: Letter sent to the server

[1609] How it works: The server feeds the letter data into a machine learning model to generate the best answers and relevant ads.

[1610] Output: Generated answers and related ads

[1611] Step 9:

[1612] The server encodes the generated conversations, responses, and advertisements as audio files and delivers them to the user terminal.

[1613] Input: Generated conversations, responses, and advertisements

[1614] Operation: The server encodes the voice data and sends it to the user's terminal.

[1615] Output: Audio file sent to the user's device

[1616] Step 10:

[1617] The user terminal plays the delivered audio file, and the user listens to it. The user can perform operations (stop, skip, play, etc.) during playback.

[1618] Input: Streamed audio file

[1619] Action: The media player on the user's device plays an audio file, and the user operates the playback controls.

[1620] Output: Played audio file and user operation log

[1621] Through the above processing steps, users can always obtain information based on their interests and concerns, allowing them to enjoy a more fulfilling experience at tourist spots and museums.

[1622] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1623] This invention is an AI radio system that combines an emotion engine that recognizes the user's emotions. Below, we will explain the system's program processing in natural language with specific examples.

[1624] User Data Collection and Initial Settings

[1625] User

[1626] First, the user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching the application for the first time.

[1627] Terminal

[1628] The entered profile information is transmitted from the user terminal to the server.

[1629] server

[1630] The server stores the received profile information in a database for later analysis.

[1631] Daily data collection

[1632] Terminal

[1633] Record usage history data (playback time, number of skips, selected topics, etc.) recorded when a user uses the application.

[1634] server

[1635] The server receives usage history data periodically sent from the device and stores it in a database. The server then analyzes this data using AI algorithms to identify the user's interests.

[1636] Generate personalized conversations

[1637] server

[1638] The server selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in technology, the server will select a topic such as "latest smartwatch features." Based on the selected topic, a natural language generation model is used to generate the conversation. At this time, an appropriate speech synthesis model is also applied to add emotion and tone.

[1639] Emotion recognition and conversation adjustment with emotion engine

[1640] Terminal

[1641] It analyzes the user's voice and input data in real time and recognizes their emotional state through an emotion engine.

[1642] server

[1643] The emotion engine receives emotional data analyzed from the device. The server adjusts the topic and conversation content in real time based on this emotional data. For example, if the user is feeling stressed, the server generates conversation content and a tone that will relax the user.

[1644] Processing letters and providing information

[1645] User

[1646] Users can submit questions or comments using the application's letter feature.

[1647] Terminal

[1648] A letter is sent from the user terminal to the server.

[1649] server

[1650] The server analyzes the received letter and generates related information and answers based on its content. For example, if a user asks for a "recommended fitness app," the server generates an answer such as "The recommended fitness app is XX." Related advertisements are also selected.

[1651] Delivering personalized ads

[1652] server

[1653] The server selects relevant advertisements based on the user's profile information and current conversation content, and encodes the generated conversation and advertisements as audio files.

[1654] Examples:

[1655] For technology-loving users, the server will select advertisements for the latest smartwatches and seamlessly integrate them into the conversation.

[1656] Content Delivery and Playback

[1657] server

[1658] The generated audio file is delivered to the user terminal.

[1659] Terminal

[1660] The user terminal plays the audio file and allows the user to perform operations (stop, skip, play, etc.).

[1661] User

[1662] Users listen to the AI ​​radio and provide playback control and feedback as needed.

[1663] Explanation based on concrete examples

[1664] For example, suppose a user sends a letter asking, "What fitness app do you recommend?" In this case, the server analyzes the letter, replies, "The recommended fitness app is ____," and generates a related advertisement (e.g., a sale on a fitness app). The generated audio file is delivered to the user's device, and when the user listens to it, the answer to the question and the advertisement information are simultaneously provided. In addition, by using an emotion engine, the server can provide a response with a tone and content that matches the emotional state the user was feeling when asking the question.

[1665] In this way, the present invention provides a personalized radio experience based on the user's interests, providing useful information while reducing visual burden. Furthermore, by integrating an emotion engine, the present invention can provide customized conversations and advertisements based on the user's current emotional state, further enhancing satisfaction.

[1666] The processing flow will be explained below.

[1667] Step 1:

[1668] The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[1669] Step 2:

[1670] The device sends the entered profile information to the server.

[1671] Step 3:

[1672] The server stores the received profile information in a database.

[1673] Step 4:

[1674] The device records usage history data such as playback time, number of skips, and selected topics when a user uses an application.

[1675] Step 5:

[1676] The terminal periodically transmits the recorded usage history data to the server.

[1677] Step 6:

[1678] The server receives the usage history data and stores it in a database.

[1679] Step 7:

[1680] The server uses AI algorithms to analyze users' interests based on their profile information and usage history data.

[1681] Step 8:

[1682] The device analyzes the user's voice and input data in real time and recognizes their emotional state through an emotion engine.

[1683] Step 9:

[1684] The terminal transmits the emotion data analyzed by the emotion engine to the server.

[1685] Step 10:

[1686] The server selects the most suitable topic for the user based on the analysis of their emotional data and interests. For example, if the user is feeling stressed, it will select relaxing topics.

[1687] Step 11:

[1688] The server generates conversation content based on the selected topic using a natural language generation model.

[1689] Step 12:

[1690] The server applies a speech synthesis model to add emotion and tone to the generated dialogue.

[1691] Step 13:

[1692] A user submits a question or comment using the in-app letter feature.

[1693] Step 14:

[1694] The terminal sends the letter from the user to the server.

[1695] Step 15:

[1696] The server receives the letter and analyzes its contents.

[1697] Step 16:

[1698] The server generates relevant information and answers based on the analysis results. For example, if a user asks, "What fitness app do you recommend?", the server will respond, "The recommended fitness app is XX."

[1699] Step 17:

[1700] The server selects relevant advertisements based on the analysis results.

[1701] Step 18:

[1702] The server encodes the generated dialogue and advertisements as audio files.

[1703] Step 19:

[1704] The server delivers the encoded audio file to the device.

[1705] Step 20:

[1706] The device plays the received audio file, allowing the user to perform operations such as stop, skip, and play during playback.

[1707] Step 21:

[1708] Users listen to what is being played and find that the content is personalized, for example, responses or advertisements tailored based on the user's current emotional state.

[1709] The above processing steps not only provide a personalized radio experience based on the user's interests, but also utilize an emotion engine to provide content optimized for the user's emotional state. This system reduces visual load while providing useful and satisfying information.

[1710] Example 2

[1711] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1712] Conventional voice information delivery systems have limited functionality for providing personalized content based on a user's profile information and usage history, resulting in issues with not being able to fully respond to user interests. Furthermore, they have not been able to adjust conversation content using real-time emotion recognition or dynamically generate content that reflects user feedback. Furthermore, there have been cases where advertising content and audio content do not always match, hindering the user experience.

[1713] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1714] In this invention, the server includes a means for collecting and saving user profile information, a means for periodically transmitting usage history data, and a means for analyzing the usage history data and profile information. This enables personalized conversation content based on the user's interests and the generation of related advertisements and their delivery as audio files. Furthermore, the user experience can be improved by recognizing the user's emotional state and adjusting the conversation content in real time.

[1715] "User profile information" is information including personal attributes of the user, such as age, sex, occupation, and hobbies.

[1716] "Means of collection and storage" refers to the technical means for receiving data entered by the user and storing it in a database or the like.

[1717] "Usage history data" refers to data such as playback time, number of skips, and selected topics that are recorded when a user uses an application.

[1718] "Means of analysis" refers to technical means using AI algorithms and machine learning models to identify user interests and concerns based on collected data.

[1719] "Means for generating conversational content" means the technical means for creating natural language responses and conversations using a generative AI model based on the analysis results.

[1720] The "means for receiving letters and analyzing their contents" refers to the technical means for receiving questions and comments sent by users and processing them using a text analysis model.

[1721] The "means for generating relevant information and answers" refers to the technical means for generating appropriate information and answers based on the analyzed content of the letter.

[1722] "Means for selecting relevant advertisements and generating audio files of conversations and advertisements" refers to the technical means for selecting highly relevant advertisements based on the user's interests and generating audio data that integrates the advertisements into the conversation content.

[1723] The "means for playing audio files" refers to the technical means for appropriately playing audio files generated on a user terminal.

[1724] The "means for recognizing emotional states" refers to technical means for analyzing the user's voice and input data in real time and identifying the user's emotions.

[1725] A "conversational content adjustment mechanism" is a technological mechanism for changing the tone or content of a conversation in real time based on a perceived emotional state.

[1726] This invention is an AI radio system that combines an emotion engine that recognizes the user's emotions. The specific processing of the system program is shown below.

[1727] The present invention is a system for providing voice information based on user operations, and has the following main functions.

[1728] User Data Collection and Initial Settings

[1729] User:

[1730] First, the user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[1731] Device:

[1732] The device establishes an Internet connection and sends an HTTP request to the server to transmit the entered profile information to the server.

[1733] server:

[1734] The server receives the HTTP request and stores the profile information in a database, using a database management system such as MySQL.

[1735] Daily data collection

[1736] Device:

[1737] Every time the application is launched, it locally records user actions (playback time, skip count, selected topics, etc.) and saves the data in JSON format.

[1738] server:

[1739] The device sends usage history data to the server at regular intervals or when a specific event occurs. The server stores the received data in a database and processes it on a cloud server for data analysis.

[1740] Generate personalized conversations

[1741] server:

[1742] The server analyzes profile information and usage history data to select the most suitable topics for each user. It uses machine learning libraries such as Python's scikit-learn and TensorFlow to generate prompts for a natural language generation model (e.g., OpenAI GPT-3) and sends them to the model via an API.

[1743] Specific prompt examples:

[1744] "Users are interested in technology, so talk in detail about the latest smartwatch features."

[1745] The server uses the received generated text to generate an audio file using a speech synthesis model (e.g., Google Text-to-Speech) taking into account emotions and tone.

[1746] Emotion recognition and conversation adjustment with emotion engine

[1747] Device:

[1748] The device captures the user's voice input through a microphone and analyzes it in real time using an emotion engine (e.g., Affectiva Emotion AI).

[1749] server:

[1750] The system receives emotional data from the device and stores it in a database. Based on the analysis results, it adjusts the conversation content accordingly. For example, if the user is feeling stressed, it will regenerate the conversation content to provide a calming tone and relaxing topics.

[1751] Processing letters and providing information

[1752] User:

[1753] Users use the letter feature within the app to enter questions or comments and press the send button.

[1754] Device:

[1755] The terminal converts the input content into JSON format and sends it to the server as an HTTP POST request.

[1756] server:

[1757] The server analyzes the received letter using a text analysis model (e.g., spaCy), generates related information and recommendations using a generative AI model (GPT-3), generates the generated answers as audio files, and selects and integrates advertisements into the audio data.

[1758] Delivering personalized ads

[1759] server:

[1760] The server selects relevant ads based on the user's profile information and current conversation content, and encodes the generated conversation content and ads as audio files.

[1761] Specific examples of behavior:

[1762] For technology-loving users, ads about the latest smartwatch features are seamlessly integrated into the conversation.

[1763] Content Delivery and Playback

[1764] server:

[1765] The server sends the generated audio file as an HTTP response to deliver it to the user terminal using an HTTP framework (e.g., Flask).

[1766] Device:

[1767] The device uses the audio playback feature (e.g., AndroidMediaPlayer) to play the audio file, and the user can perform operations such as play, stop, and skip through the app interface.

[1768] User:

[1769] The user listens to the generated audio content and submits feedback or follow-up questions as needed.

[1770] In this way, the system of the present invention delivers personalized audio content based on the user's interests and generates finely tuned responses based on real-time emotion recognition, resulting in a more satisfying user experience.

[1771] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1772] Step 1: Collecting user data and initial setup

[1773] User

[1774] Input: The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[1775] Specific behavior: A user enters profile information into an application's input form and presses the submit button.

[1776] Terminal

[1777] Data processing: The entered profile information is temporarily saved locally.

[1778] Output: This information is converted to JSON format and sent to the server as an HTTP request.

[1779] What happens: The device establishes an Internet connection and generates an HTTP POST request containing the entered profile information.

[1780] server

[1781] Input: The server receives the profile information sent from the device.

[1782] Data processing: Analyze the received information and store it in a database.

[1783] Output: Generates a save confirmation message.

[1784] Specific operation: The server stores the received profile information using a database management system such as MySQL.

[1785] Step 2: Collecting data on a daily basis

[1786] Terminal

[1787] Input: Operational data when a user uses the application (playback time, number of skips, selected topics, etc.).

[1788] Data processing: Save usage history data in JSON format to local storage.

[1789] Output: Set a trigger to periodically send the saved data to the server.

[1790] Specific operation: Detects user operations within the application and stores data locally.

[1791] server

[1792] Input: Usage history data sent periodically from the device.

[1793] Data processing: The received usage history data is stored in a database and sent to an analysis server.

[1794] Output: Generates data analysis results.

[1795] Specific operation: The server analyzes the usage history data stored in the database using an AI algorithm to identify the user's interests and concerns.

[1796] Step 3: Generate personalized conversations

[1797] server

[1798] Input: Analysis of user profile information and usage history data.

[1799] Data processing: Based on the analysis results, prompt sentences are generated for the generative AI model (e.g., GPT-3).

[1800] Output: The generated conversation.

[1801] Specific operation: The server analyzes the data using Python's scikit-learn and TensorFlow and generates a prompt. The generated prompt is sent to the generative AI model via an API and an answer is received. "The user is interested in technology, so please talk in detail about the latest smartwatch features."

[1802] server

[1803] Input: The speech received from the generative AI model.

[1804] Data processing: Applying speech synthesis models to add emotion and tone to the conversation.

[1805] Output: An audio file of the generated conversation.

[1806] What it does: The server uses a speech synthesis model such as Google Text-to-Speech to convert the generated text into an audio file.

[1807] Step 4: Emotion recognition and conversation adjustment using the emotion engine

[1808] Terminal

[1809] Input: User speech and input data.

[1810] Data processing: Analyze emotional states in real time using an emotion engine.

[1811] Output: Emotion data as the analysis result.

[1812] Specific operation: The device captures the user's voice using a microphone device and analyzes it in real time using an emotion engine (e.g., Affectiva Emotion AI).

[1813] server

[1814] Input: Emotion data sent from the device.

[1815] Data processing: Adjusting conversation content in real time based on emotional data.

[1816] Output: Improved conversation.

[1817] Specific operation: Based on the analysis results, the server re-prompts the generative AI model as necessary and corrects the conversation content.

[1818] Step 5: Processing your letter and providing information

[1819] User

[1820] Input: User comments and questions.

[1821] What happens: A user uses the in-app message feature, enters a question or comment, and presses the send button.

[1822] Terminal

[1823] Input: The letter data entered by the user.

[1824] Data processing: Generate an HTTP POST request in JSON format and send it to the server.

[1825] Output: Send to server.

[1826] Specific operation: The device converts user input into JSON format and sends it to the server via the Internet.

[1827] server

[1828] Input: Letter data sent from the terminal.

[1829] Data processing: Analyze using text analytics models (e.g., spaCy) and generate relevant answers and information using generative AI models.

[1830] Output: The generated answer and related information.

[1831] Specific operation: The server receives the letter data, analyzes it using a natural language processing model, and creates an appropriate response using a generative AI model, which then generates it as an audio file.

[1832] Step 6: Delivering personalized ads

[1833] server

[1834] Input: User profile information and current conversation.

[1835] Data processing: Selecting relevant ads and merging the conversation and ads into an audio file.

[1836] Output: An audio file containing ads.

[1837] How it works: Based on the user information and conversation content, the server selects relevant advertisements and encodes them into audio files.

[1838] Specific examples

[1839] For technology-loving users, ads for the latest smartwatches are selected and seamlessly integrated into conversations.

[1840] Step 7: Content Delivery and Playback

[1841] server

[1842] Input: The generated audio file.

[1843] Data processing: Audio files are delivered to the user's device using the HTTP framework.

[1844] Output: Audio file delivery to user device.

[1845] Specific operation: The server uses an HTTP framework such as Flask to send the audio file to the terminal as an HTTP response.

[1846] Terminal

[1847] Input: The audio file sent from the server.

[1848] Data processing: Play audio files using the audio playback function.

[1849] Output: Playback of audio content.

[1850] Specific behavior: The device plays the audio file using AndroidMediaPlayer or similar, and allows the user to control playback through the interface.

[1851] User

[1852] What it does: The user listens to the generated audio content and provides feedback or asks follow-up questions if necessary. This data is used to generate future content.

[1853] (Application example 2)

[1854] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1855] Conventional radio systems and content distribution services have difficulty providing personalized content based on a user's current emotional state. Furthermore, they are unable to respond to a user's emotional changes in real time, resulting in a poor user experience. Furthermore, in the case of advertising delivery, insufficient personalization based on the user's current situation results in a failure to attract the user's attention and a reduced effectiveness.

[1856] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1857] In this invention, the server includes means for analyzing the user's voice and input data in real time and incorporating an emotion engine that recognizes the user's emotional state, means for adjusting topics and conversation content in real time based on the emotion data, and means for receiving letters from users, analyzing the content, and generating related information and answers, thereby making it possible to provide appropriate content in real time that matches the user's emotional state.

[1858] "User profile information" is basic personal information such as age, sex, occupation, and hobbies that a user enters when using the service for the first time.

[1859] "Usage history data" refers to behavioral data such as playback time, number of skips, and selected topics when a user uses an application.

[1860] The "emotion engine" is an engine that has the function of analyzing the user's voice and input data in real time and recognizing their emotional state.

[1861] The "server" is a computer system that stores and analyzes data collected from users and generates and adjusts topics and conversation content based on their emotional state and interests.

[1862] A "speech synthesis model" is a technology for generating speech based on computer-generated text and reproducing emotion and tone.

[1863] "Emotion data" is data that indicates the user's current emotional state as analyzed by the emotion engine.

[1864] "Advertisements" are commercial information that is selected based on the user's interests and incorporated naturally into conversations.

[1865] "Letters" are messages such as questions, comments, and feedback that users send through the application.

[1866] A "topic" is a conversation subject selected by the server based on the user's interests and concerns.

[1867] The "audio file" is an audio data file in which the generated conversation content or advertisement content is encoded, and is played back on the user terminal.

[1868] In this invention, a system for exchanging data between a user terminal and a server to provide personalized content to the user will be specifically described. The detailed program processing for realizing this system will be described below.

[1869] User Data Collection and Initial Settings

[1870] After installing the application, the user enters their profile information the first time they start it. This information includes, for example, age, gender, occupation, and hobbies. This information is sent from the user's device to the server, which then stores it in a database. The hardware used at this stage is a smartphone, and the software includes front-end technology (e.g., React Native) that configures the user interface (UI) and a library (e.g., requests) that sends HTTP requests.

[1871] Daily data collection

[1872] When a user uses an application, usage history data such as playback time, number of skips, and selected topics is recorded. This data is periodically sent to a server, which stores the received data in a database. The server uses an AI algorithm to analyze this data and identify the user's interests. The software used for this analysis is a database management system (e.g., MySQL) and an AI analysis algorithm (e.g., TensorFlow).

[1873] Emotion recognition and conversation adjustment with emotion engine

[1874] The user device analyzes the user's voice and input data in real time and recognizes their emotional state using an emotion engine. The hardware used to collect voice data is the smartphone's microphone, and the software is a voice analysis library (e.g., Google Cloud Speech-to-Text API). The server analyzes the user's current emotional state based on the emotion data sent from the emotion engine and adjusts the topic and content of the conversation in real time. For example, if the user is feeling stressed, it generates a relaxing tone of voice.

[1875] Generate personalized conversations

[1876] The server selects the most suitable topic for the user based on the analysis results and generates the conversation using a natural language generation model, while also applying an appropriate speech synthesis model to add emotion and tone. For example, if the user prefers topics related to fitness, the server will generate a conversation about "recommended fitness apps."

[1877] Processing letters and providing information

[1878] Users can use the application's message feature to send questions or comments. For example, they can ask, "What fitness apps do you recommend?" These messages are sent from the user's device to a server, which analyzes the content and generates relevant information and answers. This analysis and generation is performed using a natural language processing model (e.g., GPT-3).

[1879] Delivering personalized ads

[1880] The server selects relevant ads based on the user's profile information and current conversation, and encodes the generated conversation and ads into audio files. For example, a technology-loving user might receive an ad for the latest wearable devices. The selected ads are seamlessly integrated into the conversation.

[1881] Content Delivery and Playback

[1882] The generated audio file is delivered from the server to the user's device, where it is played. When the user listens to AI Radio, they can control playback (pause, skip, play, etc.) as needed. The hardware used in this part is a smartphone, and the software is a media player library with audio playback capabilities.

[1883] Prompt Sentence Examples

[1884] For example, here is a prompt that describes a situation where the user is feeling stressed about fitness:

[1885] Example prompt sentence:

[1886] "A user who has asked a fitness question is stressed. We recommend fitness apps to that user in a relaxing tone."

[1887] In this way, it becomes possible to provide personalized content in real time according to the user's emotional state and interests. This invention not only improves the user experience, but also realizes highly personalized content delivery to increase advertising effectiveness.

[1888] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1889] Step 1:

[1890] The user installs the application and enters profile information when the application is first launched.

[1891] Examples of input: name, age, gender, occupation, hobbies, etc.

[1892] The terminal transmits the input data to the server, which stores the received data in a database.

[1893] Output: User profile information is stored in the server database.

[1894] Step 2:

[1895] When a user uses the application, usage history data such as playback time, number of skips, and selected topics is recorded.

[1896] Example inputs: playback data, skip counts, topic selection.

[1897] The terminal periodically transmits the collected data to the server, and the server stores the received data in a database.

[1898] Output: Usage history data is stored in the server database.

[1899] Step 3:

[1900] The server uses AI algorithms to analyze the collected profile information and usage history data to identify the user's interests.

[1901] Examples of input: profile information, usage history data.

[1902] The server uses AI algorithms (e.g., TensorFlow) to analyze the data.

[1903] Output: Analysis results based on user interests and concerns.

[1904] Step 4:

[1905] The server uses an emotion engine that analyzes the user's voice and input data in real time to recognize their emotional state.

[1906] Examples of input: user voice data, text input data.

[1907] The device sends the audio to an emotion engine (e.g., Google Cloud Speech-to-Text API) for emotion analysis.

[1908] Output: Emotion data indicating the user's current emotional state.

[1909] Step 5:

[1910] The server adjusts topics and conversation content in real time based on emotional data.

[1911] Examples of input: emotional data, historical usage data.

[1912] The server generates the conversation using a natural language generation model (e.g., GPT-3) and makes any necessary adjustments.

[1913] Output: Conversational content that reflects the user's emotional state.

[1914] Step 6:

[1915] Users can submit questions or comments using the application's letter feature.

[1916] Example input: A user question or comment.

[1917] The terminal sends input from the user to the server, which analyzes the content.

[1918] Output: Analysis results and related information and answers.

[1919] Step 7:

[1920] The server selects relevant advertisements based on the user's profile information and current conversation content.

[1921] Examples of input: profile information, conversation content.

[1922] The server uses AI algorithms to select the most relevant ads for the user and incorporate them into the conversation.

[1923] Output: Advertising information embedded in the conversation.

[1924] Step 8:

[1925] The server encodes the generated conversations and advertisements as audio files and delivers them to the user terminal.

[1926] Examples of input: conversation content, advertising information.

[1927] The server uses a speech synthesis model (e.g., a speech synthesis engine) to generate an audio file and send it to the device.

[1928] Output: The audio file delivered to the user's device.

[1929] Step 9:

[1930] The user terminal plays the audio file delivered from the server and allows the user to operate the file (stop, skip, play, etc.) during playback.

[1931] Example input: An audio file delivered from a server.

[1932] The device plays audio files using a built-in media player and accepts user interaction through the UI.

[1933] Output: Audio playback with user interaction.

[1934] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1935] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1936] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1937] [Fourth embodiment]

[1938] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1939] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1940] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1941] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1942] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1943] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1944] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1945] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1946] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1947] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1948] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1949] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1950] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1951] This invention is an AI radio system that develops conversations based on the user's interests and concerns. The processing of the system's program is explained in natural language with specific examples.

[1952] User Data Collection and Initial Settings

[1953] User

[1954] First, the user installs the application and enters profile information when the application is first launched.

[1955] Terminal

[1956] The entered profile information is transmitted from the user terminal to the server.

[1957] server

[1958] The server stores the received profile information in a database for later analysis.

[1959] Daily data collection

[1960] Terminal

[1961] Record usage history data (playback time, number of skips, selected topics, etc.) recorded when a user uses the application.

[1962] server

[1963] The server receives usage history data periodically sent from the device and stores it in a database. The server then analyzes this data using AI algorithms to identify the user's interests.

[1964] Generate personalized conversations

[1965] server

[1966] The server then selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in technology, the server might select a topic such as "latest smartwatch features." Based on the selected topic, the server generates the conversation using a natural language generation model. At this time, an appropriate speech synthesis model is also applied to add emotion and tone.

[1967] Processing letters and providing information

[1968] User

[1969] Users can submit questions or comments using the application's letter feature.

[1970] Terminal

[1971] A letter is sent from the user terminal to the server.

[1972] server

[1973] The server analyzes the received letter and generates related information and answers based on its content. For example, if a user asks for a "recommended fitness app," the server generates an answer such as "The recommended fitness app is XX." Related advertisements are also selected.

[1974] Delivering personalized ads

[1975] server

[1976] The server selects relevant advertisements based on the user's profile information and current conversation content, and encodes the generated conversation and advertisements as audio files.

[1977] Examples:

[1978] For technology-loving users, the server will select advertisements for the latest smartwatches and seamlessly integrate them into the conversation.

[1979] Content Delivery and Playback

[1980] server

[1981] The generated audio file is delivered to the user terminal.

[1982] Terminal

[1983] The user terminal plays the audio file, allowing the user to listen to it, and allows the user to control the playback (stop, skip, play, etc.).

[1984] User

[1985] Users listen to the AI ​​radio and provide playback control and feedback as needed.

[1986] Explanation based on concrete examples

[1987] For example, suppose a user sends a letter asking, "What fitness app do you recommend?" In this case, the server analyzes the letter, replies, "The fitness app I recommend is ____," and generates a related advertisement (e.g., a sale on a fitness app). The generated audio file is then delivered to the user's device, and when the user listens to it, the answer to the question and the advertisement information are simultaneously provided.

[1988] The system enables a personalized radio experience based on the user's interests, provides useful information while reducing visual strain, and improves advertising effectiveness through personalized advertising.

[1989] The processing flow will be explained below.

[1990] Step 1:

[1991] The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[1992] Step 2:

[1993] The device sends the entered profile information to the server.

[1994] Step 3:

[1995] The server stores the received profile information in a database.

[1996] Step 4:

[1997] The device records data about the user's use of the application (playback time, number of skips, selected topics, etc.).

[1998] Step 5:

[1999] The terminal periodically transmits the recorded usage history data to the server.

[2000] Step 6:

[2001] The server receives the usage history data and stores it in a database.

[2002] Step 7:

[2003] The server uses AI algorithms to analyze users' interests based on their profile information and usage history data.

[2004] Step 8:

[2005] The server selects topics appropriate for the user based on the analysis results.

[2006] Step 9:

[2007] The server uses a natural language generation model to generate conversational content based on the selected topic.

[2008] Step 10:

[2009] The server applies a speech synthesis model to add emotion and tone to the generated dialogue.

[2010] Step 11:

[2011] A user submits a question or comment using the in-app letter feature.

[2012] Step 12:

[2013] The terminal sends the letter from the user to the server.

[2014] Step 13:

[2015] The server receives the letter and analyzes its contents.

[2016] Step 14:

[2017] The server generates relevant information and answers based on the analysis results.

[2018] Step 15:

[2019] The server selects relevant advertisements based on the analysis results.

[2020] Step 16:

[2021] The server encodes the generated dialogue and advertisements as audio files.

[2022] Step 17:

[2023] The server delivers the encoded audio file to the user terminal.

[2024] Step 18:

[2025] The device plays the received audio file.

[2026] Step 19:

[2027] Allows the user to perform operations such as stop, skip, and play during playback.

[2028] Example 1

[2029] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2030] Conventional radio and audio content have faced challenges in providing content based on user interests, resulting in a lack of personalization for individual users. Furthermore, there were limited ways to respond appropriately to user feedback and questions, resulting in a one-way user experience. Furthermore, there was no established method for maximizing the effectiveness of advertising.

[2031] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[2032] In this invention, the server includes means for collecting and saving user profile information, means for periodically transmitting user usage history data to the server, means for analyzing user interests and concerns based on the transmitted usage history data and profile information, means for selecting a topic based on the analysis result and generating conversation content using a generative AI model, means for converting the conversation content into an audio file using an appropriate voice synthesis model, means for receiving letters from users and analyzing the content to generate related information and answers, means for selecting relevant advertisements based on the user's interests and concerns and generating the conversation and advertisements as audio files, and means for delivering the generated audio files to the user terminal. This enables the provision of personalized content based on the user's interests and concerns, improving the user experience and maximizing advertising effectiveness.

[2033] "User Profile Information" means personal and preference information provided by a user, including name, age, hobbies, etc.

[2034] "Usage history data" is data generated when a user uses an application, including play time, skip counts, selected topics, and the like.

[2035] "Server" means a computer system that collects, stores, and analyzes data sent by users and provides necessary information.

[2036] A "generative AI model" is an artificial intelligence model for generating natural language text based on input prompts.

[2037] A "speech synthesis model" is a system that includes algorithms and techniques for converting text data into speech data.

[2038] A "letter" is a text message such as a question or comment that a user sends through the application.

[2039] "Advertisements" are commercial information selected based on the user's interests and incorporated into content.

[2040] An "audio file" is a file in which audio data is stored in a digital format and is provided in a format that can be played on a user terminal.

[2041] "User terminal" means an electronic device used by a user to access the system, including a smartphone, tablet, or PC.

[2042] "Playback" refers to the act of the user terminal outputting an audio file as sound and the user listening to it.

[2043] The present invention relates to a system that provides personalized audio content based on a user's interests. This system generates and distributes content using a generative AI model and a voice synthesis model based on the user's profile information and usage history data. Specific embodiments of the system are described below.

[2044] User Data Collection and Initial Settings

[2045] User

[2046] First, a user installs the application and enters profile information such as name, age, hobbies, etc. This step lays the foundation for the system to provide a personalized experience tailored to the user.

[2047] Terminal

[2048] The device encrypts the profile information entered by the user and sends it to the server using the HTTPS protocol, ensuring the security of the user information during this process.

[2049] server

[2050] The server stores the received profile information in a database, which allows the integration and management of user-specific data for future analysis.

[2051] Specific examples

[2052] For example, if a user registers as "Yamada Taro" and selects "technology" as a hobby, this information is sent to the server and stored in a database.

[2053] ---

[2054] Daily data collection

[2055] Terminal

[2056] When a user uses an application, the device collects usage history data in real time, such as playback time, number of skips, and selected topics, allowing the application to always reflect the latest user behavior patterns.

[2057] server

[2058] The device sends the collected data to the server in batch processing at regular intervals (e.g., every 5 minutes). The server stores the received data in a database and analyzes it using AI algorithms. This allows the user's interests and concerns to be identified.

[2059] Specific examples

[2060] If a user frequently watches technology-related episodes and skips fitness-related episodes, the server may determine that the user has a strong interest in technology.

[2061] Natural Language Prompt Examples

[2062] "Data shows that users frequently watch technology-related episodes and skip fitness-related episodes."

[2063] ---

[2064] Generate personalized conversations

[2065] server

[2066] The server selects the most appropriate topic for the user based on the analyzed results, generates conversation content using a natural language generation model (e.g., OpenAI's GPT-4), and converts the generated text into speech using a speech synthesis model (e.g., Amazon Polly).

[2067] Specific examples

[2068] If the server determines from the analysis that the user is interested in technology, it selects the topic "Let's talk about the latest smartwatches." It inputs "Tell me about the latest smartwatches" as a prompt to GPT-4, and passes the generated text to Amazon Polly to generate a natural-sounding voice file.

[2069] Natural Language Prompt Examples

[2070] "Tell me about the latest smartwatch."

[2071] ---

[2072] Processing letters and providing information

[2073] User

[2074] Users can use the application's in-app message feature to send questions or comments, such as "What fitness apps do you recommend?"

[2075] Terminal

[2076] The device sends the message in real time to a server, and this information is also encrypted to ensure security.

[2077] server

[2078] The server analyzes the received letter and generates relevant information and answers based on its content. In addition, relevant advertisements are also selected and generated at the same time.

[2079] Specific examples

[2080] When a user sends a letter saying, "Please tell me your recommended fitness app," the server responds, "The recommended fitness app is XX," and also generates advertisements with information about special sales of fitness apps.

[2081] Natural Language Prompt Examples

[2082] What fitness apps do you recommend?

[2083] ---

[2084] Delivering personalized ads

[2085] server

[2086] The server selects relevant advertisements based on the user's profile information and current conversation content and encodes them as audio files.

[2087] Specific examples

[2088] If the server determines that the user is interested in technology, it selects an advertisement for the latest smartwatch, inputs a prompt into the GTP-4 saying, "Generate ad copy for the latest smartwatch," and uses Amazon Polly to convert the generated ad copy into an audio file that is seamlessly integrated into the conversation.

[2089] Natural Language Prompt Examples

[2090] "Generate ad copy for the latest smartwatch."

[2091] ---

[2092] Content Delivery and Playback

[2093] server

[2094] The generated audio file is then delivered to the user's device, taking into account the user's network conditions.

[2095] Terminal

[2096] The device plays the delivered audio file, and the user can perform operations during playback (such as stop, skip, and play).

[2097] User

[2098] Users listen to AI Radio and provide playback control and feedback as needed, allowing the system to continue to learn about the user's emerging interests and preferences.

[2099] Specific examples

[2100] The server delivers the generated audio file to the user's device, which then decodes and plays it. The user can press the "Stop" button or the "Skip" button during playback to move on to the next content.

[2101] ---

[2102] The implementation of this system will enable a personalized radio experience based on the user's interests. By appropriately utilizing generative AI models and prompts, more natural and engaging content will be provided to the user. This is expected to improve the user experience and maximize advertising effectiveness.

[2103] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2104] System program processing flow

[2105] Step 1: Collecting user data and initial setup

[2106] User

[2107] When a user installs the application and launches it for the first time, they enter their profile information, such as their name, age, and hobbies. By entering this information, the system is ready to provide a personalized experience for the user.

[2108] Terminal

[2109] The device encrypts the entered profile information and sends it to the server using the HTTPS protocol, ensuring the security of the user information during this process.

[2110] server

[2111] The server stores the received profile information in a database for later analysis.

[2112] Specific actions

[2113] For example, if a user registers as "Taro Yamada" and selects "technology" as a hobby, this information is encrypted and sent from the device to the server, which then stores the received data in the appropriate tables in the database.

[2114] Input and Output

[2115] Input: Profile information such as name, age, hobbies, etc.

[2116] Output: User profile information stored in a database

[2117] ---

[2118] Step 2: Collecting data on a daily basis

[2119] Terminal

[2120] When a user uses an application, the device collects usage history data such as playback time, number of skips, and selected topics, ensuring that the device always reflects the user's behavioral patterns in an up-to-date manner.

[2121] server

[2122] The device sends the collected usage history data at regular intervals (e.g., every 5 minutes) to the server, which receives it and stores it in a database. The server then uses AI algorithms to analyze the user's interests and concerns.

[2123] Specific actions

[2124] Every time a user listens to a radio episode, the device records data such as the start time, end time, number of skips, and selected topics. Every five minutes, the device sends this data to a server, which stores it in a database. An AI algorithm is used to analyze which topics the user is interested in.

[2125] Input and Output

[2126] Input: Usage history data such as play time, skip counts, and selected topics

[2127] Output: Analyzed user interests

[2128] ---

[2129] Step 3: Generate personalized conversations

[2130] server

[2131] The server selects the most appropriate topics for the user based on the analyzed user's interests, then generates the conversation using a natural language generation model (e.g., OpenAI's GPT-4), and converts the generated text into an audio file using an appropriate speech synthesis model (e.g., Amazon Polly).

[2132] Specific actions

[2133] For example, if the server determines that the user is interested in technology based on the analysis results, it inputs the prompt sentence "Tell me about the latest smartwatch" into GPT-4. The generated text is passed to Amazon Polly to generate an audio file.

[2134] Input and Output

[2135] Input: Topic analysis results based on user interests, prompt text

[2136] Output: Audio file (conversation content)

[2137] ---

[2138] Step 4: Processing your letter and providing information

[2139] User

[2140] Users can use the application's in-app message feature to send questions or comments, such as "What fitness apps do you recommend?"

[2141] Terminal

[2142] The device sends the message in real time to a server, and this information is also encrypted to ensure security.

[2143] server

[2144] The server analyzes the received letter and generates relevant information and answers based on its content. It also simultaneously selects and generates relevant advertisements.

[2145] Specific actions

[2146] For example, if a user sends a message asking, "What fitness apps do you recommend?", the device encrypts the question and sends it to the server, which uses GPT-4 to generate an answer and also selects and generates relevant advertisements.

[2147] Input and Output

[2148] Input: Letter from user

[2149] Output: Answers and related ads

[2150] ---

[2151] Step 5: Delivering personalized ads

[2152] server

[2153] The server selects relevant advertisements based on the user's profile information and current conversation content and encodes them as audio files.

[2154] Specific actions

[2155] For example, if the server determines that the user is interested in technology, it will generate an ad copy by inputting a prompt such as "Generate an ad copy for the latest smartwatch" into GPT-4. The generated ad copy is then converted into speech using Amazon Polly and incorporated into the conversation.

[2156] Input and Output

[2157] Input: User profile information, current conversation

[2158] Output: Audio file (conversation and advertisements)

[2159] ---

[2160] Step 6: Content Delivery and Playback

[2161] server

[2162] The server then delivers the generated audio files to the user's terminal, taking into account the user's network conditions.

[2163] Terminal

[2164] The device decodes and plays the delivered audio file, and the user can perform operations during playback (such as stop, skip, and play).

[2165] User

[2166] Users listen to AI Radio and provide playback control and feedback as needed, allowing the system to continue to learn about the user's emerging interests and preferences.

[2167] Specific actions

[2168] The server delivers the generated audio file to the user's device, which decodes and plays it. The user can press the "Stop" button during playback or the "Skip" button to move on to the next content.

[2169] Input and Output

[2170] Input: Generated audio file

[2171] Output: Played audio content

[2172] ---

[2173] This allows users to enjoy a more fulfilling radio experience through personalized content, while also maximizing advertising effectiveness and improving the user experience.

[2174] (Application example 1)

[2175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2176] Conventional voice guidance systems often provide only uniform content, without providing sufficient personalized information based on the user's interests. They also lacked the ability to deliver advertisements based on the user's interests, making efficient marketing difficult. Furthermore, the lack of effective means for analyzing usage history and profile information limited the user experience.

[2177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2178] In this invention, the server includes means for inputting user profile information, means for collecting and saving the input profile information, means for recording user usage history data, means for periodically transmitting the recorded usage history data to the server, means for analyzing user interests and concerns based on the transmitted usage history data and profile information, means for selecting topics and generating conversation content based on the analysis results, means for receiving letters from users and generating related information and responses after analyzing the content, means for selecting relevant advertisements based on the user's interests and concerns and generating the conversations and advertisements as audio files, means for delivering the generated audio files to user terminals, means for the user terminals to play the audio files, and means for generating and providing personalized audio guides for tourist spots and museums, thereby enabling improved user experience and efficient marketing.

[2179] "User Profile Information" means basic personal information and interest data provided by a User.

[2180] "Usage history data" is information related to behavior, such as operation history and playback history when a user uses an application.

[2181] A "server" is a computer system that collects, stores, and analyzes data sent by users.

[2182] "Analysis" refers to the processing of data to identify user interests based on collected profile information and usage history data.

[2183] A "topic" is a specific subject or theme that may be of interest to a user.

[2184] The "conversation content" is a voice message for dialogue with the user that is generated based on the analyzed data.

[2185] A "letter" is a message containing a question or comment from a user.

[2186] "Related information" is answers and information provided based on the user's letter.

[2187] "Advertisement" means commercial information provided based on the user's interests.

[2188] "Audio files" are generated conversations and advertisements encoded as audio data.

[2189] "Personalized audio guides at tourist spots and museums" are audio guides for tourist spots and museums that are customized according to the user's interests and concerns.

[2190] This invention is an AI radio system that develops conversations based on the user's interests and focuses on providing personalized audio guides, particularly at tourist spots and museums. The system is configured as follows:

[2191] User Data Collection and Initial Settings

[2192] The user first installs the application on their smartphone and enters their profile information when they first launch it. This profile information includes the user's name and genres of interest (e.g., history, art, technology, etc.). The entered profile information is sent from the user's device to the server, where it is stored in a database.

[2193] Daily data collection

[2194] When a user uses the application, usage history data (such as play time, number of skips, and selected topics) is recorded. This data is periodically sent from the device to a server, where it is stored and analyzed. The server then uses AI algorithms (e.g., machine learning models) to analyze this data and identify the user's interests.

[2195] Generate personalized conversations and audio guides

[2196] The server selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in history, the server selects a topic such as "historical exhibits in a specific museum." It then uses a natural language generation model (e.g., GPT-3) to generate audio guide content that includes emotion and tone. The generated content is then encoded as an audio file using a speech synthesis model (e.g., a text-to-speech engine).

[2197] Processing letters and providing information

[2198] Users can send questions and comments using the application's message function. For example, if a user asks, "What exhibits do you recommend at this museum?", the server receives the message, analyzes the content, and generates the most appropriate information and answer. At the same time, relevant advertisements are selected and generated as audio files.

[2199] Delivering personalized ads

[2200] The server selects relevant advertisements based on the user's profile information and current conversation content, and the generated conversation and advertisements are encoded as audio files and delivered to the user's device.

[2201] Content Delivery and Playback

[2202] The generated audio file is delivered from the server to the user's terminal, where the user can listen to it. The user can also perform operations (such as stopping, skipping, and playing) during playback.

[2203] Explanation based on concrete examples

[2204] For example, if a user is interested in "Japanese history," the app can provide an audio guide with the latest information on historical museum exhibits and related commentary. It can also provide personalized advertisements for books related to Japanese history.

[2205] An example prompt is:

[2206] "Prompt: User is interested in "Japanese history." Generate a detailed audio guide for a current museum exhibit. Also create related ads."

[2207] The present invention enables improved user experience and efficient marketing, and provides more comprehensive information at tourist spots and museums.

[2208] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2209] Step 1:

[2210] The user installs the application on their smartphone and enters their profile information (such as their name and genres of interest) when they first start it up. The entered profile information is sent from the user's device to the server.

[2211] Input: User profile information

[2212] What happens: A user enters information into an application and presses the "Submit" button.

[2213] Output: Profile information sent to the server

[2214] Step 2:

[2215] The server stores the received profile information in a database, which is used for analytics purposes.

[2216] Input: Profile information sent to the server

[2217] What it does: The server executes an SQL query that stores information in a database.

[2218] Output: Profile information stored in a database

[2219] Step 3:

[2220] When a user uses the application, usage history data (playback time, number of skips, selected topics, etc.) is recorded and periodically sent to the server.

[2221] Input: Application usage history data

[2222] How it works: The application records usage history in a log file or database and periodically sends it to the server.

[2223] Output: Usage history data sent to the server

[2224] Step 4:

[2225] The server uses AI algorithms (e.g., machine learning models) to analyze the user's interests and concerns based on usage history data and profile information.

[2226] Input: Submitted usage history data and profile information

[2227] How it works: The server inputs data into a machine learning model to estimate the user's interests.

[2228] Output: Analyzed user interests

[2229] Step 5:

[2230] Based on the analysis results, the server selects the most appropriate topic for the user and generates conversation content using a natural language generation model (e.g., GPT-3).

[2231] Input: Analyzed user interests

[2232] How it works: The server inputs topics into a natural language generation model and generates conversation content.

[2233] Output: Generated conversation

[2234] Step 6:

[2235] The server generates an audio file using a speech synthesis model (e.g., a text-to-speech engine) to add emotion and tone to the generated dialogue.

[2236] Input: Generated conversation

[2237] How it works: The server inputs text data into a speech synthesis model to generate speech data.

[2238] Output: Generated audio file

[2239] Step 7:

[2240] Users can submit questions and comments using the application's in-app message feature, such as "What exhibits would you recommend at this museum?"

[2241] Input: Letter from user

[2242] What happens: A user enters a question or comment into an application's input field and clicks the submit button.

[2243] Output: Letter sent to server

[2244] Step 8:

[2245] The server analyzes the received letter, generates relevant information and answers, and selects relevant advertisements if necessary.

[2246] Input: Letter sent to the server

[2247] How it works: The server feeds the letter data into a machine learning model to generate the best answers and relevant ads.

[2248] Output: Generated answers and related ads

[2249] Step 9:

[2250] The server encodes the generated conversations, responses, and advertisements as audio files and delivers them to the user terminal.

[2251] Input: Generated conversations, responses, and advertisements

[2252] Operation: The server encodes the voice data and sends it to the user's terminal.

[2253] Output: Audio file sent to the user's device

[2254] Step 10:

[2255] The user terminal plays the delivered audio file, and the user listens to it. The user can perform operations (stop, skip, play, etc.) during playback.

[2256] Input: Streamed audio file

[2257] Action: The media player on the user's device plays an audio file, and the user operates the playback controls.

[2258] Output: Played audio file and user operation log

[2259] Through the above processing steps, users can always obtain information based on their interests and concerns, allowing them to enjoy a more fulfilling experience at tourist spots and museums.

[2260] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2261] This invention is an AI radio system that combines an emotion engine that recognizes the user's emotions. Below, we will explain the system's program processing in natural language with specific examples.

[2262] User Data Collection and Initial Settings

[2263] User

[2264] First, the user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching the application for the first time.

[2265] Terminal

[2266] The entered profile information is transmitted from the user terminal to the server.

[2267] server

[2268] The server stores the received profile information in a database for later analysis.

[2269] Daily data collection

[2270] Terminal

[2271] Record usage history data (playback time, number of skips, selected topics, etc.) recorded when a user uses the application.

[2272] server

[2273] The server receives usage history data periodically sent from the device and stores it in a database. The server then analyzes this data using AI algorithms to identify the user's interests.

[2274] Generate personalized conversations

[2275] server

[2276] The server selects the most appropriate topic for the user based on the analysis results. For example, if the user is interested in technology, the server will select a topic such as "latest smartwatch features." Based on the selected topic, a natural language generation model is used to generate the conversation. At this time, an appropriate speech synthesis model is also applied to add emotion and tone.

[2277] Emotion recognition and conversation adjustment with emotion engine

[2278] Terminal

[2279] It analyzes the user's voice and input data in real time and recognizes their emotional state through an emotion engine.

[2280] server

[2281] The emotion engine receives emotional data analyzed from the device. The server adjusts the topic and conversation content in real time based on this emotional data. For example, if the user is feeling stressed, the server generates conversation content and a tone that will relax the user.

[2282] Processing letters and providing information

[2283] User

[2284] Users can submit questions or comments using the application's letter feature.

[2285] Terminal

[2286] A letter is sent from the user terminal to the server.

[2287] server

[2288] The server analyzes the received letter and generates related information and answers based on its content. For example, if a user asks for a "recommended fitness app," the server generates an answer such as "The recommended fitness app is XX." Related advertisements are also selected.

[2289] Delivering personalized ads

[2290] server

[2291] The server selects relevant advertisements based on the user's profile information and current conversation content, and encodes the generated conversation and advertisements as audio files.

[2292] Examples:

[2293] For technology-loving users, the server will select advertisements for the latest smartwatches and seamlessly integrate them into the conversation.

[2294] Content Delivery and Playback

[2295] server

[2296] The generated audio file is delivered to the user terminal.

[2297] Terminal

[2298] The user terminal plays the audio file and allows the user to perform operations (stop, skip, play, etc.).

[2299] User

[2300] Users listen to the AI ​​radio and provide playback control and feedback as needed.

[2301] Explanation based on concrete examples

[2302] For example, suppose a user sends a letter asking, "What fitness app do you recommend?" In this case, the server analyzes the letter, replies, "The recommended fitness app is ____," and generates a related advertisement (e.g., a sale on a fitness app). The generated audio file is delivered to the user's device, and when the user listens to it, the answer to the question and the advertisement information are simultaneously provided. In addition, by using an emotion engine, the server can provide a response with a tone and content that matches the emotional state the user was feeling when asking the question.

[2303] In this way, the present invention provides a personalized radio experience based on the user's interests, providing useful information while reducing visual burden. Furthermore, by integrating an emotion engine, the present invention can provide customized conversations and advertisements based on the user's current emotional state, further enhancing satisfaction.

[2304] The processing flow will be explained below.

[2305] Step 1:

[2306] The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[2307] Step 2:

[2308] The device sends the entered profile information to the server.

[2309] Step 3:

[2310] The server stores the received profile information in a database.

[2311] Step 4:

[2312] The device records usage history data such as playback time, number of skips, and selected topics when a user uses an application.

[2313] Step 5:

[2314] The terminal periodically transmits the recorded usage history data to the server.

[2315] Step 6:

[2316] The server receives the usage history data and stores it in a database.

[2317] Step 7:

[2318] The server uses AI algorithms to analyze users' interests based on their profile information and usage history data.

[2319] Step 8:

[2320] The device analyzes the user's voice and input data in real time and recognizes their emotional state through an emotion engine.

[2321] Step 9:

[2322] The terminal transmits the emotion data analyzed by the emotion engine to the server.

[2323] Step 10:

[2324] The server selects the most suitable topic for the user based on the analysis of their emotional data and interests. For example, if the user is feeling stressed, it will select relaxing topics.

[2325] Step 11:

[2326] The server generates conversation content based on the selected topic using a natural language generation model.

[2327] Step 12:

[2328] The server applies a speech synthesis model to add emotion and tone to the generated dialogue.

[2329] Step 13:

[2330] A user submits a question or comment using the in-app letter feature.

[2331] Step 14:

[2332] The terminal sends the letter from the user to the server.

[2333] Step 15:

[2334] The server receives the letter and analyzes its contents.

[2335] Step 16:

[2336] The server generates relevant information and answers based on the analysis results. For example, if a user asks, "What fitness app do you recommend?", the server will respond, "The recommended fitness app is XX."

[2337] Step 17:

[2338] The server selects relevant advertisements based on the analysis results.

[2339] Step 18:

[2340] The server encodes the generated dialogue and advertisements as audio files.

[2341] Step 19:

[2342] The server delivers the encoded audio file to the device.

[2343] Step 20:

[2344] The device plays the received audio file, allowing the user to perform operations such as stop, skip, and play during playback.

[2345] Step 21:

[2346] Users listen to what is being played and find that the content is personalized, for example, responses or advertisements tailored based on the user's current emotional state.

[2347] The above processing steps not only provide a personalized radio experience based on the user's interests, but also utilize an emotion engine to provide content optimized for the user's emotional state. This system reduces visual load while providing useful and satisfying information.

[2348] Example 2

[2349] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2350] Conventional voice information delivery systems have limited functionality for providing personalized content based on a user's profile information and usage history, resulting in issues with not being able to fully respond to user interests. Furthermore, they have not been able to adjust conversation content using real-time emotion recognition or dynamically generate content that reflects user feedback. Furthermore, there have been cases where advertising content and audio content do not always match, hindering the user experience.

[2351] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2352] In this invention, the server includes a means for collecting and saving user profile information, a means for periodically transmitting usage history data, and a means for analyzing the usage history data and profile information. This enables personalized conversation content based on the user's interests and the generation of related advertisements and their delivery as audio files. Furthermore, the user experience can be improved by recognizing the user's emotional state and adjusting the conversation content in real time.

[2353] "User profile information" is information including personal attributes of the user, such as age, sex, occupation, and hobbies.

[2354] "Means of collection and storage" refers to the technical means for receiving data entered by the user and storing it in a database or the like.

[2355] "Usage history data" refers to data such as playback time, number of skips, and selected topics that are recorded when a user uses an application.

[2356] "Means of analysis" refers to technical means using AI algorithms and machine learning models to identify user interests and concerns based on collected data.

[2357] "Means for generating conversational content" means the technical means for creating natural language responses and conversations using a generative AI model based on the analysis results.

[2358] The "means for receiving letters and analyzing their contents" refers to the technical means for receiving questions and comments sent by users and processing them using a text analysis model.

[2359] The "means for generating relevant information and answers" refers to the technical means for generating appropriate information and answers based on the analyzed content of the letter.

[2360] "Means for selecting relevant advertisements and generating audio files of conversations and advertisements" refers to the technical means for selecting highly relevant advertisements based on the user's interests and generating audio data that integrates the advertisements into the conversation content.

[2361] The "means for playing audio files" refers to the technical means for appropriately playing audio files generated on a user terminal.

[2362] The "means for recognizing emotional states" refers to technical means for analyzing the user's voice and input data in real time and identifying the user's emotions.

[2363] A "conversational content adjustment mechanism" is a technological mechanism for changing the tone or content of a conversation in real time based on a perceived emotional state.

[2364] This invention is an AI radio system that combines an emotion engine that recognizes the user's emotions. The specific processing of the system program is shown below.

[2365] The present invention is a system for providing voice information based on user operations, and has the following main functions.

[2366] User Data Collection and Initial Settings

[2367] User:

[2368] First, the user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[2369] Device:

[2370] The device establishes an Internet connection and sends an HTTP request to the server to transmit the entered profile information to the server.

[2371] server:

[2372] The server receives the HTTP request and stores the profile information in a database, using a database management system such as MySQL.

[2373] Daily data collection

[2374] Device:

[2375] Every time the application is launched, it locally records user actions (playback time, skip count, selected topics, etc.) and saves the data in JSON format.

[2376] server:

[2377] The device sends usage history data to the server at regular intervals or when a specific event occurs. The server stores the received data in a database and processes it on a cloud server for data analysis.

[2378] Generate personalized conversations

[2379] server:

[2380] The server analyzes profile information and usage history data to select the most suitable topics for each user. It uses machine learning libraries such as Python's scikit-learn and TensorFlow to generate prompts for a natural language generation model (e.g., OpenAI GPT-3) and sends them to the model via an API.

[2381] Specific prompt examples:

[2382] "Users are interested in technology, so talk in detail about the latest smartwatch features."

[2383] The server uses the received generated text to generate an audio file using a speech synthesis model (e.g., Google Text-to-Speech) taking into account emotions and tone.

[2384] Emotion recognition and conversation adjustment with emotion engine

[2385] Device:

[2386] The device captures the user's voice input through a microphone and analyzes it in real time using an emotion engine (e.g., Affectiva Emotion AI).

[2387] server:

[2388] The system receives emotional data from the device and stores it in a database. Based on the analysis results, it adjusts the conversation content accordingly. For example, if the user is feeling stressed, it will regenerate the conversation content to provide a calming tone and relaxing topics.

[2389] Processing letters and providing information

[2390] User:

[2391] Users use the letter feature within the app to enter questions or comments and press the send button.

[2392] Device:

[2393] The terminal converts the input content into JSON format and sends it to the server as an HTTP POST request.

[2394] server:

[2395] The server analyzes the received letter using a text analysis model (e.g., spaCy), generates related information and recommendations using a generative AI model (GPT-3), generates the generated answers as audio files, and selects and integrates advertisements into the audio data.

[2396] Delivering personalized ads

[2397] server:

[2398] The server selects relevant ads based on the user's profile information and current conversation content, and encodes the generated conversation content and ads as audio files.

[2399] Specific examples of behavior:

[2400] For technology-loving users, ads about the latest smartwatch features are seamlessly integrated into the conversation.

[2401] Content Delivery and Playback

[2402] server:

[2403] The server sends the generated audio file as an HTTP response to deliver it to the user terminal using an HTTP framework (e.g., Flask).

[2404] Device:

[2405] The device uses the audio playback feature (e.g., AndroidMediaPlayer) to play the audio file, and the user can perform operations such as play, stop, and skip through the app interface.

[2406] User:

[2407] The user listens to the generated audio content and submits feedback or follow-up questions as needed.

[2408] In this way, the system of the present invention delivers personalized audio content based on the user's interests and generates finely tuned responses based on real-time emotion recognition, resulting in a more satisfying user experience.

[2409] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2410] Step 1: Collecting user data and initial setup

[2411] User

[2412] Input: The user installs the application and enters profile information (age, gender, occupation, hobbies, etc.) when launching it for the first time.

[2413] Specific behavior: A user enters profile information into an application's input form and presses the submit button.

[2414] Terminal

[2415] Data processing: The entered profile information is temporarily saved locally.

[2416] Output: This information is converted to JSON format and sent to the server as an HTTP request.

[2417] What happens: The device establishes an Internet connection and generates an HTTP POST request containing the entered profile information.

[2418] server

[2419] Input: The server receives the profile information sent from the device.

[2420] Data processing: Analyze the received information and store it in a database.

[2421] Output: Generates a save confirmation message.

[2422] Specific operation: The server stores the received profile information using a database management system such as MySQL.

[2423] Step 2: Collecting data on a daily basis

[2424] Terminal

[2425] Input: Operational data when a user uses the application (playback time, number of skips, selected topics, etc.).

[2426] Data processing: Save usage history data in JSON format to local storage.

[2427] Output: Set a trigger to periodically send the saved data to the server.

[2428] Specific operation: Detects user operations within the application and stores data locally.

[2429] server

[2430] Input: Usage history data sent periodically from the device.

[2431] Data processing: The received usage history data is stored in a database and sent to an analysis server.

[2432] Output: Generates data analysis results.

[2433] Specific operation: The server analyzes the usage history data stored in the database using an AI algorithm to identify the user's interests and concerns.

[2434] Step 3: Generate personalized conversations

[2435] server

[2436] Input: Analysis of user profile information and usage history data.

[2437] Data processing: Based on the analysis results, prompt sentences are generated for the generative AI model (e.g., GPT-3).

[2438] Output: The generated conversation.

[2439] Specific operation: The server analyzes the data using Python's scikit-learn and TensorFlow and generates a prompt. The generated prompt is sent to the generative AI model via an API and an answer is received. "The user is interested in technology, so please talk in detail about the latest smartwatch features."

[2440] server

[2441] Input: The speech received from the generative AI model.

[2442] Data processing: Applying speech synthesis models to add emotion and tone to the conversation.

[2443] Output: An audio file of the generated conversation.

[2444] What it does: The server uses a speech synthesis model such as Google Text-to-Speech to convert the generated text into an audio file.

[2445] Step 4: Emotion recognition and conversation adjustment using the emotion engine

[2446] Terminal

[2447] Input: User speech and input data.

[2448] Data processing: Analyze emotional states in real time using an emotion engine.

[2449] Output: Emotion data as the analysis result.

[2450] Specific operation: The device captures the user's voice using a microphone device and analyzes it in real time using an emotion engine (e.g., Affectiva Emotion AI).

[2451] server

[2452] Input: Emotion data sent from the device.

[2453] Data processing: Adjusting conversation content in real time based on emotional data.

[2454] Output: Improved conversation.

[2455] Specific operation: Based on the analysis results, the server re-prompts the generative AI model as necessary and corrects the conversation content.

[2456] Step 5: Processing your letter and providing information

[2457] User

[2458] Input: User comments and questions.

[2459] What happens: A user uses the in-app message feature, enters a question or comment, and presses the send button.

[2460] Terminal

[2461] Input: The letter data entered by the user.

[2462] Data processing: Generate an HTTP POST request in JSON format and send it to the server.

[2463] Output: Send to server.

[2464] Specific operation: The device converts user input into JSON format and sends it to the server via the Internet.

[2465] server

[2466] Input: Letter data sent from the terminal.

[2467] Data processing: Analyze using text analytics models (e.g., spaCy) and generate relevant answers and information using generative AI models.

[2468] Output: The generated answer and related information.

[2469] Specific operation: The server receives the letter data, analyzes it using a natural language processing model, and creates an appropriate response using a generative AI model, which then generates it as an audio file.

[2470] Step 6: Delivering personalized ads

[2471] server

[2472] Input: User profile information and current conversation.

[2473] Data processing: Selecting relevant ads and merging the conversation and ads into an audio file.

[2474] Output: An audio file containing ads.

[2475] How it works: Based on the user information and conversation content, the server selects relevant advertisements and encodes them into audio files.

[2476] Specific examples

[2477] For technology-loving users, ads for the latest smartwatches are selected and seamlessly integrated into conversations.

[2478] Step 7: Content Delivery and Playback

[2479] server

[2480] Input: The generated audio file.

[2481] Data processing: Audio files are delivered to the user's device using the HTTP framework.

[2482] Output: Audio file delivery to user device.

[2483] Specific operation: The server uses an HTTP framework such as Flask to send the audio file to the terminal as an HTTP response.

[2484] Terminal

[2485] Input: The audio file sent from the server.

[2486] Data processing: Play audio files using the audio playback function.

[2487] Output: Playback of audio content.

[2488] Specific behavior: The device plays the audio file using AndroidMediaPlayer or similar, and allows the user to control playback through the interface.

[2489] User

[2490] What it does: The user listens to the generated audio content and provides feedback or asks follow-up questions if necessary. This data is used to generate future content.

[2491] (Application example 2)

[2492] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2493] Conventional radio systems and content distribution services have difficulty providing personalized content based on a user's current emotional state. Furthermore, they are unable to respond to a user's emotional changes in real time, resulting in a poor user experience. Furthermore, in the case of advertising delivery, insufficient personalization based on the user's current situation results in a failure to attract the user's attention and a reduced effectiveness.

[2494] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2495] In this invention, the server includes means for analyzing the user's voice and input data in real time and incorporating an emotion engine that recognizes the user's emotional state, means for adjusting topics and conversation content in real time based on the emotion data, and means for receiving letters from users, analyzing their content, and generating related information and answers, thereby making it possible to provide appropriate content in real time that matches the user's emotional state.

[2496] "User profile information" is basic personal information such as age, sex, occupation, and hobbies that a user enters when using the service for the first time.

[2497] "Usage history data" refers to behavioral data such as playback time, number of skips, and selected topics when a user uses an application.

[2498] The "emotion engine" is an engine that has the function of analyzing the user's voice and input data in real time and recognizing their emotional state.

[2499] The "server" is a computer system that stores and analyzes data collected from users and generates and adjusts topics and conversation content based on their emotional state and interests.

[2500] A "speech synthesis model" is a technology for generating speech based on computer-generated text and reproducing emotion and tone.

[2501] "Emotion data" is data that indicates the user's current emotional state as analyzed by the emotion engine.

[2502] "Advertisements" are commercial information that is selected based on the user's interests and incorporated naturally into conversations.

[2503] "Letters" are messages such as questions, comments, and feedback that users send through the application.

[2504] A "topic" is a conversation subject selected by the server based on the user's interests and concerns.

[2505] The "audio file" is an audio data file in which the generated conversation content or advertisement content is encoded, and is played back on the user terminal.

[2506] In this invention, a system for exchanging data between a user terminal and a server to provide personalized content to the user will be specifically described. The detailed program processing for realizing this system will be described below.

[2507] User Data Collection and Initial Settings

[2508] After installing the application, the user enters their profile information the first time they start it. This information includes, for example, age, gender, occupation, and hobbies. This information is sent from the user's device to the server, which then stores it in a database. The hardware used at this stage is a smartphone, and the software includes front-end technology (e.g., React Native) that configures the user interface (UI) and a library (e.g., requests) that sends HTTP requests.

[2509] Daily data collection

[2510] When a user uses an application, usage history data such as playback time, number of skips, and selected topics is recorded. This data is periodically sent to a server, which stores the received data in a database. The server uses an AI algorithm to analyze this data and identify the user's interests. The software used for this analysis is a database management system (e.g., MySQL) and an AI analysis algorithm (e.g., TensorFlow).

[2511] Emotion recognition and conversation adjustment with emotion engine

[2512] The user device analyzes the user's voice and input data in real time and recognizes their emotional state using an emotion engine. The hardware used to collect voice data is the smartphone's microphone, and the software is a voice analysis library (e.g., Google Cloud Speech-to-Text API). The server analyzes the user's current emotional state based on the emotion data sent from the emotion engine and adjusts the topic and content of the conversation in real time. For example, if the user is feeling stressed, it generates a relaxing tone of voice.

[2513] Generate personalized conversations

[2514] The server selects the most suitable topic for the user based on the analysis results and generates the conversation using a natural language generation model, while also applying an appropriate speech synthesis model to add emotion and tone. For example, if the user prefers topics related to fitness, the server will generate a conversation about "recommended fitness apps."

[2515] Processing letters and providing information

[2516] Users can use the application's message feature to send questions or comments. For example, they can ask, "What fitness apps do you recommend?" These messages are sent from the user's device to a server, which analyzes the content and generates relevant information and answers. This analysis and generation is performed using a natural language processing model (e.g., GPT-3).

[2517] Delivering personalized ads

[2518] The server selects relevant ads based on the user's profile information and current conversation, and encodes the generated conversation and ads into audio files. For example, a technology-loving user might receive an ad for the latest wearable devices. The selected ads are seamlessly integrated into the conversation.

[2519] Content Delivery and Playback

[2520] The generated audio file is delivered from the server to the user's device, where it is played. When the user listens to AI Radio, they can control playback (pause, skip, play, etc.) as needed. The hardware used in this part is a smartphone, and the software is a media player library with audio playback capabilities.

[2521] Prompt Sentence Examples

[2522] For example, here is a prompt that describes a situation where the user is feeling stressed about fitness:

[2523] Example prompt sentence:

[2524] "A user who has asked a fitness question is stressed. We recommend fitness apps to that user in a relaxing tone."

[2525] In this way, it becomes possible to provide personalized content in real time according to the user's emotional state and interests. This invention not only improves the user experience, but also realizes highly personalized content delivery to increase advertising effectiveness.

[2526] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2527] Step 1:

[2528] The user installs the application and enters profile information when the application is first launched.

[2529] Examples of input: name, age, gender, occupation, hobbies, etc.

[2530] The terminal transmits the input data to the server, which stores the received data in a database.

[2531] Output: User profile information is stored in the server database.

[2532] Step 2:

[2533] When a user uses the application, usage history data such as playback time, number of skips, and selected topics is recorded.

[2534] Example inputs: playback data, skip counts, topic selection.

[2535] The terminal periodically transmits the collected data to the server, and the server stores the received data in a database.

[2536] Output: Usage history data is stored in the server database.

[2537] Step 3:

[2538] The server uses AI algorithms to analyze the collected profile information and usage history data to identify the user's interests.

[2539] Examples of input: profile information, usage history data.

[2540] The server uses AI algorithms (e.g., TensorFlow) to analyze the data.

[2541] Output: Analysis results based on user interests and concerns.

[2542] Step 4:

[2543] The server uses an emotion engine that analyzes the user's voice and input data in real time to recognize their emotional state.

[2544] Examples of input: user voice data, text input data.

[2545] The device sends the audio to an emotion engine (e.g., Google Cloud Speech-to-Text API) for sentiment analysis.

[2546] Output: Emotion data indicating the user's current emotional state.

[2547] Step 5:

[2548] The server adjusts topics and conversation content in real time based on emotional data.

[2549] Examples of input: emotional data, historical usage data.

[2550] The server generates the conversation using a natural language generation model (e.g., GPT-3) and makes any necessary adjustments.

[2551] Output: Conversational content that reflects the user's emotional state.

[2552] Step 6:

[2553] Users can submit questions or comments using the application's letter feature.

[2554] Example input: A user question or comment.

[2555] The terminal sends input from the user to the server, which analyzes the content.

[2556] Output: Analysis results and related information and answers.

[2557] Step 7:

[2558] ...

Claims

1. means for inputting user profile information; means for collecting and storing the input profile information; means for recording user usage history data; means for periodically transmitting the recorded usage history data to a server; A means for analyzing the user's interests and concerns based on the transmitted usage history data and profile information; a means for selecting a topic based on the analysis result and generating conversation content; A means for receiving letters from users, analyzing the contents, and generating related information and answers; means for selecting relevant advertisements based on the user's interests and generating a voice file of the conversation and advertisements; means for delivering the generated audio file to a user terminal; The system includes a user terminal including means for playing said audio file.

2. 10. The system of claim 1, further comprising means for applying a speech synthesis model to add emotion and tone to the generated dialogue.

3. 2. The system according to claim 1, wherein the user terminal comprises means for accepting operations such as stop, skip, and play by the user during playback.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A