system

The system addresses the challenge of ineffective digital advertising by analyzing user behavior and multimodal information to deliver personalized ads at optimal times, enhancing engagement and effectiveness through continuous algorithm improvement.

JP2026074915APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing digital advertising systems struggle to provide advertisements that accurately match user interests and preferences in real time, leading to decreased effectiveness and user engagement.

Method used

A system that analyzes user behavior information and multimodal information using a generation algorithm to create personalized advertisements, delivering them at optimal times and formats while tracking user responses for algorithm improvement.

Benefits of technology

Enhances advertising effectiveness by providing personalized ads that align with user interests and preferences, improving engagement through real-time adjustments based on user behavior and emotional states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074915000001_ABST
    Figure 2026074915000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of collecting user behavior information, Means for obtaining multimodal information, A means of generating user-optimized ad content using a generation algorithm, A means for delivering the generated advertisement to the user's device, A means of tracking user responses to advertisements and using that information to improve generation algorithms, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the field of digital advertising, it is difficult to provide advertisements that suit the interests and concerns of users in real time, resulting in a problem that the effectiveness of advertisements decreases. In conventional methods, only advertisements can be generated from limited data, and there is a lack of provision of content that matches the interests of users, so there is a problem that user engagement does not improve.

Means for Solving the Problems

[0005] This invention provides a system that accurately grasps user interests by analyzing user behavior information and multimodal information, and generates optimized advertising content using a generation algorithm. This allows for the delivery of generated advertisements to user devices at the appropriate time and in the appropriate format, and improves advertising effectiveness by tracking user responses to advertisements and using the results to improve the generation algorithm.

[0006] "User behavior information" refers to data about the actions a user takes in a digital environment, including website browsing history and click data.

[0007] "Multimodal information" is a general term for data in different information formats such as text, audio, images, and videos, and is used to determine the user's interests and preferences.

[0008] A "generative algorithm" is a computational method for generating a specific output from input data, and in this case, it is used to generate user-optimized advertising content.

[0009] "User equipment" refers to electronic devices used by users to access the internet or digital services, and includes personal computers, smartphones, tablets, and other similar devices.

[0010] "Ad response tracking" refers to the process of recording and analyzing user behavior and evaluations of delivered advertisements, and is used to measure the effectiveness of advertisements and improve the generation algorithms. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, let's explain the terminology used in the following explanation.

[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] This invention is a system for providing advertisements that more accurately reflect users' interests and preferences. In this system, the server primarily performs the functions of analyzing user behavior information and multimodal information, and generates and delivers optimal advertisements using a generation algorithm.

[0033] The server first stores user activity information. This involves collecting a wide range of data, including website browsing history, social media interactions, and digital content usage. This information forms the basis for understanding user interests and preferences.

[0034] Next, the server retrieves multimodal information from the user's device. This includes photos and videos taken on the device, recorded audio, and application usage data. The retrieved information is analyzed by an AI model to contribute to creating a user profile.

[0035] The generation algorithm analyzes this information and generates ads that best match the user's preferences and behavior. For example, if a user has recently been searching for a lot of outdoor-related products, it can create video ads for camping equipment related to that.

[0036] The generated ads are delivered to the user's device at the appropriate time. For example, the server can identify when a user is relaxing and watching a video, and insert a video ad at that moment. Audio ads, on the other hand, are delivered in a way that blends seamlessly into the user's feed while they are listening to music.

[0037] User responses to ads are tracked by the server. Metrics such as click-through rates, ad viewing time, and purchase behavior are analyzed, and this data is used to further improve the generation algorithm. This ensures that delivered ads are continuously better suited to user interests and preferences, maximizing ad effectiveness.

[0038] In this way, the system of the present invention aims to provide personalized advertisements to users and enhance the effectiveness of advertising.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The server collects behavioral information from the user's device and online activity. It tracks website browsing history, click data, and social media activity to accumulate data that helps infer the user's interests.

[0042] Step 2:

[0043] The server collects multimodal information from the user's device. It retrieves data captured in the form of images, videos, and audio recordings to gain a deeper understanding of the user's interests and preferences.

[0044] Step 3:

[0045] The server uses an AI model to analyze the collected behavioral and multimodal information. This analysis generates user profiles and provides insights into which advertisements are most effective.

[0046] Step 4:

[0047] The server uses a generation algorithm to create ad content optimized for the user. This process automatically generates ads with creative elements based on the user's preferences.

[0048] Step 5:

[0049] The server determines the optimal timing and format for delivering the generated advertisements to the user's device. For example, it might identify times when users are relaxed and consuming media, and insert advertisements at those times.

[0050] Step 6:

[0051] The server monitors user reactions after ad delivery and collects data such as click-through rates and ad completion rates. This feedback is used to improve the AI ​​model, enhancing the accuracy and effectiveness of future ad generation.

[0052] (Example 1)

[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0054] In today's world, there is a need to effectively deliver advertisements that accurately match users' diverse interests and preferences. However, existing advertising systems struggle to personalize ads with high accuracy based on individual user preferences and behaviors, which limits the effectiveness of advertising.

[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0056] In this invention, the server includes means for collecting user information from multiple data sources, means for acquiring multimodal data from information devices, and means for analyzing the generated data and analyzing user attributes using machine learning algorithms. This enables the generation and delivery of optimized advertisements based on user interests and behavior, thereby improving the personalization and effectiveness of advertisements.

[0057] "User information" refers to data about a user's online and offline behavior, such as website browsing history, social media activity, and use of digital content.

[0058] "Multimodal data" is a general term for data that includes information acquired from multiple different sensors and devices, such as photos, videos, audio, and application usage data.

[0059] "Information devices" refer to electronic devices used by users, such as smartphones, tablets, and personal computers, from which data can be acquired.

[0060] A "machine learning algorithm" is a technique that uses computational models to extract patterns from data and predict future data and behavior.

[0061] "User terminals" refer to electronic devices that users use on a daily basis, and are the devices to which advertisements are delivered.

[0062] "Personalization" refers to the process of adjusting content to suit the individual user's interests and preferences, providing an optimized experience.

[0063] "Ad delivery" refers to the process of sending advertising content to a user's device and making it available for the user to view.

[0064] "Dynamic adjustment" refers to changing the system's operation in real time according to time and circumstances to achieve optimal performance.

[0065] This invention is a system for providing user-optimized advertisements. This system primarily relies on a server, which collects and analyzes information from various data sources and generates and delivers user-specific advertisements, thereby enhancing the effectiveness of the advertisements.

[0066] The server first collects the user's online activity, specifically web browsing history, social media interaction history, and digital content usage. This information is obtained using cookies, APIs, and log files and stored in a database. The obtained data serves as important foundational material for analyzing the user's interests and preferences.

[0067] Next, the device acquires multimodal data using its camera and microphone. This includes photos, videos, and audio files, and this data is processed using image recognition and audio analysis algorithms. The information obtained through this process is sent to a server to build the user's profile.

[0068] The generative AI model analyzes user behavior characteristics based on collected data and generates optimal advertisements to reach users. By using machine learning techniques, it is possible to create advertisements that are tailored to users' purchasing trends and preferences based on previously collected examples. For example, by entering a prompt such as, "Generate advertisements that suggest outdoor products that a user might be interested in, based on their recent search history and multimodal information," highly relevant ad content will be generated.

[0069] The generated ads are delivered by the server at the optimal time, tailored to the user's specific behavior patterns and lifestyle. For example, video ads are delivered at night when the user is relaxing, and audio ads are adjusted to play naturally while the user is listening to their favorite music.

[0070] In this way, the system of the present invention provides advertisements tailored to the needs of each user, thereby achieving maximum advertising effectiveness.

[0071] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0072] Step 1:

[0073] The server collects users' online behavior data. Specifically, it obtains web browsing history, social media activity logs, and digital content usage data using cookies and various APIs. This data forms the basis for creating personalized advertisements and is stored in a database. The input is user behavior data, and the output is the stored dataset.

[0074] Step 2:

[0075] The device acquires multimodal data using its camera and microphone. This process utilizes captured photos, videos, and recorded audio data. This data is processed using image recognition and audio analysis algorithms, resulting in multimodal information that reflects the user's preferences and interests. The input consists of photos, videos, and audio, while the output is the analyzed information.

[0076] Step 3:

[0077] The server integrates all collected data and builds user profiles using a generative AI model. It analyzes the data with machine learning algorithms to predict user interests, preferences, and purchasing tendencies. Inputs are integrated behavioral and multimodal information, and output is a detailed user profile.

[0078] Step 4:

[0079] The server generates the most relevant ads based on the user profile. It prompts the generation AI model with a message such as, "Generate ads suggesting outdoor products that the user might be interested in, based on their recent search history and multimodal information," and generates highly relevant ads. The input is the user profile and the prompt, and the output is the generated ad content.

[0080] Step 5:

[0081] The server delivers the generated ad content to the user's device. The ad delivery time is set to the most effective time based on the user's lifestyle. For example, a video ad might be delivered on Sunday afternoon. The input is the generated ad and its delivery schedule, and the output is the ad played on the device.

[0082] Step 6:

[0083] The server monitors user responses to ads and collects metrics such as click-through rates, viewing time, and purchase behavior. This feedback is used to further improve the generative AI model. The input is user response data, and the output is the improved ad generation algorithm.

[0084] (Application Example 1)

[0085] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0086] In modern society, as advertising personalization advances, there is a growing demand for meaningful and engaging ad delivery for users. However, traditional advertising systems have limitations in their real-time relevance to user interests and behavior, making it difficult to provide information naturally without disrupting the user's visual experience. Therefore, establishing interactive and effective ad delivery methods based on user interests, utilizing visual devices, has become a key challenge.

[0087] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0088] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, means for generating user-optimized advertising content using a generation algorithm, means for delivering the generated advertisements to the user's device, means for tracking the user's response to the advertisements and using this information to improve the generation algorithm, and means for directly overlaying advertisements onto the visual information viewed by the user. This enables the provision of optimal advertising content in real time to the visual device being used by the user, and allows for effective information transmission without compromising the visual experience.

[0089] "User behavior information" is a general term for data related to user actions, and includes behavioral data such as user click patterns, navigation history, and location information.

[0090] "Multimodal information" refers to data collected from multiple different sources, including data acquired through multiple senses such as sound, images, and text.

[0091] A "generative algorithm" is a programmatic procedure that generates output tailored to a specific purpose based on input data, and in particular, it refers to the process of creating advertisements that reflect the user's interests and preferences.

[0092] "User devices" refer to electronic information terminals that users use on a daily basis, and specifically include smartphones, tablets, smart glasses, and the like.

[0093] "User response" refers to the actions and attitudes that users show towards advertisements, and specifically includes the number of clicks, time spent on the page, and the point where users stop scrolling.

[0094] "Methods for directly overlaying advertisements onto visual information" refer to technologies that superimpose the display of advertisements onto objects viewed by users through their visual devices, creating a mechanism that allows advertisements to appear in the user's field of vision in a natural way.

[0095] In this embodiment of the invention, a system is realized in which a server collects user behavior information and multimodal information and generates advertisements using a generation algorithm. The server acquires data from user devices such as smart glasses and creates a user profile using AI software such as TENSORFLOW®. Based on this, advertisements tailored to the user's interests are generated and displayed as an overlay on the visual device. This process makes it possible to provide relevant advertisements in real time in relation to the surrounding environment that the user is viewing.

[0096] User responses to advertisements are collected through sensors and interfaces on the user's device. This allows for the measurement of the advertisement's effectiveness, and the data is used to improve the generation algorithm. For example, if it is detected that a user has been lingering their gaze on an overlaid advertisement for an extended period, the advertisement is evaluated as visually effective.

[0097] For example, when a user is walking through a shopping mall, a promotional advertisement for that store can be displayed on the reflection of the store in their glasses. In this case, the advertisement is placed in the user's field of vision in a natural way, and the information is presented in a manner that is likely to capture the user's interest.

[0098] An example of a prompt message is: "I want to design an application that overlays advertisements for stores that a user sees in their line of sight onto their glasses as they walk around. Please tell me the necessary steps and technologies to use to analyze the user's profile with AI and select the appropriate advertisement in real time."

[0099] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0100] Step 1:

[0101] Activity information is collected from the user's device. The device records the user's location, eye-tracking data, and device usage in real time. This data serves as foundational data for understanding the user's activities and interests. User activity data is obtained as input, and pre-processed data for analysis is provided as output.

[0102] Step 2:

[0103] The server retrieves multimodal information from the user's terminal. This multimodal information includes voice commands, image data from the camera, and application operation logs. The input is the user's multimodal data, and the output is data converted into a format that can be analyzed by the AI ​​model.

[0104] Step 3:

[0105] The server analyzes the collected behavioral and multimodal information. Using an AI generation algorithm, it processes this data and generates a profile based on the user's interests and behavioral patterns. The input is the analysis data obtained in the previous step, and the output is the user's interest profile.

[0106] Step 4:

[0107] The server generates optimal advertisements based on the user's profile. Using a generative AI model, it automatically creates advertisements tailored to the user's interests. Specifically, it selects advertising materials related to products and services the user has shown interest in, and generates customized advertisements through the system. The input is the user profile, and the output is the generated advertising content.

[0108] Step 5:

[0109] The server delivers the generated advertisements to the user's visual device. The device overlays the advertisements on the smart glasses' display, seamlessly blending them into the user's real-world environment. The input is the advertisement content, and the output is its display on the smart glasses.

[0110] Step 6:

[0111] The server tracks user responses to advertisements through visual devices. The terminal uses sensors to record user eye-tracking data and interface interactions, which are then analyzed to evaluate the effectiveness of the advertisements. The input is user response data, and the output is the evaluation result of the advertisement performance.

[0112] Step 7:

[0113] The server uses user response data to improve the generation algorithm. The collected data is incorporated into the algorithm as feedback, further improving it to better match user preferences in future ad generation. The input is an evaluation of the ad response, and the output is the improved ad generation algorithm.

[0114] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0115] This invention is an advertising system that combines user behavior information, multimodal information, and an emotion engine that recognizes user emotions. The purpose of this system is to generate and deliver personalized advertisements that are tailored to the user's interests and emotions.

[0116] The server first obtains activity information from the user's device. This includes website browsing history and social media activity, which is used to infer the user's interests and preferences. Next, the server collects multimodal information from the device. This allows the server to create a more detailed profile of the user through the analysis of text data, image data, and audio data.

[0117] Furthermore, in this invention, the server uses an emotion engine to analyze the user's emotions in real time. This emotion engine analyzes the user's facial expressions, tone of voice, and input data to recognize the user's emotional state. This information is a crucial element for more effectively personalizing advertisements.

[0118] For example, if the emotion engine detects a high level of excitement while a user is watching a video, the server can generate an energetic ad that matches that state. Conversely, if it detects a relaxed state, it will generate an ad with a more subdued tone.

[0119] The generation algorithm integrates all of this data to produce optimal ad content. Because this ad is tailored based on the user's interests and emotional state, it can achieve a more effective appeal.

[0120] The generated advertisements are delivered to the user's device at a time and format dynamically adjusted by the server. This process ensures optimal ad delivery based on the user's usage patterns and the type of content they are viewing.

[0121] Ultimately, the server tracks user responses to ads, analyzing click-through rates, viewing time, emotional changes, and more. This feedback can be used to improve future ad generation and enhance the overall effectiveness of advertising campaigns.

[0122] The following describes the processing flow.

[0123] Step 1:

[0124] The server collects activity information from the user's device. This includes web history, social media activity, and purchase history, and uses this data to understand the user's interests and preferences.

[0125] Step 2:

[0126] The server retrieves multimodal information from the user's device. This refers to data such as images, videos, and audio, and is used to understand the user's content consumption patterns.

[0127] Step 3:

[0128] The server uses an emotion engine to analyze the user's current emotional state. It analyzes data obtained from the device's camera and microphone, and evaluates the user's emotions through facial expressions and tone of voice.

[0129] Step 4:

[0130] The server integrates and analyzes user behavior information, multimodal information, and emotional state to generate optimal advertisements. In this process, the generation algorithm personalizes content according to the user's characteristics.

[0131] Step 5:

[0132] The server determines the optimal timing and format for delivering generated ads based on the user's usage patterns. Ads are displayed when the user is watching videos or during quiet periods.

[0133] Step 6:

[0134] The server monitors user responses to delivered advertisements. This includes tracking ad click-through rates, view completion rates, and changes in user sentiment.

[0135] Step 7:

[0136] The server establishes a feedback loop between the emotion engine and the generation algorithm based on the collected response data. This enables continuous improvement to enhance the accuracy of future ad generation and delivery.

[0137] (Example 2)

[0138] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0139] Traditional advertising systems, while capable of generating personalized ads based on user interests, were insufficient in considering users' real-time emotional states. Furthermore, they lacked adequate optimization of dynamic ad delivery based on user behavior. This made it difficult to capture user interest in ads, highlighting the need for improved advertising effectiveness.

[0140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0141] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, and means for performing sentiment analysis. This makes it possible to generate advertisements optimized based on the user's emotional state and interests and deliver them to the user's terminal in real time. This is expected to improve the effectiveness of the advertisements.

[0142] "Means for collecting user behavior information" refers to methods for obtaining data about user behavior, such as a user's web browsing history and social media activity.

[0143] "Means for acquiring multimodal information" refers to technologies that collect information in multiple media formats, such as text, images, and audio data, from a user's device.

[0144] "Methods for performing emotion analysis" refer to the process of analyzing a user's facial expressions, tone of voice, and input data to identify the user's emotional state.

[0145] "Using a generation algorithm" means using a computational method to generate user-optimized advertisements based on collected data.

[0146] "Means of delivering to user devices" refers to technologies that transmit generated advertisements to the user's terminal at the appropriate time and in the appropriate format.

[0147] "Means of tracking user responses" refer to methods for recording user behavior in response to advertisements, such as click-through rates and viewing time, in order to evaluate the effectiveness of the advertisements.

[0148] This invention is a system that generates and delivers advertisements that reflect the user's interests and emotions in real time. The server collects various information from the user's terminal and uses this information to personalize advertisements. Specifically, it obtains user behavior information, such as website browsing history and social media activity, and analyzes this information to infer the user's interests.

[0149] Furthermore, the server collects multimodal information. This information includes text, images, and audio, and is obtained from the user's device using the camera and microphone. By analyzing this data, a more detailed profile of the user can be created.

[0150] In emotion analysis, the server analyzes the user's emotional state in real time. Using an emotion engine, the AI ​​model analyzes the user's facial expressions and tone of voice to identify emotions such as joy, excitement, and relaxation. This makes it possible to adjust the content of advertisements to best suit the user's current emotions.

[0151] Ad generation utilizes a generative AI model. This model comprehensively processes collected data to automatically create ads that match the user's interests and emotions. The generated ads are then delivered to the user's device by a server at the optimal time. This delivery is dynamically adjusted according to the user's usage of the content.

[0152] For example, if a user is searching for a new recipe, the server can generate advertisements for cooking utensils that match their interests and provide the user with enjoyment and value by using cooking-related prompts as examples. In this way, advertisements based on user interests and emotions contribute significantly to improving advertising effectiveness.

[0153] An example of a prompt message would be, "The user has been detected as highly aroused while watching a video. Generate an ad for a new product that might interest them," which enhances ad personalization.

[0154] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0155] Step 1:

[0156] The user's device sends activity information to the server. Specifically, it transfers website browsing history and social media activity data from the device to the server. The server uses this data to build an initial dataset to infer the user's interests. The user's behavioral history data is used as input, and the output generates an initial set of tags indicating the user's interests.

[0157] Step 2:

[0158] The server collects multimodal information from the user's device. This information includes text, images, and audio data, captured, for example, via the camera and microphone of a smartphone. The server analyzes this data to generate a detailed user profile. The input is the user's text, images, and audio data, and the output is an enhanced version of the user profile.

[0159] Step 3:

[0160] The server initiates emotion analysis, identifying the user's emotional state in real time. It uses an emotion engine to analyze facial expressions and voice tone. The results of this analysis are output as data indicating the user's current emotional state. The input is data of the user's facial expressions and voice, and the output is a tag indicating the user's emotional state at that moment.

[0161] Step 4:

[0162] The generative AI model generates personalized advertisements based on previously obtained data (user interest tags, detailed profiles, and emotional states). The input is the user's overall profile information, and the output is advertising content optimized based on that profile. Specifically, the ad's colors and word choices are tailored to the user's emotions.

[0163] Step 5:

[0164] The server sends the generated advertisements to the user's device. The timing and format of delivery are dynamically adjusted, taking into account the user's content consumption habits and emotional state. The input is the generated ad data, and the output is the optimized advertisement displayed on the user's screen. The device presents the received advertisement to the user in the specified manner.

[0165] Step 6:

[0166] The server tracks user responses and analyzes the effectiveness of advertisements. It collects data on user clicks and viewing time to evaluate the success rate of advertisements. The input is user behavior data regarding advertisements, and the output is an evaluation of advertisement effectiveness that is reflected in the next advertisement generation. This allows the server to improve its generation algorithm.

[0167] (Application Example 2)

[0168] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0169] Traditional advertising delivery systems generate ads based only on user behavior information and limited data, resulting in ads that do not adequately address user interests. Furthermore, they fail to consider the user's real-time emotional state, hindering the maximization of ad effectiveness and user engagement.

[0170] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0171] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, means for using an emotion engine to recognize the user's emotional state, and means for using a generation algorithm to generate advertising content optimized for the user based on the emotional state. This enables the delivery of advertisements optimized based on the user's real-time emotions and interests.

[0172] "User activity information" refers to data about a user's behavior and actions on the internet, including website browsing history and social media activity history.

[0173] "Multimodal information" refers to information that encompasses multiple different types of data, such as text, images, and audio, and is used in combination with these other data.

[0174] An "emotion engine" is an algorithm or software that analyzes a user's facial expressions, voice tone, and input data to identify their emotional state.

[0175] A "generative algorithm" is a computational method for creating optimized ad content based on the user's interests and emotional state.

[0176] "User devices" refer to devices used to display advertisements and information, and include smartphones and tablet devices.

[0177] "Real-time" is a term that indicates that information is processed or delivered at the same time as, or very close to, the time it is generated.

[0178] The system of this invention combines the collection of user behavior information, acquisition of multimodal information, sentiment analysis using an emotion engine, ad optimization using a generation algorithm, and ad delivery to provide users with optimized ads in real time.

[0179] The server collects behavioral information from the user's device. Specifically, it receives web browsing history and social media activity information, and uses this data to infer interests and preferences. Next, it acquires multimodal information from devices equipped with cameras and microphones. This includes the collection of real-time audio and image data. The server analyzes this data and recognizes emotional states using an emotion engine. This emotion engine is built using machine learning libraries such as TensorFlow and PyTorch.

[0180] Based on the analyzed information, the server uses a generation algorithm to create advertisements optimized for the user's emotional state. This generation algorithm utilizes Apache® Hadoop, which can process large amounts of data at high speed, enabling the rapid generation of content that reflects the user's real-time emotional state. The generated advertisements are immediately delivered to the user's device, such as a smartphone or tablet.

[0181] As a concrete example, when the emotion engine detects a state of heightened excitement while a user is watching a video, it immediately generates and delivers advertisements for products that appeal to a sense of adventure. In this system, an example of a prompt message to the generation AI model would be: "When the user shows signs of excitement while watching a video, generate advertisements for outdoor equipment that stimulate a sense of adventure."

[0182] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0183] Step 1:

[0184] The server collects activity information from the user's device. This information includes website browsing history and social media activity records. Using this information as input, an initial dataset is generated to infer the user's interests and preferences. At this stage, the information is stored in a database and used in subsequent steps.

[0185] Step 2:

[0186] The server acquires multimodal information by utilizing the device's camera and microphone. Specifically, it takes in audio and image data as input and transfers it to the server in a format that can be analyzed in real time. This data is used to identify the user's facial expressions and voice tone. The output here is a numerical profile of facial expression data and voice characteristics.

[0187] Step 3:

[0188] The server analyzes the user's multimodal information using an emotion engine. It uses the data obtained in the previous step as input and determines the emotional state using machine learning models based on TensorFlow or PyTorch. The output is an evaluation of the emotional state, such as excitement or relaxation, which is then used in the next processing step to generate advertisements.

[0189] Step 4:

[0190] The server runs a generation algorithm based on the obtained emotional state and behavioral information. The input consists of the emotional state evaluation results and data indicating the user's interests. Using Apache Hadoop, the data is processed to generate optimal ad content tailored to the user's real-time emotions. As a result, targeted ads are output.

[0191] Step 5:

[0192] Finally, the server delivers the generated advertisements to the user's device. In this step, the advertisements are inserted in real time, taking into account the user's current content viewing status and device performance. The output is the appropriate advertisement displayed on the user's device screen. The server also tracks the user's reaction when the advertisement is delivered and feeds this feedback back to the server.

[0193] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0194] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0195] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0196] [Second Embodiment]

[0197] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0198] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0199] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0200] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0201] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0202] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0203] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0204] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0205] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0206] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0207] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0208] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0209] This invention is a system for providing advertisements that more accurately reflect users' interests and preferences. In this system, the server primarily performs the functions of analyzing user behavior information and multimodal information, and generates and delivers optimal advertisements using a generation algorithm.

[0210] The server first stores user activity information. This involves collecting a wide range of data, including website browsing history, social media interactions, and digital content usage. This information forms the basis for understanding user interests and preferences.

[0211] Next, the server retrieves multimodal information from the user's device. This includes photos and videos taken on the device, recorded audio, and application usage data. The retrieved information is analyzed by an AI model to contribute to creating a user profile.

[0212] The generation algorithm analyzes this information and generates ads that best match the user's preferences and behavior. For example, if a user has recently been searching for a lot of outdoor-related products, it can create video ads for camping equipment related to that.

[0213] The generated ads are delivered to the user's device at the appropriate time. For example, the server can identify when a user is relaxing and watching a video, and insert a video ad at that moment. Audio ads, on the other hand, are delivered in a way that blends seamlessly into the user's feed while they are listening to music.

[0214] User responses to ads are tracked by the server. Metrics such as click-through rates, ad viewing time, and purchase behavior are analyzed, and this data is used to further improve the generation algorithm. This ensures that delivered ads are continuously better suited to user interests and preferences, maximizing ad effectiveness.

[0215] In this way, the system of the present invention aims to provide personalized advertisements to users and enhance the effectiveness of advertising.

[0216] The following describes the processing flow.

[0217] Step 1:

[0218] The server collects behavioral information from the user's device and online activity. It tracks website browsing history, click data, and social media activity to accumulate data that helps infer the user's interests.

[0219] Step 2:

[0220] The server collects multimodal information from the user's device. It retrieves data captured in the form of images, videos, and audio recordings to gain a deeper understanding of the user's interests and preferences.

[0221] Step 3:

[0222] The server uses an AI model to analyze the collected behavioral and multimodal information. This analysis generates user profiles and provides insights into which advertisements are most effective.

[0223] Step 4:

[0224] The server uses a generation algorithm to create ad content optimized for the user. This process automatically generates ads with creative elements based on the user's preferences.

[0225] Step 5:

[0226] The server determines the optimal timing and format for delivering the generated advertisements to the user's device. For example, it might identify times when users are relaxed and consuming media, and insert advertisements at those times.

[0227] Step 6:

[0228] The server monitors user reactions after ad delivery and collects data such as click-through rates and ad completion rates. This feedback is used to improve the AI ​​model, enhancing the accuracy and effectiveness of future ad generation.

[0229] (Example 1)

[0230] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0231] In today's world, there is a need to effectively deliver advertisements that accurately match users' diverse interests and preferences. However, existing advertising systems struggle to personalize ads with high accuracy based on individual user preferences and behaviors, which limits the effectiveness of advertising.

[0232] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0233] In this invention, the server includes means for collecting user information from multiple data sources, means for acquiring multimodal data from information devices, and means for analyzing the generated data and analyzing user attributes using machine learning algorithms. This enables the generation and delivery of optimized advertisements based on user interests and behavior, thereby improving the personalization and effectiveness of advertisements.

[0234] "User information" refers to data about a user's online and offline behavior, such as website browsing history, social media activity, and use of digital content.

[0235] "Multimodal data" is a general term for data that includes information acquired from multiple different sensors and devices, such as photos, videos, audio, and application usage data.

[0236] "Information devices" refer to electronic devices used by users, such as smartphones, tablets, and personal computers, from which data can be acquired.

[0237] A "machine learning algorithm" is a technique that uses computational models to extract patterns from data and predict future data and behavior.

[0238] "User terminals" refer to electronic devices that users use on a daily basis, and are the devices to which advertisements are delivered.

[0239] "Personalization" refers to the process of adjusting content to suit the individual user's interests and preferences, providing an optimized experience.

[0240] "Ad delivery" refers to the process of sending advertising content to a user's device and making it available for the user to view.

[0241] "Dynamic adjustment" refers to changing the system's operation in real time according to time and circumstances to achieve optimal performance.

[0242] This invention is a system for providing user-optimized advertisements. This system primarily relies on a server, which collects and analyzes information from various data sources and generates and delivers user-specific advertisements, thereby enhancing the effectiveness of the advertisements.

[0243] The server first collects the user's online activity, specifically web browsing history, social media interaction history, and digital content usage. This information is obtained using cookies, APIs, and log files and stored in a database. The obtained data serves as important foundational material for analyzing the user's interests and preferences.

[0244] Next, the device acquires multimodal data using its camera and microphone. This includes photos, videos, and audio files, and this data is processed using image recognition and audio analysis algorithms. The information obtained through this process is sent to a server to build the user's profile.

[0245] The generative AI model analyzes user behavior characteristics based on collected data and generates optimal advertisements to reach users. By using machine learning techniques, it is possible to create advertisements that are tailored to users' purchasing trends and preferences based on previously collected examples. For example, by entering a prompt such as, "Generate advertisements that suggest outdoor products that a user might be interested in, based on their recent search history and multimodal information," highly relevant ad content will be generated.

[0246] The generated ads are delivered by the server at the optimal time, tailored to the user's specific behavior patterns and lifestyle. For example, video ads are delivered at night when the user is relaxing, and audio ads are adjusted to play naturally while the user is listening to their favorite music.

[0247] In this way, the system of the present invention provides advertisements tailored to the needs of each user, thereby achieving maximum advertising effectiveness.

[0248] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0249] Step 1:

[0250] The server collects users' online behavior data. Specifically, it obtains web browsing history, social media activity logs, and digital content usage data using cookies and various APIs. This data forms the basis for creating personalized advertisements and is stored in a database. The input is user behavior data, and the output is the stored dataset.

[0251] Step 2:

[0252] The device acquires multimodal data using its camera and microphone. This process utilizes captured photos, videos, and recorded audio data. This data is processed using image recognition and audio analysis algorithms, resulting in multimodal information that reflects the user's preferences and interests. The input consists of photos, videos, and audio, while the output is the analyzed information.

[0253] Step 3:

[0254] The server integrates all collected data and builds user profiles using a generative AI model. It analyzes the data with machine learning algorithms to predict user interests, preferences, and purchasing tendencies. Inputs are integrated behavioral and multimodal information, and output is a detailed user profile.

[0255] Step 4:

[0256] The server generates the most relevant ads based on the user profile. It prompts the generation AI model with a message such as, "Generate ads suggesting outdoor products that the user might be interested in, based on their recent search history and multimodal information," and generates highly relevant ads. The input is the user profile and the prompt, and the output is the generated ad content.

[0257] Step 5:

[0258] The server delivers the generated ad content to the user's device. The ad delivery time is set to the most effective time based on the user's lifestyle. For example, a video ad might be delivered on Sunday afternoon. The input is the generated ad and its delivery schedule, and the output is the ad played on the device.

[0259] Step 6:

[0260] The server monitors user responses to ads and collects metrics such as click-through rates, viewing time, and purchase behavior. This feedback is used to further improve the generative AI model. The input is user response data, and the output is the improved ad generation algorithm.

[0261] (Application Example 1)

[0262] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0263] In modern society, as advertising personalization advances, there is a growing demand for meaningful and engaging ad delivery for users. However, traditional advertising systems have limitations in their real-time relevance to user interests and behavior, making it difficult to provide information naturally without disrupting the user's visual experience. Therefore, establishing interactive and effective ad delivery methods based on user interests, utilizing visual devices, has become a key challenge.

[0264] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0265] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, means for generating user-optimized advertising content using a generation algorithm, means for delivering the generated advertisements to the user's device, means for tracking the user's response to the advertisements and using this information to improve the generation algorithm, and means for directly overlaying advertisements onto the visual information viewed by the user. This enables the provision of optimal advertising content in real time to the visual device being used by the user, and allows for effective information transmission without compromising the visual experience.

[0266] "User behavior information" is a general term for data related to user actions, and includes behavioral data such as user click patterns, navigation history, and location information.

[0267] "Multimodal information" refers to data collected from multiple different sources, including data acquired through multiple senses such as sound, images, and text.

[0268] A "generative algorithm" is a programmatic procedure that generates output tailored to a specific purpose based on input data, and in particular, it refers to the process of creating advertisements that reflect the user's interests and preferences.

[0269] "User devices" refer to electronic information terminals that users use on a daily basis, and specifically include smartphones, tablets, smart glasses, and the like.

[0270] "User response" refers to the actions and attitudes that users show towards advertisements, and specifically includes the number of clicks, time spent on the page, and the point where users stop scrolling.

[0271] "Methods for directly overlaying advertisements onto visual information" refer to technologies that superimpose the display of advertisements onto objects viewed by users through their visual devices, creating a mechanism that allows advertisements to appear in the user's field of vision in a natural way.

[0272] In this embodiment of the invention, a system is realized in which a server collects user behavior information and multimodal information and generates advertisements using a generation algorithm. The server acquires data from user devices such as smart glasses and creates a user profile using AI software such as TensorFlow. Based on this, it generates advertisements tailored to the user's interests and displays them as an overlay on the visual device. This process makes it possible to provide relevant advertisements in real time in relation to the surrounding environment that the user is viewing.

[0273] User responses to advertisements are collected through sensors and interfaces on the user's device. This allows for the measurement of the advertisement's effectiveness, and the data is used to improve the generation algorithm. For example, if it is detected that a user has been lingering their gaze on an overlaid advertisement for an extended period, the advertisement is evaluated as visually effective.

[0274] For example, when a user is walking through a shopping mall, a promotional advertisement for that store can be displayed on the reflection of the store in their glasses. In this case, the advertisement is placed in the user's field of vision in a natural way, and the information is presented in a manner that is likely to capture the user's interest.

[0275] An example of a prompt message is: "I want to design an application that overlays advertisements for stores that a user sees in their line of sight onto their glasses as they walk around. Please tell me the necessary steps and technologies to use to analyze the user's profile with AI and select the appropriate advertisement in real time."

[0276] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0277] Step 1:

[0278] Activity information is collected from the user's device. The device records the user's location, eye-tracking data, and device usage in real time. This data serves as foundational data for understanding the user's activities and interests. User activity data is obtained as input, and pre-processed data for analysis is provided as output.

[0279] Step 2:

[0280] The server retrieves multimodal information from the user's terminal. This multimodal information includes voice commands, image data from the camera, and application operation logs. The input is the user's multimodal data, and the output is data converted into a format that can be analyzed by the AI ​​model.

[0281] Step 3:

[0282] The server analyzes the collected operation information and multimodal information. Using an AI generation algorithm, it processes this data to generate a profile based on the user's interests and behavior patterns. The input is the analysis data obtained in the previous step, and the output is the user's interest profile.

[0283] Step 4:

[0284] The server generates an optimal advertisement based on the profile. Using a generation AI model, it automatically creates an advertisement tailored to the user's interests. Specifically, it selects advertising materials related to the products or services that the user has shown interest in and generates a customized advertisement in the system. The input is the user profile, and the output is the generated advertisement content.

[0285] Step 5:

[0286] The server distributes the generated advertisement to the user's visual device. The terminal overlays and displays the advertisement on the display of the smart glasses, seamlessly integrating it into the real environment that the user is viewing. The input is the advertisement content, and the output is the display on the smart glasses.

[0287] Step 6:

[0288] The server tracks the user's reaction to the advertisement through the visual device. The terminal uses sensors to record the user's eye line data and interface operations, analyzes them to evaluate the effectiveness of the advertisement. The input is the user's response data, and the output is the evaluation result of the advertisement performance.

[0289] Step 7:

[0290] The server uses user response data to improve the generation algorithm. The collected data is incorporated into the algorithm as feedback, further improving it to better match user preferences in future ad generation. The input is an evaluation of the ad response, and the output is the improved ad generation algorithm.

[0291] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0292] This invention is an advertising system that combines user behavior information, multimodal information, and an emotion engine that recognizes user emotions. The purpose of this system is to generate and deliver personalized advertisements that are tailored to the user's interests and emotions.

[0293] The server first obtains activity information from the user's device. This includes website browsing history and social media activity, which is used to infer the user's interests and preferences. Next, the server collects multimodal information from the device. This allows the server to create a more detailed profile of the user through the analysis of text data, image data, and audio data.

[0294] Furthermore, in this invention, the server uses an emotion engine to analyze the user's emotions in real time. This emotion engine analyzes the user's facial expressions, tone of voice, and input data to recognize the user's emotional state. This information is a crucial element for more effectively personalizing advertisements.

[0295] For example, if the emotion engine detects a high level of excitement while a user is watching a video, the server can generate an energetic ad that matches that state. Conversely, if it detects a relaxed state, it will generate an ad with a more subdued tone.

[0296] The generation algorithm integrates all of this data to produce optimal ad content. Because this ad is tailored based on the user's interests and emotional state, it can achieve a more effective appeal.

[0297] The generated advertisements are delivered to the user's device at a time and format dynamically adjusted by the server. This process ensures optimal ad delivery based on the user's usage patterns and the type of content they are viewing.

[0298] Ultimately, the server tracks user responses to ads, analyzing click-through rates, viewing time, emotional changes, and more. This feedback can be used to improve future ad generation and enhance the overall effectiveness of advertising campaigns.

[0299] The following describes the processing flow.

[0300] Step 1:

[0301] The server collects activity information from the user's device. This includes web history, social media activity, and purchase history, and uses this data to understand the user's interests and preferences.

[0302] Step 2:

[0303] The server retrieves multimodal information from the user's device. This refers to data such as images, videos, and audio, and is used to understand the user's content consumption patterns.

[0304] Step 3:

[0305] The server uses an emotion engine to analyze the user's current emotional state. It analyzes data obtained from the device's camera and microphone, and evaluates the user's emotions through facial expressions and tone of voice.

[0306] Step 4:

[0307] The server integrates and analyzes the user's operation information, multimodal information, and emotional state to generate optimal advertisements. At this time, the generation algorithm personalizes the content according to the user's characteristics.

[0308] Step 5:

[0309] Based on the user's usage situation, the server determines the optimal timing and format for delivering the generated advertisements. When the user is watching a video or targeting quiet time periods, the advertisements are displayed.

[0310] Step 6:

[0311] The server monitors the user's reaction to the delivered advertisements. This includes activities such as tracking the click-through rate and view completion rate of the advertisements, as well as changes in the user's emotions.

[0312] Step 7:

[0313] Based on the collected reaction data, the server establishes a feedback loop for the emotion engine and the generation algorithm. This enables continuous improvement to enhance the accuracy of future advertisement generation and delivery.

[0314] (Example 2)

[0315] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0316] In a conventional advertising system, although it is possible to generate personalized advertisements based on the user's interests and concerns, the generation of advertisements considering the user's real-time emotional state is insufficient. Also, the optimization of dynamic advertisement delivery according to the user's usage situation is not sufficient. As a result, it is difficult to arouse the user's interest in the advertisement, and an improvement in the advertisement effect is required.

[0317] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0318] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, and means for performing sentiment analysis. This makes it possible to generate advertisements optimized based on the user's emotional state and interests and deliver them to the user's terminal in real time. This is expected to improve the effectiveness of the advertisements.

[0319] "Means for collecting user behavior information" refers to methods for obtaining data about user behavior, such as a user's web browsing history and social media activity.

[0320] "Means for acquiring multimodal information" refers to technologies that collect information in multiple media formats, such as text, images, and audio data, from a user's device.

[0321] "Methods for performing emotion analysis" refer to the process of analyzing a user's facial expressions, tone of voice, and input data to identify the user's emotional state.

[0322] "Using a generation algorithm" means using a computational method to generate user-optimized advertisements based on collected data.

[0323] "Means of delivering to user devices" refers to technologies that transmit generated advertisements to the user's terminal at the appropriate time and in the appropriate format.

[0324] "Means of tracking user responses" refer to methods for recording user behavior in response to advertisements, such as click-through rates and viewing time, in order to evaluate the effectiveness of the advertisements.

[0325] This invention is a system that generates and delivers advertisements that reflect the user's interests and emotions in real time. The server collects various information from the user's terminal and uses this information to personalize advertisements. Specifically, it obtains user behavior information, such as website browsing history and social media activity, and analyzes this information to infer the user's interests.

[0326] Furthermore, the server collects multimodal information. This information includes text, images, and audio, and is obtained from the user's device using the camera and microphone. By analyzing this data, a more detailed profile of the user can be created.

[0327] In emotion analysis, the server analyzes the user's emotional state in real time. Using an emotion engine, the AI ​​model analyzes the user's facial expressions and tone of voice to identify emotions such as joy, excitement, and relaxation. This makes it possible to adjust the content of advertisements to best suit the user's current emotions.

[0328] Ad generation utilizes a generative AI model. This model comprehensively processes collected data to automatically create ads that match the user's interests and emotions. The generated ads are then delivered to the user's device by a server at the optimal time. This delivery is dynamically adjusted according to the user's usage of the content.

[0329] For example, if a user is searching for a new recipe, the server can generate advertisements for cooking utensils that match their interests and provide the user with enjoyment and value by using cooking-related prompts as examples. In this way, advertisements based on user interests and emotions contribute significantly to improving advertising effectiveness.

[0330] An example of a prompt message would be, "The user has been detected as highly aroused while watching a video. Generate an ad for a new product that might interest them," which enhances ad personalization.

[0331] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0332] Step 1:

[0333] The user's device sends activity information to the server. Specifically, it transfers website browsing history and social media activity data from the device to the server. The server uses this data to build an initial dataset to infer the user's interests. The user's behavioral history data is used as input, and the output generates an initial set of tags indicating the user's interests.

[0334] Step 2:

[0335] The server collects multimodal information from the user's device. This information includes text, images, and audio data, captured, for example, via the camera and microphone of a smartphone. The server analyzes this data to generate a detailed user profile. The input is the user's text, images, and audio data, and the output is an enhanced version of the user profile.

[0336] Step 3:

[0337] The server initiates emotion analysis, identifying the user's emotional state in real time. It uses an emotion engine to analyze facial expressions and voice tone. The results of this analysis are output as data indicating the user's current emotional state. The input is data of the user's facial expressions and voice, and the output is a tag indicating the user's emotional state at that moment.

[0338] Step 4:

[0339] The generative AI model generates personalized advertisements based on previously obtained data (user interest tags, detailed profiles, and emotional states). The input is the user's overall profile information, and the output is advertising content optimized based on that profile. Specifically, the ad's colors and word choices are tailored to the user's emotions.

[0340] Step 5:

[0341] The server sends the generated advertisements to the user's device. The timing and format of delivery are dynamically adjusted, taking into account the user's content consumption habits and emotional state. The input is the generated ad data, and the output is the optimized advertisement displayed on the user's screen. The device presents the received advertisement to the user in the specified manner.

[0342] Step 6:

[0343] The server tracks user responses and analyzes the effectiveness of advertisements. It collects data on user clicks and viewing time to evaluate the success rate of advertisements. The input is user behavior data regarding advertisements, and the output is an evaluation of advertisement effectiveness that is reflected in the next advertisement generation. This allows the server to improve its generation algorithm.

[0344] (Application Example 2)

[0345] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0346] Traditional advertising delivery systems generate ads based only on user behavior information and limited data, resulting in ads that do not adequately address user interests. Furthermore, they fail to consider the user's real-time emotional state, hindering the maximization of ad effectiveness and user engagement.

[0347] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0348] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, means for using an emotion engine to recognize the user's emotional state, and means for using a generation algorithm to generate advertising content optimized for the user based on the emotional state. This enables the delivery of advertisements optimized based on the user's real-time emotions and interests.

[0349] "User activity information" refers to data about a user's behavior and actions on the internet, including website browsing history and social media activity history.

[0350] "Multimodal information" refers to information that encompasses multiple different types of data, such as text, images, and audio, and is used in combination with these other data.

[0351] An "emotion engine" is an algorithm or software that analyzes a user's facial expressions, voice tone, and input data to identify their emotional state.

[0352] A "generative algorithm" is a computational method for creating optimized ad content based on the user's interests and emotional state.

[0353] "User devices" refer to devices used to display advertisements and information, and include smartphones and tablet devices.

[0354] "Real-time" is a term that indicates that information is processed or delivered at the same time as, or very close to, the time it is generated.

[0355] The system of this invention combines the collection of user behavior information, acquisition of multimodal information, sentiment analysis using an emotion engine, ad optimization using a generation algorithm, and ad delivery to provide users with optimized ads in real time.

[0356] The server collects behavioral information from the user's device. Specifically, it receives web browsing history and social media activity information, and uses this data to infer interests and preferences. Next, it acquires multimodal information from devices equipped with cameras and microphones. This includes the collection of real-time audio and image data. The server analyzes this data and recognizes emotional states using an emotion engine. This emotion engine is built using machine learning libraries such as TensorFlow and PyTorch.

[0357] Based on the analyzed information, the server uses a generation algorithm to generate advertisements optimized for the user's emotional state. This generation algorithm utilizes Apache Hadoop, which can process large amounts of data at high speed, enabling the rapid generation of content that reflects the user's real-time emotional state. The generated advertisements are immediately delivered to the user's device, such as a smartphone or tablet.

[0358] As a concrete example, when the emotion engine detects a state of heightened excitement while a user is watching a video, it immediately generates and delivers advertisements for products that appeal to a sense of adventure. In this system, an example of a prompt message to the generation AI model would be: "When the user shows signs of excitement while watching a video, generate advertisements for outdoor equipment that stimulate a sense of adventure."

[0359] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0360] Step 1:

[0361] The server collects activity information from the user's device. This information includes website browsing history and social media activity records. Using this information as input, an initial dataset is generated to infer the user's interests and preferences. At this stage, the information is stored in a database and used in subsequent steps.

[0362] Step 2:

[0363] The server acquires multimodal information by utilizing the device's camera and microphone. Specifically, it takes in audio and image data as input and transfers it to the server in a format that can be analyzed in real time. This data is used to identify the user's facial expressions and voice tone. The output here is a numerical profile of facial expression data and voice characteristics.

[0364] Step 3:

[0365] The server analyzes the user's multimodal information using an emotion engine. It uses the data obtained in the previous step as input and determines the emotional state using machine learning models based on TensorFlow or PyTorch. The output is an evaluation of the emotional state, such as excitement or relaxation, which is then used in the next processing step to generate advertisements.

[0366] Step 4:

[0367] The server runs a generation algorithm based on the obtained emotional state and behavioral information. The input consists of the emotional state evaluation results and data indicating the user's interests. Using Apache Hadoop, the data is processed to generate optimal ad content tailored to the user's real-time emotions. As a result, targeted ads are output.

[0368] Step 5:

[0369] Finally, the server delivers the generated advertisements to the user's device. In this step, the advertisements are inserted in real time, taking into account the user's current content viewing status and device performance. The output is the appropriate advertisement displayed on the user's device screen. The server also tracks the user's reaction when the advertisement is delivered and feeds this feedback back to the server.

[0370] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0371] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0372] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0373] [Third Embodiment]

[0374] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0375] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0376] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0377] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0378] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0379] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0380] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0381] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0382] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0383] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0384] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0385] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0386] This invention is a system for providing advertisements that more accurately reflect users' interests and preferences. In this system, the server primarily performs the functions of analyzing user behavior information and multimodal information, and generates and delivers optimal advertisements using a generation algorithm.

[0387] The server first stores user activity information. This involves collecting a wide range of data, including website browsing history, social media interactions, and digital content usage. This information forms the basis for understanding user interests and preferences.

[0388] Next, the server retrieves multimodal information from the user's device. This includes photos and videos taken on the device, recorded audio, and application usage data. The retrieved information is analyzed by an AI model to contribute to creating a user profile.

[0389] The generation algorithm analyzes this information and generates ads that best match the user's preferences and behavior. For example, if a user has recently been searching for a lot of outdoor-related products, it can create video ads for camping equipment related to that.

[0390] The generated ads are delivered to the user's device at the appropriate time. For example, the server can identify when a user is relaxing and watching a video, and insert a video ad at that moment. Audio ads, on the other hand, are delivered in a way that blends seamlessly into the user's feed while they are listening to music.

[0391] User responses to ads are tracked by the server. Metrics such as click-through rates, ad viewing time, and purchase behavior are analyzed, and this data is used to further improve the generation algorithm. This ensures that delivered ads are continuously better suited to user interests and preferences, maximizing ad effectiveness.

[0392] In this way, the system of the present invention aims to provide personalized advertisements to users and enhance the effectiveness of advertising.

[0393] The following describes the processing flow.

[0394] Step 1:

[0395] The server collects behavioral information from the user's device and online activity. It tracks website browsing history, click data, and social media activity to accumulate data that helps infer the user's interests.

[0396] Step 2:

[0397] The server collects multimodal information from the user's device. It retrieves data captured in the form of images, videos, and audio recordings to gain a deeper understanding of the user's interests and preferences.

[0398] Step 3:

[0399] The server uses an AI model to analyze the collected behavioral and multimodal information. This analysis generates user profiles and provides insights into which advertisements are most effective.

[0400] Step 4:

[0401] The server uses a generation algorithm to create ad content optimized for the user. This process automatically generates ads with creative elements based on the user's preferences.

[0402] Step 5:

[0403] The server determines the optimal timing and format for delivering the generated advertisements to the user's device. For example, it might identify times when users are relaxed and consuming media, and insert advertisements at those times.

[0404] Step 6:

[0405] The server monitors user reactions after ad delivery and collects data such as click-through rates and ad completion rates. This feedback is used to improve the AI ​​model, enhancing the accuracy and effectiveness of future ad generation.

[0406] (Example 1)

[0407] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0408] In today's world, there is a need to effectively deliver advertisements that accurately match users' diverse interests and preferences. However, existing advertising systems struggle to personalize ads with high accuracy based on individual user preferences and behaviors, which limits the effectiveness of advertising.

[0409] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0410] In this invention, the server includes means for collecting user information from multiple data sources, means for acquiring multimodal data from information devices, and means for analyzing the generated data and analyzing user attributes using machine learning algorithms. This enables the generation and delivery of optimized advertisements based on user interests and behavior, thereby improving the personalization and effectiveness of advertisements.

[0411] "User information" refers to data about a user's online and offline behavior, such as website browsing history, social media activity, and use of digital content.

[0412] "Multimodal data" is a general term for data that includes information acquired from multiple different sensors and devices, such as photos, videos, audio, and application usage data.

[0413] "Information devices" refer to electronic devices used by users, such as smartphones, tablets, and personal computers, from which data can be acquired.

[0414] A "machine learning algorithm" is a technique that uses computational models to extract patterns from data and predict future data and behavior.

[0415] "User terminals" refer to electronic devices that users use on a daily basis, and are the devices to which advertisements are delivered.

[0416] "Personalization" refers to the process of adjusting content to suit the individual user's interests and preferences, providing an optimized experience.

[0417] "Ad delivery" refers to the process of sending advertising content to a user's device and making it available for the user to view.

[0418] "Dynamic adjustment" refers to changing the system's operation in real time according to time and circumstances to achieve optimal performance.

[0419] This invention is a system for providing user-optimized advertisements. This system primarily relies on a server, which collects and analyzes information from various data sources and generates and delivers user-specific advertisements, thereby enhancing the effectiveness of the advertisements.

[0420] The server first collects the user's online activity, specifically web browsing history, social media interaction history, and digital content usage. This information is obtained using cookies, APIs, and log files and stored in a database. The obtained data serves as important foundational material for analyzing the user's interests and preferences.

[0421] Next, the device acquires multimodal data using its camera and microphone. This includes photos, videos, and audio files, and this data is processed using image recognition and audio analysis algorithms. The information obtained through this process is sent to a server to build the user's profile.

[0422] The generative AI model analyzes user behavior characteristics based on collected data and generates optimal advertisements to reach users. By using machine learning techniques, it is possible to create advertisements that are tailored to users' purchasing trends and preferences based on previously collected examples. For example, by entering a prompt such as, "Generate advertisements that suggest outdoor products that a user might be interested in, based on their recent search history and multimodal information," highly relevant ad content will be generated.

[0423] The generated ads are delivered by the server at the optimal time, tailored to the user's specific behavior patterns and lifestyle. For example, video ads are delivered at night when the user is relaxing, and audio ads are adjusted to play naturally while the user is listening to their favorite music.

[0424] In this way, the system of the present invention provides advertisements tailored to the needs of each user, thereby achieving maximum advertising effectiveness.

[0425] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0426] Step 1:

[0427] The server collects users' online behavior data. Specifically, it obtains web browsing history, social media activity logs, and digital content usage data using cookies and various APIs. This data forms the basis for creating personalized advertisements and is stored in a database. The input is user behavior data, and the output is the stored dataset.

[0428] Step 2:

[0429] The device acquires multimodal data using its camera and microphone. This process utilizes captured photos, videos, and recorded audio data. This data is processed using image recognition and audio analysis algorithms, resulting in multimodal information that reflects the user's preferences and interests. The input consists of photos, videos, and audio, while the output is the analyzed information.

[0430] Step 3:

[0431] The server integrates all collected data and builds user profiles using a generative AI model. It analyzes the data with machine learning algorithms to predict user interests, preferences, and purchasing tendencies. Inputs are integrated behavioral and multimodal information, and output is a detailed user profile.

[0432] Step 4:

[0433] The server generates the most relevant ads based on the user profile. It prompts the generation AI model with a message such as, "Generate ads suggesting outdoor products that the user might be interested in, based on their recent search history and multimodal information," and generates highly relevant ads. The input is the user profile and the prompt, and the output is the generated ad content.

[0434] Step 5:

[0435] The server delivers the generated ad content to the user's device. The ad delivery time is set to the most effective time based on the user's lifestyle. For example, a video ad might be delivered on Sunday afternoon. The input is the generated ad and its delivery schedule, and the output is the ad played on the device.

[0436] Step 6:

[0437] The server monitors user responses to ads and collects metrics such as click-through rates, viewing time, and purchase behavior. This feedback is used to further improve the generative AI model. The input is user response data, and the output is the improved ad generation algorithm.

[0438] (Application Example 1)

[0439] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0440] In modern society, as advertising personalization advances, there is a growing demand for meaningful and engaging ad delivery for users. However, traditional advertising systems have limitations in their real-time relevance to user interests and behavior, making it difficult to provide information naturally without disrupting the user's visual experience. Therefore, establishing interactive and effective ad delivery methods based on user interests, utilizing visual devices, has become a key challenge.

[0441] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0442] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, means for generating user-optimized advertising content using a generation algorithm, means for delivering the generated advertisements to the user's device, means for tracking the user's response to the advertisements and using this information to improve the generation algorithm, and means for directly overlaying advertisements onto the visual information viewed by the user. This enables the provision of optimal advertising content in real time to the visual device being used by the user, and allows for effective information transmission without compromising the visual experience.

[0443] "User behavior information" is a general term for data related to user actions, and includes behavioral data such as user click patterns, navigation history, and location information.

[0444] "Multimodal information" refers to data collected from multiple different sources, including data acquired through multiple senses such as sound, images, and text.

[0445] A "generative algorithm" is a programmatic procedure that generates output tailored to a specific purpose based on input data, and in particular, it refers to the process of creating advertisements that reflect the user's interests and preferences.

[0446] "User devices" refer to electronic information terminals that users use on a daily basis, and specifically include smartphones, tablets, smart glasses, and the like.

[0447] "User response" refers to the actions and attitudes that users show towards advertisements, and specifically includes the number of clicks, time spent on the page, and the point where users stop scrolling.

[0448] "Methods for directly overlaying advertisements onto visual information" refer to technologies that superimpose the display of advertisements onto objects viewed by users through their visual devices, creating a mechanism that allows advertisements to appear in the user's field of vision in a natural way.

[0449] In this embodiment of the invention, a system is realized in which a server collects user behavior information and multimodal information and generates advertisements using a generation algorithm. The server acquires data from user devices such as smart glasses and creates a user profile using AI software such as TensorFlow. Based on this, it generates advertisements tailored to the user's interests and displays them as an overlay on the visual device. This process makes it possible to provide relevant advertisements in real time in relation to the surrounding environment that the user is viewing.

[0450] User responses to advertisements are collected through sensors and interfaces on the user's device. This allows for the measurement of the advertisement's effectiveness, and the data is used to improve the generation algorithm. For example, if it is detected that a user has been lingering their gaze on an overlaid advertisement for an extended period, the advertisement is evaluated as visually effective.

[0451] For example, when a user is walking through a shopping mall, a promotional advertisement for that store can be displayed on the reflection of the store in their glasses. In this case, the advertisement is placed in the user's field of vision in a natural way, and the information is presented in a manner that is likely to capture the user's interest.

[0452] An example of a prompt message is: "I want to design an application that overlays advertisements for stores that a user sees in their line of sight onto their glasses as they walk around. Please tell me the necessary steps and technologies to use to analyze the user's profile with AI and select the appropriate advertisement in real time."

[0453] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0454] Step 1:

[0455] Activity information is collected from the user's device. The device records the user's location, eye-tracking data, and device usage in real time. This data serves as foundational data for understanding the user's activities and interests. User activity data is obtained as input, and pre-processed data for analysis is provided as output.

[0456] Step 2:

[0457] The server retrieves multimodal information from the user's terminal. This multimodal information includes voice commands, image data from the camera, and application operation logs. The input is the user's multimodal data, and the output is data converted into a format that can be analyzed by the AI ​​model.

[0458] Step 3:

[0459] The server analyzes the collected behavioral and multimodal information. Using an AI generation algorithm, it processes this data and generates a profile based on the user's interests and behavioral patterns. The input is the analysis data obtained in the previous step, and the output is the user's interest profile.

[0460] Step 4:

[0461] The server generates optimal advertisements based on the user's profile. Using a generative AI model, it automatically creates advertisements tailored to the user's interests. Specifically, it selects advertising materials related to products and services the user has shown interest in, and generates customized advertisements through the system. The input is the user profile, and the output is the generated advertising content.

[0462] Step 5:

[0463] The server delivers the generated advertisements to the user's visual device. The device overlays the advertisements on the smart glasses' display, seamlessly blending them into the user's real-world environment. The input is the advertisement content, and the output is its display on the smart glasses.

[0464] Step 6:

[0465] The server tracks user responses to advertisements through visual devices. The terminal uses sensors to record user eye-tracking data and interface interactions, which are then analyzed to evaluate the effectiveness of the advertisements. The input is user response data, and the output is the evaluation result of the advertisement performance.

[0466] Step 7:

[0467] The server uses user response data to improve the generation algorithm. The collected data is incorporated into the algorithm as feedback, further improving it to better match user preferences in future ad generation. The input is an evaluation of the ad response, and the output is the improved ad generation algorithm.

[0468] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0469] This invention is an advertising system that combines user behavior information, multimodal information, and an emotion engine that recognizes user emotions. The purpose of this system is to generate and deliver personalized advertisements that are tailored to the user's interests and emotions.

[0470] The server first obtains activity information from the user's device. This includes website browsing history and social media activity, which is used to infer the user's interests and preferences. Next, the server collects multimodal information from the device. This allows the server to create a more detailed profile of the user through the analysis of text data, image data, and audio data.

[0471] Furthermore, in this invention, the server uses an emotion engine to analyze the user's emotions in real time. This emotion engine analyzes the user's facial expressions, tone of voice, and input data to recognize the user's emotional state. This information is a crucial element for more effectively personalizing advertisements.

[0472] For example, if the emotion engine detects a high level of excitement while a user is watching a video, the server can generate an energetic ad that matches that state. Conversely, if it detects a relaxed state, it will generate an ad with a more subdued tone.

[0473] The generation algorithm integrates all of this data to produce optimal ad content. Because this ad is tailored based on the user's interests and emotional state, it can achieve a more effective appeal.

[0474] The generated advertisements are delivered to the user's device at a time and format dynamically adjusted by the server. This process ensures optimal ad delivery based on the user's usage patterns and the type of content they are viewing.

[0475] Ultimately, the server tracks user responses to ads, analyzing click-through rates, viewing time, emotional changes, and more. This feedback can be used to improve future ad generation and enhance the overall effectiveness of advertising campaigns.

[0476] The following describes the processing flow.

[0477] Step 1:

[0478] The server collects activity information from the user's device. This includes web history, social media activity, and purchase history, and uses this data to understand the user's interests and preferences.

[0479] Step 2:

[0480] The server retrieves multimodal information from the user's device. This refers to data such as images, videos, and audio, and is used to understand the user's content consumption patterns.

[0481] Step 3:

[0482] The server uses an emotion engine to analyze the user's current emotional state. It analyzes data obtained from the device's camera and microphone, and evaluates the user's emotions through facial expressions and tone of voice.

[0483] Step 4:

[0484] The server integrates and analyzes user behavior information, multimodal information, and emotional state to generate optimal advertisements. In this process, the generation algorithm personalizes content according to the user's characteristics.

[0485] Step 5:

[0486] The server determines the optimal timing and format for delivering generated ads based on the user's usage patterns. Ads are displayed when the user is watching videos or during quiet periods.

[0487] Step 6:

[0488] The server monitors user responses to delivered advertisements. This includes tracking ad click-through rates, view completion rates, and changes in user sentiment.

[0489] Step 7:

[0490] The server establishes a feedback loop between the emotion engine and the generation algorithm based on the collected response data. This enables continuous improvement to enhance the accuracy of future ad generation and delivery.

[0491] (Example 2)

[0492] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0493] Traditional advertising systems, while capable of generating personalized ads based on user interests, were insufficient in considering users' real-time emotional states. Furthermore, they lacked adequate optimization of dynamic ad delivery based on user behavior. This made it difficult to capture user interest in ads, highlighting the need for improved advertising effectiveness.

[0494] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0495] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, and means for performing sentiment analysis. This makes it possible to generate advertisements optimized based on the user's emotional state and interests and deliver them to the user's terminal in real time. This is expected to improve the effectiveness of the advertisements.

[0496] "Means for collecting user behavior information" refers to methods for obtaining data about user behavior, such as a user's web browsing history and social media activity.

[0497] "Means for acquiring multimodal information" refers to technologies that collect information in multiple media formats, such as text, images, and audio data, from a user's device.

[0498] "Methods for performing emotion analysis" refer to the process of analyzing a user's facial expressions, tone of voice, and input data to identify the user's emotional state.

[0499] "Using a generation algorithm" means using a computational method to generate user-optimized advertisements based on collected data.

[0500] "Means of delivering to user devices" refers to technologies that transmit generated advertisements to the user's terminal at the appropriate time and in the appropriate format.

[0501] "Means of tracking user responses" refer to methods for recording user behavior in response to advertisements, such as click-through rates and viewing time, in order to evaluate the effectiveness of the advertisements.

[0502] This invention is a system that generates and delivers advertisements that reflect the user's interests and emotions in real time. The server collects various information from the user's terminal and uses this information to personalize advertisements. Specifically, it obtains user behavior information, such as website browsing history and social media activity, and analyzes this information to infer the user's interests.

[0503] Furthermore, the server collects multimodal information. This information includes text, images, and audio, and is obtained from the user's device using the camera and microphone. By analyzing this data, a more detailed profile of the user can be created.

[0504] In emotion analysis, the server analyzes the user's emotional state in real time. Using an emotion engine, the AI ​​model analyzes the user's facial expressions and tone of voice to identify emotions such as joy, excitement, and relaxation. This makes it possible to adjust the content of advertisements to best suit the user's current emotions.

[0505] Ad generation utilizes a generative AI model. This model comprehensively processes collected data to automatically create ads that match the user's interests and emotions. The generated ads are then delivered to the user's device by a server at the optimal time. This delivery is dynamically adjusted according to the user's usage of the content.

[0506] For example, if a user is searching for a new recipe, the server can generate advertisements for cooking utensils that match their interests and provide the user with enjoyment and value by using cooking-related prompts as examples. In this way, advertisements based on user interests and emotions contribute significantly to improving advertising effectiveness.

[0507] An example of a prompt message would be, "The user has been detected as highly aroused while watching a video. Generate an ad for a new product that might interest them," which enhances ad personalization.

[0508] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0509] Step 1:

[0510] The user's device sends activity information to the server. Specifically, it transfers website browsing history and social media activity data from the device to the server. The server uses this data to build an initial dataset to infer the user's interests. The user's behavioral history data is used as input, and the output generates an initial set of tags indicating the user's interests.

[0511] Step 2:

[0512] The server collects multimodal information from the user's device. This information includes text, images, and audio data, captured, for example, via the camera and microphone of a smartphone. The server analyzes this data to generate a detailed user profile. The input is the user's text, images, and audio data, and the output is an enhanced version of the user profile.

[0513] Step 3:

[0514] The server initiates emotion analysis, identifying the user's emotional state in real time. It uses an emotion engine to analyze facial expressions and voice tone. The results of this analysis are output as data indicating the user's current emotional state. The input is data of the user's facial expressions and voice, and the output is a tag indicating the user's emotional state at that moment.

[0515] Step 4:

[0516] The generative AI model generates personalized advertisements based on previously obtained data (user interest tags, detailed profiles, and emotional states). The input is the user's overall profile information, and the output is advertising content optimized based on that profile. Specifically, the ad's colors and word choices are tailored to the user's emotions.

[0517] Step 5:

[0518] The server sends the generated advertisements to the user's device. The timing and format of delivery are dynamically adjusted, taking into account the user's content consumption habits and emotional state. The input is the generated ad data, and the output is the optimized advertisement displayed on the user's screen. The device presents the received advertisement to the user in the specified manner.

[0519] Step 6:

[0520] The server tracks user responses and analyzes the effectiveness of advertisements. It collects data on user clicks and viewing time to evaluate the success rate of advertisements. The input is user behavior data regarding advertisements, and the output is an evaluation of advertisement effectiveness that is reflected in the next advertisement generation. This allows the server to improve its generation algorithm.

[0521] (Application Example 2)

[0522] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0523] Traditional advertising delivery systems generate ads based only on user behavior information and limited data, resulting in ads that do not adequately address user interests. Furthermore, they fail to consider the user's real-time emotional state, hindering the maximization of ad effectiveness and user engagement.

[0524] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0525] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, means for using an emotion engine to recognize the user's emotional state, and means for using a generation algorithm to generate advertising content optimized for the user based on the emotional state. This enables the delivery of advertisements optimized based on the user's real-time emotions and interests.

[0526] "User activity information" refers to data about a user's behavior and actions on the internet, including website browsing history and social media activity history.

[0527] "Multimodal information" refers to information that encompasses multiple different types of data, such as text, images, and audio, and is used in combination with these other data.

[0528] An "emotion engine" is an algorithm or software that analyzes a user's facial expressions, voice tone, and input data to identify their emotional state.

[0529] A "generative algorithm" is a computational method for creating optimized ad content based on the user's interests and emotional state.

[0530] "User devices" refer to devices used to display advertisements and information, and include smartphones and tablet devices.

[0531] "Real-time" is a term that indicates that information is processed or delivered at the same time as, or very close to, the time it is generated.

[0532] The system of this invention combines the collection of user behavior information, acquisition of multimodal information, sentiment analysis using an emotion engine, ad optimization using a generation algorithm, and ad delivery to provide users with optimized ads in real time.

[0533] The server collects behavioral information from the user's device. Specifically, it receives web browsing history and social media activity information, and uses this data to infer interests and preferences. Next, it acquires multimodal information from devices equipped with cameras and microphones. This includes the collection of real-time audio and image data. The server analyzes this data and recognizes emotional states using an emotion engine. This emotion engine is built using machine learning libraries such as TensorFlow and PyTorch.

[0534] Based on the analyzed information, the server uses a generation algorithm to generate advertisements optimized for the user's emotional state. This generation algorithm utilizes Apache Hadoop, which can process large amounts of data at high speed, enabling the rapid generation of content that reflects the user's real-time emotional state. The generated advertisements are immediately delivered to the user's device, such as a smartphone or tablet.

[0535] As a concrete example, when the emotion engine detects a state of heightened excitement while a user is watching a video, it immediately generates and delivers advertisements for products that appeal to a sense of adventure. In this system, an example of a prompt message to the generation AI model would be: "When the user shows signs of excitement while watching a video, generate advertisements for outdoor equipment that stimulate a sense of adventure."

[0536] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0537] Step 1:

[0538] The server collects activity information from the user's device. This information includes website browsing history and social media activity records. Using this information as input, an initial dataset is generated to infer the user's interests and preferences. At this stage, the information is stored in a database and used in subsequent steps.

[0539] Step 2:

[0540] The server acquires multimodal information by utilizing the device's camera and microphone. Specifically, it takes in audio and image data as input and transfers it to the server in a format that can be analyzed in real time. This data is used to identify the user's facial expressions and voice tone. The output here is a numerical profile of facial expression data and voice characteristics.

[0541] Step 3:

[0542] The server analyzes the user's multimodal information using an emotion engine. It uses the data obtained in the previous step as input and determines the emotional state using machine learning models based on TensorFlow or PyTorch. The output is an evaluation of the emotional state, such as excitement or relaxation, which is then used in the next processing step to generate advertisements.

[0543] Step 4:

[0544] The server runs a generation algorithm based on the obtained emotional state and behavioral information. The input consists of the emotional state evaluation results and data indicating the user's interests. Using Apache Hadoop, the data is processed to generate optimal ad content tailored to the user's real-time emotions. As a result, targeted ads are output.

[0545] Step 5:

[0546] Finally, the server delivers the generated advertisements to the user's device. In this step, the advertisements are inserted in real time, taking into account the user's current content viewing status and device performance. The output is the appropriate advertisement displayed on the user's device screen. The server also tracks the user's reaction when the advertisement is delivered and feeds this feedback back to the server.

[0547] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0548] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0549] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0550] [Fourth Embodiment]

[0551] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0552] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0553] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0554] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0555] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0556] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0557] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0558] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0559] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0560] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0561] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0562] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0563] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0564] This invention is a system for providing advertisements that more accurately reflect users' interests and preferences. In this system, the server primarily performs the functions of analyzing user behavior information and multimodal information, and generates and delivers optimal advertisements using a generation algorithm.

[0565] The server first stores user activity information. This involves collecting a wide range of data, including website browsing history, social media interactions, and digital content usage. This information forms the basis for understanding user interests and preferences.

[0566] Next, the server retrieves multimodal information from the user's device. This includes photos and videos taken on the device, recorded audio, and application usage data. The retrieved information is analyzed by an AI model to contribute to creating a user profile.

[0567] The generation algorithm analyzes this information and generates ads that best match the user's preferences and behavior. For example, if a user has recently been searching for a lot of outdoor-related products, it can create video ads for camping equipment related to that.

[0568] The generated ads are delivered to the user's device at the appropriate time. For example, the server can identify when a user is relaxing and watching a video, and insert a video ad at that moment. Audio ads, on the other hand, are delivered in a way that blends seamlessly into the user's feed while they are listening to music.

[0569] User responses to ads are tracked by the server. Metrics such as click-through rates, ad viewing time, and purchase behavior are analyzed, and this data is used to further improve the generation algorithm. This ensures that delivered ads are continuously better suited to user interests and preferences, maximizing ad effectiveness.

[0570] In this way, the system of the present invention aims to provide personalized advertisements to users and enhance the effectiveness of advertising.

[0571] The following describes the processing flow.

[0572] Step 1:

[0573] The server collects behavioral information from the user's device and online activity. It tracks website browsing history, click data, and social media activity to accumulate data that helps infer the user's interests.

[0574] Step 2:

[0575] The server collects multimodal information from the user's device. It retrieves data captured in the form of images, videos, and audio recordings to gain a deeper understanding of the user's interests and preferences.

[0576] Step 3:

[0577] The server uses an AI model to analyze the collected behavioral and multimodal information. This analysis generates user profiles and provides insights into which advertisements are most effective.

[0578] Step 4:

[0579] The server uses a generation algorithm to create ad content optimized for the user. This process automatically generates ads with creative elements based on the user's preferences.

[0580] Step 5:

[0581] The server determines the optimal timing and format for delivering the generated advertisements to the user's device. For example, it might identify times when users are relaxed and consuming media, and insert advertisements at those times.

[0582] Step 6:

[0583] The server monitors user reactions after ad delivery and collects data such as click-through rates and ad completion rates. This feedback is used to improve the AI ​​model, enhancing the accuracy and effectiveness of future ad generation.

[0584] (Example 1)

[0585] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0586] In today's world, there is a need to effectively deliver advertisements that accurately match users' diverse interests and preferences. However, existing advertising systems struggle to personalize ads with high accuracy based on individual user preferences and behaviors, which limits the effectiveness of advertising.

[0587] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0588] In this invention, the server includes means for collecting user information from multiple data sources, means for acquiring multimodal data from information devices, and means for analyzing the generated data and analyzing user attributes using machine learning algorithms. This enables the generation and delivery of optimized advertisements based on user interests and behavior, thereby improving the personalization and effectiveness of advertisements.

[0589] "User information" refers to data about a user's online and offline behavior, such as website browsing history, social media activity, and use of digital content.

[0590] "Multimodal data" is a general term for data that includes information acquired from multiple different sensors and devices, such as photos, videos, audio, and application usage data.

[0591] "Information devices" refer to electronic devices used by users, such as smartphones, tablets, and personal computers, from which data can be acquired.

[0592] A "machine learning algorithm" is a technique that uses computational models to extract patterns from data and predict future data and behavior.

[0593] "User terminals" refer to electronic devices that users use on a daily basis, and are the devices to which advertisements are delivered.

[0594] "Personalization" refers to the process of adjusting content to suit the individual user's interests and preferences, providing an optimized experience.

[0595] "Ad delivery" refers to the process of sending advertising content to a user's device and making it available for the user to view.

[0596] "Dynamic adjustment" refers to changing the system's operation in real time according to time and circumstances to achieve optimal performance.

[0597] This invention is a system for providing user-optimized advertisements. This system primarily relies on a server, which collects and analyzes information from various data sources and generates and delivers user-specific advertisements, thereby enhancing the effectiveness of the advertisements.

[0598] The server first collects the user's online activity, specifically web browsing history, social media interaction history, and digital content usage. This information is obtained using cookies, APIs, and log files and stored in a database. The obtained data serves as important foundational material for analyzing the user's interests and preferences.

[0599] Next, the device acquires multimodal data using its camera and microphone. This includes photos, videos, and audio files, and this data is processed using image recognition and audio analysis algorithms. The information obtained through this process is sent to a server to build the user's profile.

[0600] The generative AI model analyzes user behavior characteristics based on collected data and generates optimal advertisements to reach users. By using machine learning techniques, it is possible to create advertisements that are tailored to users' purchasing trends and preferences based on previously collected examples. For example, by entering a prompt such as, "Generate advertisements that suggest outdoor products that a user might be interested in, based on their recent search history and multimodal information," highly relevant ad content will be generated.

[0601] The generated ads are delivered by the server at the optimal time, tailored to the user's specific behavior patterns and lifestyle. For example, video ads are delivered at night when the user is relaxing, and audio ads are adjusted to play naturally while the user is listening to their favorite music.

[0602] In this way, the system of the present invention provides advertisements tailored to the needs of each user, thereby achieving maximum advertising effectiveness.

[0603] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0604] Step 1:

[0605] The server collects users' online behavior data. Specifically, it obtains web browsing history, social media activity logs, and digital content usage data using cookies and various APIs. This data forms the basis for creating personalized advertisements and is stored in a database. The input is user behavior data, and the output is the stored dataset.

[0606] Step 2:

[0607] The device acquires multimodal data using its camera and microphone. This process utilizes captured photos, videos, and recorded audio data. This data is processed using image recognition and audio analysis algorithms, resulting in multimodal information that reflects the user's preferences and interests. The input consists of photos, videos, and audio, while the output is the analyzed information.

[0608] Step 3:

[0609] The server integrates all collected data and builds user profiles using a generative AI model. It analyzes the data with machine learning algorithms to predict user interests, preferences, and purchasing tendencies. Inputs are integrated behavioral and multimodal information, and output is a detailed user profile.

[0610] Step 4:

[0611] The server generates the most relevant ads based on the user profile. It prompts the generation AI model with a message such as, "Generate ads suggesting outdoor products that the user might be interested in, based on their recent search history and multimodal information," and generates highly relevant ads. The input is the user profile and the prompt, and the output is the generated ad content.

[0612] Step 5:

[0613] The server delivers the generated ad content to the user's device. The ad delivery time is set to the most effective time based on the user's lifestyle. For example, a video ad might be delivered on Sunday afternoon. The input is the generated ad and its delivery schedule, and the output is the ad played on the device.

[0614] Step 6:

[0615] The server monitors user responses to ads and collects metrics such as click-through rates, viewing time, and purchase behavior. This feedback is used to further improve the generative AI model. The input is user response data, and the output is the improved ad generation algorithm.

[0616] (Application Example 1)

[0617] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0618] In modern society, as advertising personalization advances, there is a growing demand for meaningful and engaging ad delivery for users. However, traditional advertising systems have limitations in their real-time relevance to user interests and behavior, making it difficult to provide information naturally without disrupting the user's visual experience. Therefore, establishing interactive and effective ad delivery methods based on user interests, utilizing visual devices, has become a key challenge.

[0619] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0620] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, means for generating user-optimized advertising content using a generation algorithm, means for delivering the generated advertisements to the user's device, means for tracking the user's response to the advertisements and using this information to improve the generation algorithm, and means for directly overlaying advertisements onto the visual information viewed by the user. This enables the provision of optimal advertising content in real time to the visual device being used by the user, and allows for effective information transmission without compromising the visual experience.

[0621] "User behavior information" is a general term for data related to user actions, and includes behavioral data such as user click patterns, navigation history, and location information.

[0622] "Multimodal information" refers to data collected from multiple different sources, including data acquired through multiple senses such as sound, images, and text.

[0623] A "generative algorithm" is a programmatic procedure that generates output tailored to a specific purpose based on input data, and in particular, it refers to the process of creating advertisements that reflect the user's interests and preferences.

[0624] "User devices" refer to electronic information terminals that users use on a daily basis, and specifically include smartphones, tablets, smart glasses, and the like.

[0625] "User response" refers to the actions and attitudes that users show towards advertisements, and specifically includes the number of clicks, time spent on the page, and the point where users stop scrolling.

[0626] "Methods for directly overlaying advertisements onto visual information" refer to technologies that superimpose the display of advertisements onto objects viewed by users through their visual devices, creating a mechanism that allows advertisements to appear in the user's field of vision in a natural way.

[0627] In this embodiment of the invention, a system is realized in which a server collects user behavior information and multimodal information and generates advertisements using a generation algorithm. The server acquires data from user devices such as smart glasses and creates a user profile using AI software such as TensorFlow. Based on this, it generates advertisements tailored to the user's interests and displays them as an overlay on the visual device. This process makes it possible to provide relevant advertisements in real time in relation to the surrounding environment that the user is viewing.

[0628] User responses to advertisements are collected through sensors and interfaces on the user's device. This allows for the measurement of the advertisement's effectiveness, and the data is used to improve the generation algorithm. For example, if it is detected that a user has been lingering their gaze on an overlaid advertisement for an extended period, the advertisement is evaluated as visually effective.

[0629] For example, when a user is walking through a shopping mall, a promotional advertisement for that store can be displayed on the reflection of the store in their glasses. In this case, the advertisement is placed in the user's field of vision in a natural way, and the information is presented in a manner that is likely to capture the user's interest.

[0630] An example of a prompt message is: "I want to design an application that overlays advertisements for stores that a user sees in their line of sight onto their glasses as they walk around. Please tell me the necessary steps and technologies to use to analyze the user's profile with AI and select the appropriate advertisement in real time."

[0631] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0632] Step 1:

[0633] Activity information is collected from the user's device. The device records the user's location, eye-tracking data, and device usage in real time. This data serves as foundational data for understanding the user's activities and interests. User activity data is obtained as input, and pre-processed data for analysis is provided as output.

[0634] Step 2:

[0635] The server retrieves multimodal information from the user's terminal. This multimodal information includes voice commands, image data from the camera, and application operation logs. The input is the user's multimodal data, and the output is data converted into a format that can be analyzed by the AI ​​model.

[0636] Step 3:

[0637] The server analyzes the collected behavioral and multimodal information. Using an AI generation algorithm, it processes this data and generates a profile based on the user's interests and behavioral patterns. The input is the analysis data obtained in the previous step, and the output is the user's interest profile.

[0638] Step 4:

[0639] The server generates optimal advertisements based on the user's profile. Using a generative AI model, it automatically creates advertisements tailored to the user's interests. Specifically, it selects advertising materials related to products and services the user has shown interest in, and generates customized advertisements through the system. The input is the user profile, and the output is the generated advertising content.

[0640] Step 5:

[0641] The server delivers the generated advertisements to the user's visual device. The device overlays the advertisements on the smart glasses' display, seamlessly blending them into the user's real-world environment. The input is the advertisement content, and the output is its display on the smart glasses.

[0642] Step 6:

[0643] The server tracks user responses to advertisements through visual devices. The terminal uses sensors to record user eye-tracking data and interface interactions, which are then analyzed to evaluate the effectiveness of the advertisements. The input is user response data, and the output is the evaluation result of the advertisement performance.

[0644] Step 7:

[0645] The server uses user response data to improve the generation algorithm. The collected data is incorporated into the algorithm as feedback, further improving it to better match user preferences in future ad generation. The input is an evaluation of the ad response, and the output is the improved ad generation algorithm.

[0646] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0647] This invention is an advertising system that combines user behavior information, multimodal information, and an emotion engine that recognizes user emotions. The purpose of this system is to generate and deliver personalized advertisements that are tailored to the user's interests and emotions.

[0648] The server first obtains activity information from the user's device. This includes website browsing history and social media activity, which is used to infer the user's interests and preferences. Next, the server collects multimodal information from the device. This allows the server to create a more detailed profile of the user through the analysis of text data, image data, and audio data.

[0649] Furthermore, in this invention, the server uses an emotion engine to analyze the user's emotions in real time. This emotion engine analyzes the user's facial expressions, tone of voice, and input data to recognize the user's emotional state. This information is a crucial element for more effectively personalizing advertisements.

[0650] For example, if the emotion engine detects a high level of excitement while a user is watching a video, the server can generate an energetic ad that matches that state. Conversely, if it detects a relaxed state, it will generate an ad with a more subdued tone.

[0651] The generation algorithm integrates all of this data to produce optimal ad content. Because this ad is tailored based on the user's interests and emotional state, it can achieve a more effective appeal.

[0652] The generated advertisements are delivered to the user's device at a time and format dynamically adjusted by the server. This process ensures optimal ad delivery based on the user's usage patterns and the type of content they are viewing.

[0653] Ultimately, the server tracks user responses to ads, analyzing click-through rates, viewing time, emotional changes, and more. This feedback can be used to improve future ad generation and enhance the overall effectiveness of advertising campaigns.

[0654] The following describes the processing flow.

[0655] Step 1:

[0656] The server collects activity information from the user's device. This includes web history, social media activity, and purchase history, and uses this data to understand the user's interests and preferences.

[0657] Step 2:

[0658] The server retrieves multimodal information from the user's device. This refers to data such as images, videos, and audio, and is used to understand the user's content consumption patterns.

[0659] Step 3:

[0660] The server uses an emotion engine to analyze the user's current emotional state. It analyzes data obtained from the device's camera and microphone, and evaluates the user's emotions through facial expressions and tone of voice.

[0661] Step 4:

[0662] The server integrates and analyzes user behavior information, multimodal information, and emotional state to generate optimal advertisements. In this process, the generation algorithm personalizes content according to the user's characteristics.

[0663] Step 5:

[0664] The server determines the optimal timing and format for delivering generated ads based on the user's usage patterns. Ads are displayed when the user is watching videos or during quiet periods.

[0665] Step 6:

[0666] The server monitors user responses to delivered advertisements. This includes tracking ad click-through rates, view completion rates, and changes in user sentiment.

[0667] Step 7:

[0668] The server establishes a feedback loop between the emotion engine and the generation algorithm based on the collected response data. This enables continuous improvement to enhance the accuracy of future ad generation and delivery.

[0669] (Example 2)

[0670] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0671] Traditional advertising systems, while capable of generating personalized ads based on user interests, were insufficient in considering users' real-time emotional states. Furthermore, they lacked adequate optimization of dynamic ad delivery based on user behavior. This made it difficult to capture user interest in ads, highlighting the need for improved advertising effectiveness.

[0672] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0673] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, and means for performing sentiment analysis. This makes it possible to generate advertisements optimized based on the user's emotional state and interests and deliver them to the user's terminal in real time. This is expected to improve the effectiveness of the advertisements.

[0674] "Means for collecting user behavior information" refers to methods for obtaining data about user behavior, such as a user's web browsing history and social media activity.

[0675] "Means for acquiring multimodal information" refers to technologies that collect information in multiple media formats, such as text, images, and audio data, from a user's device.

[0676] "Methods for performing emotion analysis" refer to the process of analyzing a user's facial expressions, tone of voice, and input data to identify the user's emotional state.

[0677] "Using a generation algorithm" means using a computational method to generate user-optimized advertisements based on collected data.

[0678] "Means of delivering to user devices" refers to technologies that transmit generated advertisements to the user's terminal at the appropriate time and in the appropriate format.

[0679] "Means of tracking user responses" refer to methods for recording user behavior in response to advertisements, such as click-through rates and viewing time, in order to evaluate the effectiveness of the advertisements.

[0680] This invention is a system that generates and delivers advertisements that reflect the user's interests and emotions in real time. The server collects various information from the user's terminal and uses this information to personalize advertisements. Specifically, it obtains user behavior information, such as website browsing history and social media activity, and analyzes this information to infer the user's interests.

[0681] Furthermore, the server collects multimodal information. This information includes text, images, and audio, and is obtained from the user's device using the camera and microphone. By analyzing this data, a more detailed profile of the user can be created.

[0682] In emotion analysis, the server analyzes the user's emotional state in real time. Using an emotion engine, the AI ​​model analyzes the user's facial expressions and tone of voice to identify emotions such as joy, excitement, and relaxation. This makes it possible to adjust the content of advertisements to best suit the user's current emotions.

[0683] Ad generation utilizes a generative AI model. This model comprehensively processes collected data to automatically create ads that match the user's interests and emotions. The generated ads are then delivered to the user's device by a server at the optimal time. This delivery is dynamically adjusted according to the user's usage of the content.

[0684] For example, if a user is searching for a new recipe, the server can generate advertisements for cooking utensils that match their interests and provide the user with enjoyment and value by using cooking-related prompts as examples. In this way, advertisements based on user interests and emotions contribute significantly to improving advertising effectiveness.

[0685] An example of a prompt message would be, "The user has been detected as highly aroused while watching a video. Generate an ad for a new product that might interest them," which enhances ad personalization.

[0686] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0687] Step 1:

[0688] The user's device sends activity information to the server. Specifically, it transfers website browsing history and social media activity data from the device to the server. The server uses this data to build an initial dataset to infer the user's interests. The user's behavioral history data is used as input, and the output generates an initial set of tags indicating the user's interests.

[0689] Step 2:

[0690] The server collects multimodal information from the user's device. This information includes text, images, and audio data, captured, for example, via the camera and microphone of a smartphone. The server analyzes this data to generate a detailed user profile. The input is the user's text, images, and audio data, and the output is an enhanced version of the user profile.

[0691] Step 3:

[0692] The server initiates emotion analysis, identifying the user's emotional state in real time. It uses an emotion engine to analyze facial expressions and voice tone. The results of this analysis are output as data indicating the user's current emotional state. The input is data of the user's facial expressions and voice, and the output is a tag indicating the user's emotional state at that moment.

[0693] Step 4:

[0694] The generative AI model generates personalized advertisements based on previously obtained data (user interest tags, detailed profiles, and emotional states). The input is the user's overall profile information, and the output is advertising content optimized based on that profile. Specifically, the ad's colors and word choices are tailored to the user's emotions.

[0695] Step 5:

[0696] The server sends the generated advertisements to the user's device. The timing and format of delivery are dynamically adjusted, taking into account the user's content consumption habits and emotional state. The input is the generated ad data, and the output is the optimized advertisement displayed on the user's screen. The device presents the received advertisement to the user in the specified manner.

[0697] Step 6:

[0698] The server tracks user responses and analyzes the effectiveness of advertisements. It collects data on user clicks and viewing time to evaluate the success rate of advertisements. The input is user behavior data regarding advertisements, and the output is an evaluation of advertisement effectiveness that is reflected in the next advertisement generation. This allows the server to improve its generation algorithm.

[0699] (Application Example 2)

[0700] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0701] Traditional advertising delivery systems generate ads based only on user behavior information and limited data, resulting in ads that do not adequately address user interests. Furthermore, they fail to consider the user's real-time emotional state, hindering the maximization of ad effectiveness and user engagement.

[0702] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0703] In this invention, the server includes means for collecting user behavior information, means for acquiring multimodal information, means for using an emotion engine to recognize the user's emotional state, and means for using a generation algorithm to generate advertising content optimized for the user based on the emotional state. This enables the delivery of advertisements optimized based on the user's real-time emotions and interests.

[0704] "User activity information" refers to data about a user's behavior and actions on the internet, including website browsing history and social media activity history.

[0705] "Multimodal information" refers to information that encompasses multiple different types of data, such as text, images, and audio, and is used in combination with these other data.

[0706] An "emotion engine" is an algorithm or software that analyzes a user's facial expressions, voice tone, and input data to identify their emotional state.

[0707] A "generative algorithm" is a computational method for creating optimized ad content based on the user's interests and emotional state.

[0708] "User devices" refer to devices used to display advertisements and information, and include smartphones and tablet devices.

[0709] "Real-time" is a term that indicates that information is processed or delivered at the same time as, or very close to, the time it is generated.

[0710] The system of this invention combines the collection of user behavior information, acquisition of multimodal information, sentiment analysis using an emotion engine, ad optimization using a generation algorithm, and ad delivery to provide users with optimized ads in real time.

[0711] The server collects behavioral information from the user's device. Specifically, it receives web browsing history and social media activity information, and uses this data to infer interests and preferences. Next, it acquires multimodal information from devices equipped with cameras and microphones. This includes the collection of real-time audio and image data. The server analyzes this data and recognizes emotional states using an emotion engine. This emotion engine is built using machine learning libraries such as TensorFlow and PyTorch.

[0712] Based on the analyzed information, the server uses a generation algorithm to generate advertisements optimized for the user's emotional state. This generation algorithm utilizes Apache Hadoop, which can process large amounts of data at high speed, enabling the rapid generation of content that reflects the user's real-time emotional state. The generated advertisements are immediately delivered to the user's device, such as a smartphone or tablet.

[0713] As a concrete example, when the emotion engine detects a state of heightened excitement while a user is watching a video, it immediately generates and delivers advertisements for products that appeal to a sense of adventure. In this system, an example of a prompt message to the generation AI model would be: "When the user shows signs of excitement while watching a video, generate advertisements for outdoor equipment that stimulate a sense of adventure."

[0714] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0715] Step 1:

[0716] The server collects activity information from the user's device. This information includes website browsing history and social media activity records. Using this information as input, an initial dataset is generated to infer the user's interests and preferences. At this stage, the information is stored in a database and used in subsequent steps.

[0717] Step 2:

[0718] The server acquires multimodal information by utilizing the device's camera and microphone. Specifically, it takes in audio and image data as input and transfers it to the server in a format that can be analyzed in real time. This data is used to identify the user's facial expressions and voice tone. The output here is a numerical profile of facial expression data and voice characteristics.

[0719] Step 3:

[0720] The server analyzes the user's multimodal information using an emotion engine. It uses the data obtained in the previous step as input and determines the emotional state using machine learning models based on TensorFlow or PyTorch. The output is an evaluation of the emotional state, such as excitement or relaxation, which is then used in the next processing step to generate advertisements.

[0721] Step 4:

[0722] The server runs a generation algorithm based on the obtained emotional state and behavioral information. The input consists of the emotional state evaluation results and data indicating the user's interests. Using Apache Hadoop, the data is processed to generate optimal ad content tailored to the user's real-time emotions. As a result, targeted ads are output.

[0723] Step 5:

[0724] Finally, the server delivers the generated advertisements to the user's device. In this step, the advertisements are inserted in real time, taking into account the user's current content viewing status and device performance. The output is the appropriate advertisement displayed on the user's device screen. The server also tracks the user's reaction when the advertisement is delivered and feeds this feedback back to the server.

[0725] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0726] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0727] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0728] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0729] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0730] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0731] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0732] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0733] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0734] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0735] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0736] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0737] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0738] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0739] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0740] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0741] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0742] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0743] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0744] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0745] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0746] The following is further disclosed regarding the embodiments described above.

[0747] (Claim 1)

[0748] A means of collecting user behavior information,

[0749] Means for obtaining multimodal information,

[0750] A means of generating user-optimized ad content using a generation algorithm,

[0751] A means for delivering the generated advertisement to the user's device,

[0752] A means of tracking user responses to advertisements and using that information to improve generation algorithms,

[0753] A system that includes this.

[0754] (Claim 2)

[0755] The system according to claim 1, characterized in that the generated advertising content reflects the user's areas of interest.

[0756] (Claim 3)

[0757] The system according to claim 1, comprising means for dynamically adjusting the timing and format of ad delivery based on the status of the information device being used by the user.

[0758] "Example 1"

[0759] (Claim 1)

[0760] A means of collecting user information from multiple data sources,

[0761] A means of acquiring multimodal data from an information device,

[0762] A method for analyzing generated data and analyzing user attributes using machine learning algorithms,

[0763] A means of generating user-specific advertisements based on analysis results,

[0764] A means of selecting the timing for delivering the generated advertisement to the user's device and sending the advertisement in an appropriate format,

[0765] A means of tracking user behavior towards ads and providing feedback to improve ad generation algorithms,

[0766] A system that includes this.

[0767] (Claim 2)

[0768] The system according to claim 1, characterized in that the generated advertisements are personalized to reflect the user's interests and behavioral tendencies.

[0769] (Claim 3)

[0770] The system according to claim 1, characterized in that it dynamically adjusts the time and format of ad delivery according to the user's behavior history and environment.

[0771] "Application Example 1"

[0772] (Claim 1)

[0773] A means of collecting user behavior information,

[0774] Means for obtaining multimodal information,

[0775] A means of generating user-optimized ad content using a generation algorithm,

[0776] A means for delivering the generated advertisement to the user's device,

[0777] A means of tracking user responses to advertisements and using that information to improve generation algorithms,

[0778] A method of directly overlaying advertisements onto the visual information that users see,

[0779] A system that includes this.

[0780] (Claim 2)

[0781] The system according to claim 1, characterized in that the generated advertisement content reflects the user's areas of interest, and displays the advertisement using the visual device of a personal device.

[0782] (Claim 3)

[0783] The system according to claim 1, comprising means for dynamically adjusting the timing and format of ad delivery based on the status of the information device being used by the user, and adapting it to the display environment of the visual device.

[0784] "Example 2 of combining an emotion engine"

[0785] (Claim 1)

[0786] A means of collecting user behavior information,

[0787] Means for obtaining multimodal information,

[0788] Methods for performing emotion analysis,

[0789] A means for generating optimized advertising content based on the user's emotional state and interests using a generation algorithm,

[0790] A means for delivering the generated advertisement to the user's device,

[0791] A means of tracking user responses to advertisements and using this information to improve generation algorithms and sentiment analysis,

[0792] A system that includes this.

[0793] (Claim 2)

[0794] The system according to claim 1, characterized in that the generated advertisement content reflects the user's emotional state and areas of interest.

[0795] (Claim 3)

[0796] The system according to claim 1, comprising means for dynamically adjusting the timing and format of ad delivery based on the status of the information device being used by the user and the user's emotional state.

[0797] "Application example 2 when combining with an emotional engine"

[0798] (Claim 1)

[0799] A means of collecting user behavior information,

[0800] Means for obtaining multimodal information,

[0801] A means of using an emotion engine to recognize the user's emotional state,

[0802] A means of using a generation algorithm to generate ad content optimized for the user based on their emotional state,

[0803] A means for delivering generated advertisements to user devices in real time,

[0804] The means used to track user responses to delivered advertisements and improve generation algorithms,

[0805] A system that includes this.

[0806] (Claim 2)

[0807] The system according to claim 1, characterized in that the generated advertisement content reflects information about the user's emotional state.

[0808] (Claim 3)

[0809] The system according to claim 1, comprising means for dynamically and in real time adjusting the timing and format of ad delivery based on the operating status of the information device used by the user. [Explanation of Symbols]

[0810] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of collecting user behavior information, Means for obtaining multimodal information, A means of generating user-optimized ad content using a generation algorithm, A means for delivering the generated advertisement to the user's device, A means of tracking user responses to advertisements and using that information to improve generation algorithms, A system that includes this.

2. The system according to claim 1, characterized in that the generated advertising content reflects the user's areas of interest.

3. The system according to claim 1, comprising means for dynamically adjusting the timing and format of ad delivery based on the status of the information device being used by the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A