System
A system using generative AI to analyze content context and insert relevant ads addresses the issue of irrelevant internet ads, improving user experience and ad effectiveness.
Patent Information
- Application Number
- JP2024123983
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Internet advertisements often lack relevance to the content they accompany, leading to user annoyance and decreased effectiveness, making it difficult for advertisers to deliver impactful ads.
A system that utilizes a generative AI model to analyze the context of web page articles and video scenes, selecting and inserting relevant advertisements at optimal times, thereby enhancing the alignment between ad content and user experience.
Improves the effectiveness of advertisements by ensuring they are contextually relevant, thus enhancing user engagement and advertiser outcomes.
Smart Images

Figure 2026022466000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Currently, internet advertisements are often displayed based on user attribute information, which often results in advertisements that have little relevance to the content. As a result, not only do users feel annoyed and bothered, but the effectiveness of advertisements also decreases. In this situation, it is difficult for advertisers to effectively advertise, so there is a demand for increased relevance between advertisements and content and an improved user experience. [Means for solving the problem]
[0005] The present invention relates to a system that includes a means for acquiring articles from web pages accessed by a user, a means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model, a means for selecting relevant advertisements based on the analysis results and creating a list of advertisements, a means for inserting the selected advertisements into the articles at appropriate times and locations to reconstruct the articles, and a means for transmitting the reconstructed articles to the user's device. This system enables advertisements appropriate to the content context to be displayed at natural timing, regardless of the user's attribute information, thereby improving the user experience and increasing the effectiveness of advertisements. The system also includes a means for analyzing the context of individual scenes in a video during playback using a generative AI model, selecting advertisements relevant to the scenes based on the analysis results, and displaying them at specific times in the video, thereby achieving a similar effect in video content. Furthermore, the system includes a means for displaying the reconstructed articles received by the user's device and displaying advertisements corresponding to each sentence at appropriate times, further improving the accuracy and timing of advertisement display.
[0006] "Article" refers to the text information and its content posted on a web page.
[0007] A "sentence" is an individual sentence or semantically coherent unit that makes up an article.
[0008] "Generative AI model" refers to an algorithm and its implementation that uses artificial intelligence technology to analyze natural language and understand context.
[0009] "Context" refers to the meaning and connections of information in a given text or conversation and the background information.
[0010] "Analysis" refers to the process of using generative AI models to understand the meaning and structure of text or content.
[0011] "Advertising" refers to promotional content such as banners, text, and videos displayed for the purpose of promoting a particular product or service.
[0012] "List" refers to an organized list of selected advertisements or its data structure.
[0013] "Reconstruction" refers to the process of inserting ads into original articles or content and rearranging them into a new form.
[0014] "Users" refers to people who view web pages or videos.
[0015] "Device" refers to a device such as a computer, smartphone, or tablet that a user uses to view web pages or videos.
[0016] "Video" refers to dynamic video content in general, including formats that are streamed over the web.
[0017] A "scene" refers to a specific moment or sequence of images within a video. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is a system that displays advertisements for web page articles and video content at optimal timing based on not only the user's attribute information but also the context of the content itself. This system operates as follows between the server, terminal, and user.
[0040] 1. Get and split articles
[0041] The server retrieves the articles from the web pages that the user visits.
[0042] The server divides the retrieved article into sentences, for example, "Doing moderate exercise every day is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0043] 2. Context Analysis
[0044] The server uses a generative AI model to analyze the context of each sentence.
[0045] The server obtains topic and sentiment information from the generative AI model. For example, it determines that the first sentence is a "topic about exercise" and the second sentence is a "topic about diet."
[0046] 3. Ad selection
[0047] Based on the analysis results, the server searches for relevant advertisements from an advertisement database.
[0048] The server selects advertisements for fitness gyms for sentences about "exercise" and advertisements for health foods for sentences about "diet."
[0049] 4. Restructuring the article
[0050] The server incorporates the selected advertisements into the original article to generate a reconstructed article.
[0051] The server sends this article to the user's terminal.
[0052] 5. Display of articles and advertisements
[0053] The terminal displays the received reconstructed article on the user's browser.
[0054] The device displays advertisements related to each sentence at the appropriate time.
[0055] For example, if a user is viewing an article about health, an ad for a fitness gym could be displayed immediately after the sentence, "Doing moderate exercise every day is good for your physical and mental health," and an ad for health foods could be displayed immediately after the sentence, "A balanced diet will boost your immune system." This allows users to see highly relevant ads in a natural way, increasing the effectiveness of the ads.
[0056] Video analysis and ad insertion
[0057] As the video plays, the server analyzes each individual scene in the video using a generative AI model.
[0058] The server selects the appropriate advertisement for each scene and generates instructions to insert the advertisement at a specific timing in the video.
[0059] For example, if a user is watching a cooking recipe video, an advertisement related to the content of the scene will be inserted, such as an advertisement for cooking utensils, immediately after the cooking scene ends. This allows users to see advertisements that are relevant to the content of the video, increasing the effectiveness of the advertisements.
[0060] This invention is a system that optimizes the display of advertisements to users by taking into account the relevance of the content of web articles and videos as a whole, thereby improving the user experience and enabling advertisers to expect better results.
[0061] The processing flow will be explained below.
[0062] Step 1:
[0063] A user visits a specific web page.
[0064] Step 2:
[0065] The device (user's browser) sends the request to the server.
[0066] Step 3:
[0067] The server receives the request and retrieves the article from the target web page.
[0068] Step 4:
[0069] The server divides the text data of the acquired article into sentences.
[0070] Step 5:
[0071] The server feeds each sentence into a generative AI model and analyzes the context, which includes information such as topic and sentiment.
[0072] Step 6:
[0073] Based on the analysis results obtained from the generative AI model, the server searches the advertising database for advertisements related to each sentence.
[0074] Step 7:
[0075] The server selects advertisements for fitness gyms for sentences about "exercise" and advertisements for health foods for sentences about "diet."
[0076] Step 8:
[0077] The server lists the selected ads and determines the best time and place to display them.
[0078] Step 9:
[0079] The server reconstructs the article so that it includes advertisements and transmits it to the user's terminal.
[0080] Step 10:
[0081] The terminal displays the retrieved reconstructed article on the user's browser.
[0082] Step 11:
[0083] The device displays relevant ads immediately after each sentence, allowing users to see relevant ads in a natural way as they continue reading the article.
[0084] Steps for Video
[0085] Step 1:
[0086] A user begins watching a particular video.
[0087] Step 2:
[0088] The terminal sends the request to the server.
[0089] Step 3:
[0090] The server receives the request and prepares the video for analysis.
[0091] Step 4:
[0092] The server uses a generative AI model to analyze the context of each scene in the video.
[0093] Step 5:
[0094] Based on the analysis results, the server searches the advertisement database for advertisements related to each scene.
[0095] Step 6:
[0096] The server selects advertisements for cooking utensils for the "cooking scene" and advertisements for sports equipment for the "exercise scene."
[0097] Step 7:
[0098] The server creates timing instructions containing advertisements and reconfigures them to appear at specific moments in the video.
[0099] Step 8:
[0100] The server transmits the reconstructed video data to the user's terminal.
[0101] Step 9:
[0102] The device will display relevant ads at designated times while the video is playing, allowing users to see ads that are highly relevant to the content of the video in a natural way.
[0103] Example 1
[0104] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0105] With conventional methods of displaying advertisements on web pages and video content, advertisements are often displayed uniformly based on user attribute information, and are not properly associated with the context or scene of the content, which results in a lack of user interest and low advertising effectiveness. Furthermore, if the timing of advertisement insertion is inappropriate, the user experience may be impaired.
[0106] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0107] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences; means for analyzing the context of each sentence using a generative AI model; means for selecting relevant advertisements based on the analysis results and creating an advertisement list; means for inserting the selected advertisements into the article at appropriate times and locations to reconstruct it; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in the video using a generative AI model while the video is being played; means for selecting advertisements related to the scenes based on the analysis results and displaying them at specific times in the video; and means for displaying the reconstructed article received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times. This makes it possible to display advertisements that interest the user based on the context of the content, improving the effectiveness of the advertisements and the user experience.
[0108] A "server" is a central computer system that manages the data of web pages and video content accessed by users and performs various processing.
[0109] A "terminal" is a device used by a user to display web pages, reconstructed articles, and video content.
[0110] A "user" is a person who accesses and views web pages and video content.
[0111] An "article" is the text content on a web page, the information that users view.
[0112] A "sentence" is the smallest text unit that makes up an article, and is a sentence that completes one idea.
[0113] "Context" refers to the context and background of a particular sentence, and is an element that is analyzed by generative AI models.
[0114] A "generative AI model" is an artificial intelligence model for natural language processing and text generation, and is used for contextual analysis of sentences and advertisement selection.
[0115] An "advertisement" is a visual or audio message used by a business or service to promote its products or information.
[0116] An "ad list" is a collection of multiple advertisements selected based on the analyzed context.
[0117] A "reconstructed article" is a newly created text content created by appropriately inserting selected advertisements into the original article.
[0118] A "video scene" refers to a specific scene or segment within video content, and is the unit used as the basis for selecting advertisements.
[0119] "Timing" refers to the right moment or time to display an ad, and is an important factor in ensuring a good user experience.
[0120] This invention is a system for optimizing the display of advertisements on web pages and video content, and operates mainly between a server, a terminal, and a user as follows.
[0121] Configuration and Processing Overview
[0122] Retrieving and splitting articles
[0123] The server retrieves articles from web pages accessed by users. The retrieved articles are saved in HTML or text format and broken down into sentences. Specifically, natural language processing libraries such as nltk and spaCy can be used with programming languages such as Python. For example, the following can be done using a script:
[0124] python
[0125] import spacy
[0126] nlp = spacy.load("en_core_web_sm")
[0127] text = "Moderate exercise every day is good for your physical and mental health. A balanced diet strengthens your immune system."
[0128] doc = nlp(text)
[0129] sentences = [sent.text for sent in doc.sents]
[0130] print(sentences) Output: ['Daily moderate exercise is good for your physical and mental health.', 'A balanced diet strengthens your immune system.']
[0131] Contextual Analysis
[0132] The server uses a generative AI model (specifically, OpenAI's GPT-4) to analyze the context of each sentence. This analysis extracts the topic and sentiment information of each sentence. For example, for the sentence "Doing moderate exercise every day is good for your physical and mental health," the server sends the following prompt to the generative AI model:
[0133] Sample prompt: "Analyze the context of the following sentence to extract topic and emotional information. Sentence: 'Doing moderate exercise every day is good for your physical and mental health.'"
[0134] Based on the response from the generative AI model, "exercise topics" and "positive sentiment" are obtained.
[0135] Ad selection
[0136] The server then searches for relevant advertisements from its advertising database based on the analysis results. The advertisements are selected based on specific topics and emotions. For example, for the context of "exercise," it would select advertisements for fitness gyms.
[0137] Generate and submit restructured articles
[0138] The server then embeds the selected ads into the original article, generating a reconstructed article by inserting the ads at the appropriate places and times. The reconstructed article is then saved in HTML format and sent to the user's device.
[0139] View articles and ads
[0140] The device then displays the received reconstructed article in the user's browser. Specifically, it uses an HTML rendering engine to display the reconstructed article and advertisements, allowing the user to view the advertisements in a natural way.
[0141] Application to video content
[0142] The server analyzes scenes in the video as they play and selects appropriate ads. It uses a generative AI model to analyze the context of the scene and determine the timing of ad insertion. For example, while playing a cooking recipe video, it can generate instructions such as "display an ad for cooking utensils after the cooking scene."
[0143] Sample prompt: "Select the appropriate ad based on the content of the following video scene. Scene description: 'A scene introducing a cooking recipe.'"
[0144] Specific examples
[0145] If a user is viewing an article about health, an ad for a fitness gym could be displayed immediately after the sentence, "Doing moderate exercise every day is good for your physical and mental health," and an ad for health foods could be displayed immediately after the sentence, "A balanced diet will boost your immune system." This allows the user to naturally see more relevant ads, improving the effectiveness of the ads.
[0146] In the case of video content, if a user is watching a cooking recipe video, advertisements that match the content of the scene will be inserted, such as an advertisement for cooking utensils, immediately after the cooking scene ends.
[0147] As described above, this system utilizes generative AI models to effectively display ads that match users' interests based on context, thereby improving user experience and advertising effectiveness.
[0148] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0149] Step 1:
[0150] Retrieving articles
[0151] The server retrieves HTML content from the web page accessed by the user. It extracts the article portion from the retrieved HTML content and stores it in text format. The input is the web page URL, and the output is the extracted article text. Specifically, it uses the requests or BeautifulSoup library to retrieve the required data from the web page.
[0152] Step 2:
[0153] Article division
[0154] The server splits the retrieved articles into sentences. The input is the extracted article text, and the output is a list of split sentences. Specifically, it splits sentences using natural language processing libraries such as nltk and spaCy.
[0155] Step 3:
[0156] Contextual Analysis
[0157] The server uses a generative AI model to analyze the context of each sentence. The input is a list of segmented sentences, and the output is a list of sentences with topic and emotional information added. Specifically, the server sends the following prompt to the generative AI model:
[0158] Sample prompt: "Analyze the context of the following sentence to extract topic and emotional information. Sentence: 'Doing moderate exercise every day is good for your physical and mental health.'"
[0159] Based on the response from the generative AI model, topic and sentiment information is assigned to each sentence.
[0160] Step 4:
[0161] Ad selection
[0162] The server selects relevant advertisements based on the analysis results and creates an advertisement list. The input is a list of sentences with topic and emotion information, and the output is a list of advertisements. Specifically, it searches an advertisement database and selects the most suitable advertisement. For example, it matches a fitness gym advertisement to a sentence with the topic "exercise."
[0163] Step 5:
[0164] Generate restructured articles
[0165] The server inserts the selected advertisements into the article at the appropriate time and place to generate a reconstructed article. The input is the advertisement list and the original article text, and the output is the reconstructed article text. Specifically, the server inserts advertisements after each sentence.
[0166] Step 6:
[0167] Submitting a restructured article
[0168] The server sends the reconstructed article to the user's device. The input is the reconstructed article text, and the output is a notification of completion of transmission to the user's device. Specifically, the data is sent using the HTTP protocol.
[0169] Step 7:
[0170] View articles and ads
[0171] The terminal displays the received reconstructed article in the user's browser. The input is the reconstructed article text, and the output is the web page displayed to the user. Specifically, it uses an HTML rendering engine to display the article and advertisements in the browser.
[0172] Step 8:
[0173] Video scene analysis
[0174] The server analyzes individual scenes in the video as it plays. The input is the video data, and the output is the analyzed scene information. Specifically, it uses a generative AI model to analyze the context of the scene.
[0175] Sample prompt: "Select the appropriate ad based on the content of the following video scene. Scene description: 'A scene introducing a cooking recipe.'"
[0176] Step 9:
[0177] Ad selection and insertion instruction generation
[0178] The server selects appropriate advertisements based on the analyzed scene information and generates instructions to insert the advertisements at specific times in the video. The input is the analyzed scene information and an advertisement database, and the output is insertion instructions. Specifically, the server selects the optimal advertisement for each scene and determines the appropriate timing.
[0179] The above is the processing flow of this system. By clearly indicating the specific operations and inputs / outputs performed at each step, the processing flow of the entire system becomes clear.
[0180] (Application example 1)
[0181] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0182] Conventional advertising display systems have difficulty displaying ads that are relevant to the context of the content being viewed by the user or to specific scenes at the appropriate time. Furthermore, there is a problem of reduced advertising effectiveness due to the lack of relevance between the content of articles or videos and the ads. Furthermore, this can lead to a poor user experience, making it difficult for advertisers to provide effective ads.
[0183] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0184] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; means for reconstructing the article by inserting the selected advertisements into the article at appropriate times and places; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in a video accessed by the user using a generative AI model while the video is being played; and means for selecting advertisements related to scenes in the video based on the analysis results and displaying them at specific times in the video. This makes it possible to display advertisements that are highly relevant to the content viewed by the user at appropriate times.
[0185] A "server" is a computer system that provides data and resources for joint use over a network.
[0186] "User" means any person, organization, or device that uses a computer system or network service.
[0187] A "web page" is an information page published on the Internet, written in a markup language such as HTML, and viewable using a browser.
[0188] "Article" refers to a single work or document published in a newspaper, magazine, web page, etc.
[0189] A "sentence" is the smallest unit of a sentence or piece of writing, a series of words that has complete meaning.
[0190] "Context" refers to the relationship between adjacent words or sentences, and the overall situation in which it is placed.
[0191] A "generative AI model" is an artificial intelligence model trained using neural networks on large datasets, and is capable of generating, understanding, and analyzing text.
[0192] "Advertising" refers to information created to introduce products and services and encourage consumers to purchase them.
[0193] "Video" refers to moving video content that is made up of a series of images and sounds.
[0194] A "scene" is a single continuous action or event in video content that does not span a specific time period.
[0195] A "list" is a data structure that lists and organizes multiple items.
[0196] As an embodiment of the present invention, a system that performs the following processing steps is constructed.
[0197] First, the server retrieves the article from the web page accessed by the user. The retrieved article is then divided into sentences. The context of each sentence is then analyzed using a generative AI model. For example, the natural language processing model "GPT-4" is used to extract the topic and sentiment information of each sentence.
[0198] After the context analysis is completed, the server selects relevant advertisements from the advertisement database based on the analysis results, lists the selected advertisements, and reconstructs the article based on them. The reconstructed article is then sent to the user's device.
[0199] Next, while the video being viewed by the user is being played, the server analyzes each scene in the video using a generative AI model. It selects the appropriate ad for each scene and generates instructions to display it at a specific moment in the video. These instructions are then sent to the user's device, and the ad is displayed at the appropriate time.
[0200] The specific configuration for realizing this system involves running the following program. First, the server retrieves articles from a webpage using the Python libraries "requests" and "BeautifulSoup," and then splits the articles into sentences using the "split_by_sentence" function. Next, contextual analysis is performed using the generative AI model "GPT-4." The analysis is performed by sending a prompt sentence using the "GPT-4" API and obtaining a response to it. The following prompt sentence is used as an example of this analysis.
[0201] "Analyze the context of this sentence and identify the topic and emotion: Regular exercise every day is good for your physical and mental health."
[0202] "Analyze the context of this sentence and identify the topic and emotion: A balanced diet boosts your immunity."
[0203] Based on the analysis results, the server selects the most suitable advertisement from the advertisement database and reconstructs the article, which is then sent to the user's device and displayed in the browser.
[0204] For video analysis, each scene in the video is identified and the context of each scene is analyzed using a generative AI model. For example, if a user is watching a cooking recipe video, an advertisement for cooking utensils will be displayed immediately after the cooking scene ends. For context analysis for each scene, a prompt in the form of "Analyze the context of this scene and identify related advertisements:" is used.
[0205] In this way, it is possible to build a system that dynamically displays highly relevant advertisements at the appropriate time based on the content the user is viewing, thereby improving the user experience and maximizing advertising effectiveness.
[0206] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0207] Step 1:
[0208] The server retrieves the article of the web page accessed by the user. As input, it receives the URL of the web page accessed by the user. As output, it returns the text data of the entire article of the web page. Specifically, the server uses the "requests" library to retrieve the HTML of the web page and "BeautifulSoup" to extract the text data.
[0209] Step 2:
[0210] The server splits the retrieved articles into sentences. As input, it receives the text data of the entire article. As output, it returns a list of text split into sentences. Specifically, the server uses Python's string manipulation functions to split the article using delimiters such as "."
[0211] Step 3:
[0212] The server analyzes the context of each sentence using a generative AI model. As input, it receives a text list divided into sentences. As output, it returns the context information (topic and sentiment information) of each sentence. Specifically, the server sends a prompt sentence to the API of the generative AI model "GPT-4" and obtains the analysis results.
[0213] Step 4:
[0214] The server selects relevant advertisements based on the results of the context analysis. As input, it receives the context information of each sentence. As output, it returns a list of relevant advertisements. Specifically, the server matches the context information with an advertisement database to select the most suitable advertisement.
[0215] Step 5:
[0216] The server reconstructs the article by inserting the selected advertisements at the appropriate time and place. As input, it receives the segmented sentences and a list of selected advertisements. As output, it returns the reconstructed article with the advertisements inserted. Specifically, the server executes logic to insert advertisements at the appropriate place for each sentence.
[0217] Step 6:
[0218] The server sends the reconstructed article to the user's device. As input, it receives the reconstructed article. As output, it sends the article to the user's device. Specifically, the server uses an HTTP request to display the reconstructed article in the user's browser.
[0219] Step 7:
[0220] The server analyzes the context of each individual scene in a video while the video accessed by the user is being played. As input, it receives video data and timestamps. As output, it returns contextual information for each scene. Specifically, the server uses a generative AI model to analyze each scene in the video.
[0221] Step 8:
[0222] The server selects advertisements relevant to the scenes based on the analysis results and generates instructions to display them at specific times in the video. It receives contextual information for each scene as input. It returns instructions to display the advertisements as output. Specifically, the server compares the analysis results for each scene with an advertisement database, selects appropriate advertisements, and sends instructions to the user's device for the timing of their display.
[0223] Step 9:
[0224] The device displays the reconstructed article in the user's browser and displays advertisements corresponding to each sentence at the appropriate time. As input, it receives the reconstructed article and advertisement display instructions received from the server. As output, it displays the article and advertisements in the user's browser. Specifically, the device controls the browser display using HTML and JavaScript and displays advertisements at the appropriate time.
[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0226] This invention is a system that displays advertisements at optimal times based on not only user attribute information but also the context of the content itself and the user's emotional state in web page articles and video content. By adding an emotion engine, this system can recognize user emotions and further improve the accuracy and effectiveness of advertisement display.
[0227] 1. Get and split articles
[0228] When a user accesses a specific web page, the device (user's browser) sends the request to the server. The server receives the request and retrieves the article from the target web page. The server then divides the retrieved article into sentences. For example, it may divide the sentences into "Daily moderate exercise is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0229] 2. Context analysis and emotional state recognition
[0230] The server uses a generative AI model to analyze the context of each sentence. Context includes topic and emotional information. The server also uses an emotion engine to detect the user's emotional state. For example, the server can analyze facial expressions and tone of voice using a camera or microphone.
[0231] 3. Ad selection
[0232] Based on the analysis results, the server searches the advertisement database for advertisements related to each sentence and the user's emotional state. For example, if the sentence is related to "exercise" and the user is in a "positive" emotional state, an advertisement for a fitness gym will be selected. The analysis results of the emotion engine play an important role in advertisement selection.
[0233] 4. Restructure and submit your article
[0234] The server then incorporates the selected advertisements into the original article to generate a reconstructed article, which is then sent to the user's device. The reconstructed article also includes advertisements based on the user's emotional state.
[0235] 5. Display of articles and advertisements
[0236] The device displays the reconstructed article in the user's browser. The device displays relevant advertisements immediately after each sentence. For example, a fitness gym advertisement is displayed after the sentence "Daily moderate exercise is good for your physical and mental health," while a health food advertisement is displayed after the sentence "A balanced diet strengthens your immune system." As the user continues reading the article, they can naturally see highly relevant advertisements based on their emotional state and context.
[0237] Video analysis and ad insertion
[0238] When a user starts watching a particular video, the device sends the request to the server, which receives the request and prepares the video for analysis. The server uses a generative AI model to analyze the context of each scene in the video and an emotion engine to recognize the user's emotional state.
[0239] Based on the analysis results, the server searches the advertisement database for advertisements relevant to each scene and the user's emotional state. For example, if a cooking scene is displayed and the user finds it "interesting," an advertisement for cooking utensils will be selected. The server creates timing instructions including the selected advertisement and reconstructs the video to be displayed at a specific moment. The reconstructed video data is sent to the user's device. The device displays relevant advertisements at the specified timing for the video being played, allowing the user to naturally view advertisements based on the content of the video and the user's emotional state.
[0240] This invention is a system that can take into account the user's emotional state, increase the effectiveness of advertising, and improve the user experience by combining advertising display related to the content of web articles and entire videos with an emotion engine.
[0241] The processing flow will be explained below.
[0242] Step 1:
[0243] A user visits a specific web page.
[0244] Step 2:
[0245] The device (user's browser) sends the request to the server.
[0246] Step 3:
[0247] The server receives the request and retrieves the article from the target web page.
[0248] Step 4:
[0249] The server divides the text data of the acquired article into sentences.
[0250] Step 5:
[0251] The server feeds each sentence into a generative AI model to analyze the context, which includes topic and sentiment information.
[0252] Step 6:
[0253] The server uses an emotion engine to detect the user's emotional state by analyzing the user's facial expressions and tone of voice using the device's camera and microphone.
[0254] Step 7:
[0255] Based on the analysis results of the generative AI model and the results of the emotion engine, the server selects advertisements from the advertising database that are relevant to each sentence and the user's emotional state.
[0256] Step 8:
[0257] The server lists the selected ads and determines the best time and place to display them.
[0258] Step 9:
[0259] The server reconstructs the article so that it includes advertisements and transmits it to the user's terminal.
[0260] Step 10:
[0261] The terminal displays the received reconstructed article on the user's browser.
[0262] Step 11:
[0263] The device displays relevant ads immediately after each sentence, allowing users to naturally see relevant ads based on their context and emotional state as they continue reading the article.
[0264] Steps for Video
[0265] Step 1:
[0266] A user begins watching a particular video.
[0267] Step 2:
[0268] The terminal sends the request to the server.
[0269] Step 3:
[0270] The server receives the request and prepares the video for analysis.
[0271] Step 4:
[0272] The server uses a generative AI model to analyze the context of each scene in the video.
[0273] Step 5:
[0274] The server uses an emotion engine to set the user's emotional state by analyzing the user's emotional response to the video they are watching.
[0275] Step 6:
[0276] The server searches the advertising database for advertisements relevant to the scene based on the analysis results obtained from the generative AI model and emotion engine.
[0277] Step 7:
[0278] The server selects advertisements for cooking tools for the "cooking scene" and advertisements for sports equipment for the "exercise scene." The server adjusts the content and timing based on the user's emotional state.
[0279] Step 8:
[0280] The server creates timing instructions containing advertisements and reconfigures them to appear at specific moments in the video.
[0281] Step 9:
[0282] The server transmits the reconstructed video data to the user's terminal.
[0283] Step 10:
[0284] The device will display relevant ads at designated times while the video is playing, allowing users to see ads that are highly relevant to the content of the video in a natural way.
[0285] Example 2
[0286] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0287] Conventional advertising systems typically display ads based on user attribute information, but this method fails to fully consider each user's emotional state or the context of the content, resulting in ineffective ad display. Furthermore, even in video content, the timing and content of ad insertions often fail to match the user's emotions, resulting in a poor user experience. It is necessary to provide a system that can resolve these issues and display more effective and relevant ads to users.
[0288] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0289] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; means for acquiring emotional information using an emotion engine that recognizes the user's emotional state; means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; means for reconstructing the article by inserting the selected advertisements into the article at appropriate times and locations; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in the video using a generative AI model while the video is being played; and means for selecting advertisements related to the scenes based on the analysis results and displaying them at specific times in the video. This enables more effective and relevant advertisements to be displayed based on the user's emotional state and the context of the content.
[0290] "User" means an individual who accesses the Website or Application and uses or consumes its content.
[0291] A "device" is a device, such as a computer, smartphone, or tablet, that a user uses to connect to the Internet and view web pages and videos.
[0292] A "server" is a computer system that stores, manages, and transmits and receives data over a network, and is a device that provides content in response to user requests.
[0293] A "web page" is a unit of information that is published on the Internet and can be viewed by users through a browser; it is a digital page that consists of text, images, videos, etc.
[0294] An "article" is text-based information content posted on a web page, and may be in the form of news, blogs, commentaries, etc.
[0295] A "sentence" is a complete sentence in an article, the smallest meaningful unit of text.
[0296] A "generative AI model" is a machine learning model that uses artificial intelligence to generate, analyze, translate, and perform other tasks on text, and is based on natural language processing technology.
[0297] "Context" refers to the meaning or background that a particular part of a sentence or image has in relation to other parts or the whole.
[0298] "Analysis" is the process of breaking down data or information to understand its components and patterns.
[0299] An "emotion engine" is a technology that analyzes external information such as a user's facial expressions and tone of voice to determine the user's emotional state.
[0300] "Emotional state" refers to the type and intensity of emotion a user feels in a particular situation, and includes states such as "positive," "negative," and "interesting."
[0301] "Advertising" is promotional content that presents information about products and services and stimulates consumers' desire to purchase them.
[0302] A "list" refers to an ordered arrangement of related items.
[0303] "Timing" refers to the moment in time when a particular event or action occurs.
[0304] A "location" refers to a specific location within a web page or video.
[0305] "Reconstruction" means constructing the original content in a new form by inserting new elements or rearranging them.
[0306] "Video" is digital media content that combines a sequence of images and sounds to tell a visual and audio story.
[0307] A "scene" refers to a specific sequence of images and sounds in a video, and is a unit that depicts a specific situation or activity.
[0308] This invention is a system that displays advertisements at optimal times based on not only user attribute information but also the context of the content itself and the user's emotional state in web page articles and video content. By adding an emotion engine, this system can recognize user emotions and further improve the accuracy and effectiveness of advertisement display.
[0309] First, when a user accesses a specific web page, their device (the user's browser) sends the request to a server. For example, if a user accesses a news site to read an article about health, this request is sent. The server receives the request and retrieves the article from the target web page. Next, the server uses a generative AI model (for example, a model using natural language processing technology) to divide the retrieved article into sentences. Specifically, it divides the article into sentences such as "Daily moderate exercise is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0310] The server uses a generative AI model to analyze the context of each sentence. This context includes topic and emotional information. The server also uses an emotion engine (e.g., technology that analyzes facial expressions and vocal tones) to recognize the user's emotional state. Specifically, the server analyzes facial expressions and vocal tones captured in real time through the device's camera and microphone to determine whether the user is in a "positive," "negative," "interested," or other emotional state.
[0311] Next, the server searches the advertisement database for advertisements related to each sentence and the user's emotional state based on the analysis results. For example, if the sentence is related to "exercise" and the user is in a "positive" emotional state, advertisements for fitness gyms will be selected. The selected advertisements will be listed.
[0312] The server then incorporates the selected advertisements into the original article to generate a reconstructed article. For example, it might insert an advertisement for a fitness gym after the sentence, "Doing moderate exercise every day is good for your physical and mental health." The reconstructed article data is then sent to the terminal.
[0313] The device then displays the reconstructed article in the user's browser. Relevant advertisements are displayed immediately after each sentence, allowing users to naturally see highly relevant advertisements based on their emotional state and context as they read the article. For example, an advertisement for a health food product is displayed after the sentence, "A balanced diet boosts your immune system."
[0314] When a user starts watching a specific video, the device sends a request to the server. The server receives the request and prepares the video for analysis. The server uses a generative AI model to analyze the context of each scene in the video and an emotion engine to recognize the user's emotional state and suggest advertisements. The specific analysis takes into account the content displayed in each scene of the video and the user's reactions (facial expressions, tone of voice, etc.).
[0315] For example, if a user finds a cooking scene "interesting," an advertisement for cooking utensils will be selected. The server reconstructs the selected advertisement to be displayed at a specific moment in the video, and sends the reconstructed video data to the device. The device then displays the relevant advertisement at a specified timing during video playback. While watching the video, the user can naturally see advertisements based on their emotional state and the context of the video.
[0316] For example, the following prompt might be used: "Explain how to analyze the article content of a web page a user views and use an emotion engine to display relevant ads."
[0317] By implementing this system, it will be possible to display more effective and relevant advertisements that take into consideration the user's emotional state and the context of the content, which is expected to improve the effectiveness of advertisements and the user experience.
[0318] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0319] Step 1:
[0320] When a user accesses a specific web page, the device (the user's browser) sends the request to the server. The input is the URL of the web page being accessed, and the output is the article data for the web page. In concrete terms, when a user accesses the "health" category on a news site, the browser requests "https: / / example.com / health" from the server. The server analyzes the request and retrieves the article data from the specified URL.
[0321] Step 2:
[0322] The server splits the retrieved article data into sentences. The input is the article data of the web page, and the output is text data split into sentences. Specifically, the server uses the "Sentence Splitter API" to split the article text into sentences such as "Doing moderate exercise every day is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0323] Step 3:
[0324] The server uses a generative AI model (e.g., a model using natural language processing technology) to analyze the context of each sentence. The input is text data divided into sentences, and the output is the context data for each sentence. Specifically, the server sends each sentence to the generative AI model, which extracts the topics "health" and "exercise" and positive sentiment information from the sentence "exercise is good for your health."
[0325] Step 4:
[0326] The server recognizes the user's emotional state using an emotion engine (e.g., technology that analyzes facial expressions and voice tones). The input is real-time data acquired through the device's camera and microphone, and the output is the user's emotional state data. Specifically, the server analyzes the user's facial expressions and voice tone to determine whether the user's emotional state is "positive," "negative," "interested," etc.
[0327] Step 5:
[0328] The server searches for relevant advertisements from the advertisement database based on the analysis results (contextual data and emotional state data). The input is contextual data and emotional state data, and the output is a list of selected advertisements. Specifically, the server selects advertisements for fitness gyms if the sentence is related to "exercise" and the user is in a "positive" emotional state.
[0329] Step 6:
[0330] The server incorporates the selected advertisements into the original article and generates a reconstructed article. The input is text data divided into sentences and a list of advertisements, and the output is the reconstructed article data. Specifically, the server adds an advertisement for a fitness gym after the sentence, "Doing moderate exercise every day is good for your physical and mental health."
[0331] Step 7:
[0332] The server sends the reconstructed article data to the terminal. The input is the reconstructed article data, and the output is the data sent to the user's terminal. In concrete terms, the server sends the reconstructed article data to the terminal and it is displayed in the specified browser.
[0333] Step 8:
[0334] The terminal displays the received reconstructed article in the user's browser. The input is the reconstructed article data, and the output is the information displayed in the user's browser. In concrete terms, the terminal displays the reconstructed article in the browser, and an advertisement for a health food is displayed after the sentence, "A balanced diet boosts your immune system."
[0335] Step 9:
[0336] When a user starts watching a particular video, the device sends a request to the server. The input is the URL of the video being watched, and the output is the video data. In concrete terms, when a user requests "https: / / example.com / cooking" to watch a cooking tutorial video, the device sends this information to the server.
[0337] Step 10:
[0338] The server receives the request and prepares the target video for analysis. The input is the video data, and the output is data for each scene in the video. Specifically, the server divides the video into frames and identifies each scene.
[0339] Step 11:
[0340] The server uses a generative AI model to analyze the context of each scene in the video. The input is the video data for each scene, and the output is the contextual data for each scene. Specifically, the server sends each scene to the generative AI model, identifies the context, and extracts information such as "cooking scene" or "exercise scene."
[0341] Step 12:
[0342] The server uses an emotion engine to recognize the user's emotional state. The input is real-time data acquired through the device's camera and microphone, and the output is the user's emotional state data. Specifically, the server analyzes the user's facial expressions and tone of voice while watching a video to determine whether the user is in an "interesting" emotional state.
[0343] Step 13:
[0344] The server searches for relevant advertisements from the advertisement database based on the analysis results (scene context data and emotional state data). The input is the scene context data and emotional state data, and the output is a list of selected advertisements. Specifically, when a cooking scene is displayed and the user feels that it is "interesting," the server selects an advertisement for cooking utensils.
[0345] Step 14:
[0346] The server inserts the selected advertisements into specific scenes of the video and reconstructs it. The input is the video data for each scene and the advertisement list, and the output is the reconstructed video data. Specifically, the server inserts advertisements at the specified timing and reconstructs the video.
[0347] Step 15:
[0348] The server transmits the reconstructed video data to the terminal. The input is the reconstructed video data, and the output is the data to be transmitted to the user's terminal. In concrete terms, the server transmits the reconstructed video data to the terminal.
[0349] Step 16:
[0350] The device displays relevant advertisements at specified times during the video being played. The input is the reconstructed video data, and the output is a video with advertisements that is displayed to the user. Specifically, the device displays advertisements for cooking utensils during specific scenes in the video.
[0351] (Application example 2)
[0352] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0353] Conventional advertising systems for web pages and videos display ads based on user attribute information and the general context of the content, but this does not fully consider the user's emotional state or the specific context of the moment. As a result, ads are often ineffective or not perceived as natural by users. Furthermore, even when using smart glasses or head-mounted displays, there is a lack of technology that can analyze the user's emotional state in real time and display appropriate ads. This has led to a demand for improved user experience.
[0354] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0355] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; sentiment analysis means for recognizing an emotional state from the context of the acquired articles; means for selecting relevant advertisements based on the analysis results and the emotional state and creating a list of advertisements; means for reconstructing the articles by inserting the selected advertisements into the articles at appropriate times and places; means for transmitting the reconstructed articles to the user's terminal; and means for displaying the reconstructed articles received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times. This enables more effective and natural advertisement display based on the user's emotional state and the specific context of the content.
[0356] A "User" is an individual who accesses a particular web page or video content.
[0357] A "web page" is a unit of information that is published on the Internet and consists of text, images, videos, etc.
[0358] An "article" refers to a sentence or text content that is published on a web page.
[0359] A "sentence" is a unit of text that has an independent meaning.
[0360] A "generative AI model" is an artificial intelligence model used to perform natural language processing and contextual analysis.
[0361] "Emotion analysis means" is a technology for recognizing a user's emotional state from facial expressions, tone of voice, etc.
[0362] "Advertising" refers to messages or content that convey information about products or services and encourage desired behavior.
[0363] A "list" refers to an ordered arrangement of related items.
[0364] "Timing" refers to the point in time at which a particular event or action should occur.
[0365] "Reconstruction" refers to inserting advertisements into existing content and rearranging it into a new form.
[0366] A "terminal" is a device used by a user, such as a computer, smart glasses, or head-mounted display.
[0367] "Video" is a media format that expresses movement by playing a series of still images.
[0368] A "scene" refers to a part of a video that depicts a specific situation or situation.
[0369] "Smart glasses" are a wearable device in the form of glasses worn by the user that can display information and analyze emotions.
[0370] A "head-mounted display" is a display device worn on the head that provides visual information.
[0371] "Real time" means that things happen at the same speed as real time.
[0372] To implement this invention, a system including a server, a terminal, and a user is required. The system can display advertisements related to the content of articles and videos on a web page at optimal timing.
[0373] The server retrieves articles from web pages accessed by users. The retrieved articles are divided into sentences, and the context of each sentence is analyzed using a generative AI model. At the same time, the user's emotional state is recognized using a sentiment analysis method. Based on the analysis results and the user's emotional state, relevant ads are selected and a list of ads is created.
[0374] The selected advertisements are inserted into the article at the appropriate time and place, and the article is reconstructed. This reconstructed article is then sent to the user's device. The user's device displays the received reconstructed article, and the advertisements corresponding to each sentence are displayed at the appropriate time.
[0375] Furthermore, while the video is playing, the server analyzes the context of each scene in the video using a generative AI model. The user's emotional state is recognized from the scene in the video, and an advertisement related to the scene is selected based on the analysis results and the emotional state. The selected advertisement is displayed at a specific time in the video.
[0376] When using smart glasses or head-mounted displays, the user's emotional state is analyzed using a built-in camera and microphone, and advertisements are selected and displayed in real time based on the emotional state and content context.
[0377] The hardware used includes smart glasses or a head-mounted display, a built-in camera and microphone, and the software used includes a sentiment analysis engine, a generative AI model (e.g., GPT-4), a web server (e.g., Node.js), and an advertising database (e.g., MongoDB).
[0378] As a concrete example, when a user is reading an article on a web page that reads, "Doing moderate exercise every day is good for your physical and mental health," the generative AI model analyzes the context of this sentence, and if the emotion engine recognizes the user's "positive" emotional state, an advertisement for a "fitness gym" will be selected. The generative AI model is analyzed using the following prompt sentence:
[0379] Example prompt sentence:
[0380] Analyze the context and emotional state of the following sentences and select the appropriate ad.
[0381] Document: Regular exercise every day is good for your physical and mental health.
[0382] Emotion: Positive
[0383] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0384] Step 1:
[0385] The server receives a request for the web page accessed by the user and retrieves the article from the web page.
[0386] Input: User request (Web page URL)
[0387] Output: Text data of the retrieved article
[0388] Specific operation: The server analyzes the HTTP request, retrieves the web page content from the specified URL, extracts the text data of the article, and passes it to the next step.
[0389] Step 2:
[0390] The server divides the retrieved article into sentences.
[0391] Input: Article text data
[0392] Output: A list of text data split into sentences
[0393] How it works: The server uses a text analysis algorithm to split the article into sentences. You can use a Python natural language processing library, for example.
[0394] Step 3:
[0395] The server analyzes the context of each sentence using a generative AI model.
[0396] Input: A list of text data split into sentences
[0397] Output: Analysis result data on the context of each sentence
[0398] How it works: It calls an API to operate the generative AI model and analyzes the context of each sentence, generating structured data that includes the topic and sentiment information of the sentence.
[0399] Step 4:
[0400] The server utilizes emotion analysis means to recognize the user's emotional state.
[0401] Input: User's facial expression data (camera) or voice data (microphone)
[0402] Output: User's emotional state (e.g. positive, negative, interesting, etc.)
[0403] How it works: The server sends data from the camera and microphone to the emotion analysis engine, which analyzes the user's current emotional state. The emotion analysis engine identifies the user's emotional state based on their facial expressions and tone of voice.
[0404] Step 5:
[0405] The server selects relevant advertisements from an advertisement database based on the analysis results and the emotional state.
[0406] Input: Contextual information for each sentence and the user's emotional state
[0407] Output: A list of relevant ads
[0408] Specific operations: Query the advertisement database and select advertisements that match the context of the sentence and the user's emotional state. Retrieve information about the selected advertisements in list form.
[0409] Step 6:
[0410] The server reconstructs the article by inserting the selected advertisements at the appropriate time and place.
[0411] Input: Context information for each sentence, list of ads
[0412] Output: Reconstructed article
[0413] Specific operation: The server inserts advertisements into the original text data and restructures the article so that the advertisements are displayed immediately after the sentence or at natural times along the flow of the content.
[0414] Step 7:
[0415] The server sends the reconstructed article to the user's terminal.
[0416] Input: Reconstructed article
[0417] Output: Send article to user's device
[0418] What it does: Converts the reconstructed article into HTML or another appropriate format and sends it to the user's browser or device.
[0419] Step 8:
[0420] The terminal displays the received reconstructed article and displays advertisements corresponding to each sentence at appropriate times.
[0421] Input: Reconstructed article
[0422] Output: Articles and advertisements displayed in the browser
[0423] Specific operation: The user's device receives the reconstructed article and displays it in a web browser, with relevant advertisements displayed immediately after each sentence of the article.
[0424] Step 9:
[0425] As the video plays, the server analyzes the context of each scene in the video using a generative AI model.
[0426] Input: Video file
[0427] Output: Context analysis results of video scenes
[0428] How it works: The server analyzes the video data frame by frame and understands the context of each scene based on a generative AI model.
[0429] Step 10:
[0430] The server recognizes the emotional state from scenes in the video and selects advertisements related to the scenes based on the analysis results and the emotional state.
[0431] Input: Context analysis results of video scenes, user emotional state
[0432] Output: A list of ads related to the scene
[0433] Specific operation: The server uses an emotion analysis engine to analyze the user's emotional state in real time and selects advertisements from the advertisement database that correspond to the context of the video scene.
[0434] Step 11:
[0435] The server displays the selected advertisement at a specific time in the video.
[0436] Input: A list of ads related to the scene
[0437] Output: Video data with ads inserted
[0438] Specific operation: The server reconstructs the video data and places the selected advertisement at a specific time so that it is naturally incorporated.
[0439] Step 12:
[0440] The device displays relevant advertisements at designated times in relation to the video being played.
[0441] Input: Video data with ads inserted
[0442] Output: Videos and ads watched by users
[0443] Specific operation: The user's device plays the reconstructed video and displays the advertisement at the specified time.
[0444] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0445] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0446] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0447] [Second embodiment]
[0448] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0449] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0450] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0451] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0452] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0453] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0454] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0455] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0456] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0457] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0458] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0459] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0460] This invention is a system that displays advertisements for web page articles and video content at optimal timing based on not only the user's attribute information but also the context of the content itself. This system operates as follows between the server, terminal, and user.
[0461] 1. Get and split articles
[0462] The server retrieves the articles from the web pages that the user visits.
[0463] The server divides the retrieved article into sentences, for example, "Doing moderate exercise every day is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0464] 2. Context Analysis
[0465] The server uses a generative AI model to analyze the context of each sentence.
[0466] The server obtains topic and sentiment information from the generative AI model. For example, it determines that the first sentence is a "topic about exercise" and the second sentence is a "topic about diet."
[0467] 3. Ad selection
[0468] Based on the analysis results, the server searches for relevant advertisements from an advertisement database.
[0469] The server selects advertisements for fitness gyms for sentences about "exercise" and advertisements for health foods for sentences about "diet."
[0470] 4. Restructuring the article
[0471] The server incorporates the selected advertisements into the original article to generate a reconstructed article.
[0472] The server sends this article to the user's terminal.
[0473] 5. Display of articles and advertisements
[0474] The terminal displays the received reconstructed article on the user's browser.
[0475] The device displays advertisements related to each sentence at the appropriate time.
[0476] For example, if a user is viewing an article about health, an ad for a fitness gym could be displayed immediately after the sentence, "Doing moderate exercise every day is good for your physical and mental health," and an ad for health foods could be displayed immediately after the sentence, "A balanced diet will boost your immune system." This allows users to see highly relevant ads in a natural way, increasing the effectiveness of the ads.
[0477] Video analysis and ad insertion
[0478] As the video plays, the server analyzes each individual scene in the video using a generative AI model.
[0479] The server selects the appropriate advertisement for each scene and generates instructions to insert the advertisement at a specific timing in the video.
[0480] For example, if a user is watching a cooking recipe video, an advertisement related to the content of the scene will be inserted, such as an advertisement for cooking utensils, immediately after the cooking scene ends. This allows users to see advertisements that are relevant to the content of the video, increasing the effectiveness of the advertisements.
[0481] This invention is a system that optimizes the display of advertisements to users by taking into account the relevance of the content of web articles and videos as a whole, thereby improving the user experience and enabling advertisers to expect better results.
[0482] The processing flow will be explained below.
[0483] Step 1:
[0484] A user visits a specific web page.
[0485] Step 2:
[0486] The device (user's browser) sends the request to the server.
[0487] Step 3:
[0488] The server receives the request and retrieves the article from the target web page.
[0489] Step 4:
[0490] The server divides the text data of the acquired article into sentences.
[0491] Step 5:
[0492] The server feeds each sentence into a generative AI model and analyzes the context, which includes information such as topic and sentiment.
[0493] Step 6:
[0494] Based on the analysis results obtained from the generative AI model, the server searches the advertising database for advertisements related to each sentence.
[0495] Step 7:
[0496] The server selects advertisements for fitness gyms for sentences about "exercise" and advertisements for health foods for sentences about "diet."
[0497] Step 8:
[0498] The server lists the selected ads and determines the best time and place to display them.
[0499] Step 9:
[0500] The server reconstructs the article so that it includes advertisements and transmits it to the user's terminal.
[0501] Step 10:
[0502] The terminal displays the retrieved reconstructed article on the user's browser.
[0503] Step 11:
[0504] The device displays relevant ads immediately after each sentence, allowing users to see relevant ads in a natural way as they continue reading the article.
[0505] Steps for Video
[0506] Step 1:
[0507] A user begins watching a particular video.
[0508] Step 2:
[0509] The terminal sends the request to the server.
[0510] Step 3:
[0511] The server receives the request and prepares the video for analysis.
[0512] Step 4:
[0513] The server uses a generative AI model to analyze the context of each scene in the video.
[0514] Step 5:
[0515] Based on the analysis results, the server searches the advertisement database for advertisements related to each scene.
[0516] Step 6:
[0517] The server selects advertisements for cooking utensils for the "cooking scene" and advertisements for sports equipment for the "exercise scene."
[0518] Step 7:
[0519] The server creates timing instructions containing advertisements and reconfigures them to appear at specific moments in the video.
[0520] Step 8:
[0521] The server transmits the reconstructed video data to the user's terminal.
[0522] Step 9:
[0523] The device will display relevant ads at designated times while the video is playing, allowing users to see ads that are highly relevant to the content of the video in a natural way.
[0524] Example 1
[0525] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0526] With conventional methods of displaying advertisements on web pages and video content, advertisements are often displayed uniformly based on user attribute information, and are not properly associated with the context or scene of the content, which results in a lack of user interest and low advertising effectiveness. Furthermore, if the timing of advertisement insertion is inappropriate, the user experience may be impaired.
[0527] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0528] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences; means for analyzing the context of each sentence using a generative AI model; means for selecting relevant advertisements based on the analysis results and creating an advertisement list; means for inserting the selected advertisements into the article at appropriate times and locations to reconstruct it; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in the video using a generative AI model while the video is being played; means for selecting advertisements related to the scenes based on the analysis results and displaying them at specific times in the video; and means for displaying the reconstructed article received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times. This makes it possible to display advertisements that interest the user based on the context of the content, improving the effectiveness of the advertisements and the user experience.
[0529] A "server" is a central computer system that manages the data of web pages and video content accessed by users and performs various processing.
[0530] A "terminal" is a device used by a user to display web pages, reconstructed articles, and video content.
[0531] A "user" is a person who accesses and views web pages and video content.
[0532] An "article" is the text content on a web page, the information that users view.
[0533] A "sentence" is the smallest text unit that makes up an article, and is a sentence that completes one idea.
[0534] "Context" refers to the context and background of a particular sentence, and is an element that is analyzed by generative AI models.
[0535] A "generative AI model" is an artificial intelligence model for natural language processing and text generation, and is used for contextual analysis of sentences and advertisement selection.
[0536] An "advertisement" is a visual or audio message used by a business or service to promote its products or information.
[0537] An "ad list" is a collection of multiple advertisements selected based on the analyzed context.
[0538] A "reconstructed article" is a newly created text content created by appropriately inserting selected advertisements into the original article.
[0539] A "video scene" refers to a specific scene or segment within video content, and is the unit used as the basis for selecting advertisements.
[0540] "Timing" refers to the right moment or time to display an ad, and is an important factor in ensuring a good user experience.
[0541] This invention is a system for optimizing the display of advertisements on web pages and video content, and operates mainly between a server, a terminal, and a user as follows.
[0542] Configuration and Processing Overview
[0543] Retrieving and splitting articles
[0544] The server retrieves articles from web pages accessed by users. The retrieved articles are saved in HTML or text format and broken down into sentences. Specifically, natural language processing libraries such as nltk and spaCy can be used with programming languages such as Python. For example, the following can be done using a script:
[0545] python
[0546] import spacy
[0547] nlp = spacy.load("en_core_web_sm")
[0548] text = "Moderate exercise every day is good for your physical and mental health. A balanced diet strengthens your immune system."
[0549] doc = nlp(text)
[0550] sentences = [sent.text for sent in doc.sents]
[0551] print(sentences) Output: ['Daily moderate exercise is good for your physical and mental health.', 'A balanced diet strengthens your immune system.']
[0552] Contextual Analysis
[0553] The server uses a generative AI model (specifically, OpenAI's GPT-4) to analyze the context of each sentence. This analysis extracts the topic and sentiment information of each sentence. For example, for the sentence "Doing moderate exercise every day is good for your physical and mental health," the server sends the following prompt to the generative AI model:
[0554] Sample prompt: "Analyze the context of the following sentence to extract topic and emotional information. Sentence: 'Doing moderate exercise every day is good for your physical and mental health.'"
[0555] Based on the response from the generative AI model, "exercise topics" and "positive sentiment" are obtained.
[0556] Ad selection
[0557] The server then searches for relevant advertisements from its advertising database based on the analysis results. The advertisements are selected based on specific topics and emotions. For example, for the context of "exercise," it would select advertisements for fitness gyms.
[0558] Generate and submit restructured articles
[0559] The server then embeds the selected ads into the original article, generating a reconstructed article by inserting the ads at the appropriate places and times. The reconstructed article is then saved in HTML format and sent to the user's device.
[0560] View articles and ads
[0561] The device then displays the received reconstructed article in the user's browser. Specifically, it uses an HTML rendering engine to display the reconstructed article and advertisements, allowing the user to view the advertisements in a natural way.
[0562] Application to video content
[0563] The server analyzes scenes in the video as they play and selects appropriate ads. It uses a generative AI model to analyze the context of the scene and determine the timing of ad insertion. For example, while playing a cooking recipe video, it can generate instructions such as "display an ad for cooking utensils after the cooking scene."
[0564] Sample prompt: "Select the appropriate ad based on the content of the following video scene. Scene description: 'A scene introducing a cooking recipe.'"
[0565] Specific examples
[0566] If a user is viewing an article about health, an ad for a fitness gym could be displayed immediately after the sentence, "Doing moderate exercise every day is good for your physical and mental health," and an ad for health foods could be displayed immediately after the sentence, "A balanced diet will boost your immune system." This allows the user to naturally see more relevant ads, improving the effectiveness of the ads.
[0567] In the case of video content, if a user is watching a cooking recipe video, advertisements that match the content of the scene will be inserted, such as an advertisement for cooking utensils, immediately after the cooking scene ends.
[0568] As described above, this system utilizes generative AI models to effectively display ads that match users' interests based on context, thereby improving user experience and advertising effectiveness.
[0569] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0570] Step 1:
[0571] Retrieving articles
[0572] The server retrieves HTML content from the web page accessed by the user. It extracts the article portion from the retrieved HTML content and stores it in text format. The input is the web page URL, and the output is the extracted article text. Specifically, it uses the requests or BeautifulSoup library to retrieve the required data from the web page.
[0573] Step 2:
[0574] Article division
[0575] The server splits the retrieved articles into sentences. The input is the extracted article text, and the output is a list of split sentences. Specifically, it splits sentences using natural language processing libraries such as nltk and spaCy.
[0576] Step 3:
[0577] Contextual Analysis
[0578] The server uses a generative AI model to analyze the context of each sentence. The input is a list of segmented sentences, and the output is a list of sentences with topic and emotional information added. Specifically, the server sends the following prompt to the generative AI model:
[0579] Sample prompt: "Analyze the context of the following sentence to extract topic and emotional information. Sentence: 'Doing moderate exercise every day is good for your physical and mental health.'"
[0580] Based on the response from the generative AI model, topic and sentiment information is assigned to each sentence.
[0581] Step 4:
[0582] Ad selection
[0583] The server selects relevant advertisements based on the analysis results and creates an advertisement list. The input is a list of sentences with topic and emotion information, and the output is a list of advertisements. Specifically, it searches an advertisement database and selects the most suitable advertisement. For example, it matches a fitness gym advertisement to a sentence with the topic "exercise."
[0584] Step 5:
[0585] Generate restructured articles
[0586] The server inserts the selected advertisements into the article at the appropriate time and place to generate a reconstructed article. The input is the advertisement list and the original article text, and the output is the reconstructed article text. Specifically, the server inserts advertisements after each sentence.
[0587] Step 6:
[0588] Submitting a restructured article
[0589] The server sends the reconstructed article to the user's device. The input is the reconstructed article text, and the output is a notification of completion of transmission to the user's device. Specifically, the data is sent using the HTTP protocol.
[0590] Step 7:
[0591] View articles and ads
[0592] The terminal displays the received reconstructed article in the user's browser. The input is the reconstructed article text, and the output is the web page displayed to the user. Specifically, it uses an HTML rendering engine to display the article and advertisements in the browser.
[0593] Step 8:
[0594] Video scene analysis
[0595] The server analyzes individual scenes in the video as it plays. The input is the video data, and the output is the analyzed scene information. Specifically, it uses a generative AI model to analyze the context of the scene.
[0596] Sample prompt: "Select the appropriate ad based on the content of the following video scene. Scene description: 'A scene introducing a cooking recipe.'"
[0597] Step 9:
[0598] Ad selection and insertion instruction generation
[0599] The server selects appropriate advertisements based on the analyzed scene information and generates instructions to insert the advertisements at specific times in the video. The input is the analyzed scene information and an advertisement database, and the output is insertion instructions. Specifically, the server selects the optimal advertisement for each scene and determines the appropriate timing.
[0600] The above is the processing flow of this system. By clearly indicating the specific operations and inputs / outputs performed at each step, the processing flow of the entire system becomes clear.
[0601] (Application example 1)
[0602] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0603] Conventional advertising display systems have difficulty displaying ads that are relevant to the context of the content being viewed by the user or to specific scenes at the appropriate time. Furthermore, there is a problem of reduced advertising effectiveness due to the lack of relevance between the content of articles or videos and the ads. Furthermore, this can lead to a poor user experience, making it difficult for advertisers to provide effective ads.
[0604] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0605] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; means for reconstructing the article by inserting the selected advertisements into the article at appropriate times and places; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in a video accessed by the user using a generative AI model while the video is being played; and means for selecting advertisements related to scenes in the video based on the analysis results and displaying them at specific times in the video. This makes it possible to display advertisements that are highly relevant to the content viewed by the user at appropriate times.
[0606] A "server" is a computer system that provides data and resources for joint use over a network.
[0607] "User" means any person, organization, or device that uses a computer system or network service.
[0608] A "web page" is an information page published on the Internet, written in a markup language such as HTML, and viewable using a browser.
[0609] "Article" refers to a single work or document published in a newspaper, magazine, web page, etc.
[0610] A "sentence" is the smallest unit of a sentence or piece of writing, a series of words that has complete meaning.
[0611] "Context" refers to the relationship between adjacent words or sentences, and the overall situation in which it is placed.
[0612] A "generative AI model" is an artificial intelligence model trained using neural networks on large datasets, and is capable of generating, understanding, and analyzing text.
[0613] "Advertising" refers to information created to introduce products and services and encourage consumers to purchase them.
[0614] "Video" refers to moving video content that is made up of a series of images and sounds.
[0615] A "scene" is a single continuous action or event in video content that does not span a specific time period.
[0616] A "list" is a data structure that lists and organizes multiple items.
[0617] As an embodiment of the present invention, a system that performs the following processing steps is constructed.
[0618] First, the server retrieves the article from the web page accessed by the user. The retrieved article is then divided into sentences. The context of each sentence is then analyzed using a generative AI model. For example, the natural language processing model "GPT-4" is used to extract the topic and sentiment information of each sentence.
[0619] After the context analysis is completed, the server selects relevant advertisements from the advertisement database based on the analysis results, lists the selected advertisements, and reconstructs the article based on them. The reconstructed article is then sent to the user's device.
[0620] Next, while the video being viewed by the user is being played, the server analyzes each scene in the video using a generative AI model. It selects the appropriate ad for each scene and generates instructions to display it at a specific moment in the video. These instructions are then sent to the user's device, and the ad is displayed at the appropriate time.
[0621] The specific configuration for realizing this system involves running the following program. First, the server retrieves articles from a webpage using the Python libraries "requests" and "BeautifulSoup," and then splits the articles into sentences using the "split_by_sentence" function. Next, contextual analysis is performed using the generative AI model "GPT-4." The analysis is performed by sending a prompt sentence using the "GPT-4" API and obtaining a response to it. The following prompt sentence is used as an example of this analysis.
[0622] "Analyze the context of this sentence and identify the topic and emotion: Regular exercise every day is good for your physical and mental health."
[0623] "Analyze the context of this sentence and identify the topic and emotion: A balanced diet boosts your immunity."
[0624] Based on the analysis results, the server selects the most suitable advertisement from the advertisement database and reconstructs the article, which is then sent to the user's device and displayed in the browser.
[0625] For video analysis, each scene in the video is identified and the context of each scene is analyzed using a generative AI model. For example, if a user is watching a cooking recipe video, an advertisement for cooking utensils will be displayed immediately after the cooking scene ends. For context analysis for each scene, a prompt in the form of "Analyze the context of this scene and identify related advertisements:" is used.
[0626] In this way, it is possible to build a system that dynamically displays highly relevant advertisements at the appropriate time based on the content the user is viewing, thereby improving the user experience and maximizing advertising effectiveness.
[0627] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0628] Step 1:
[0629] The server retrieves the article of the web page accessed by the user. As input, it receives the URL of the web page accessed by the user. As output, it returns the text data of the entire article of the web page. Specifically, the server uses the "requests" library to retrieve the HTML of the web page and "BeautifulSoup" to extract the text data.
[0630] Step 2:
[0631] The server splits the retrieved articles into sentences. As input, it receives the text data of the entire article. As output, it returns a list of text split into sentences. Specifically, the server uses Python's string manipulation functions to split the article using delimiters such as "."
[0632] Step 3:
[0633] The server analyzes the context of each sentence using a generative AI model. As input, it receives a text list divided into sentences. As output, it returns the context information (topic and sentiment information) of each sentence. Specifically, the server sends a prompt sentence to the API of the generative AI model "GPT-4" and obtains the analysis results.
[0634] Step 4:
[0635] The server selects relevant advertisements based on the results of the context analysis. As input, it receives the context information of each sentence. As output, it returns a list of relevant advertisements. Specifically, the server matches the context information with an advertisement database to select the most suitable advertisement.
[0636] Step 5:
[0637] The server reconstructs the article by inserting the selected advertisements at the appropriate time and place. As input, it receives the segmented sentences and a list of selected advertisements. As output, it returns the reconstructed article with the advertisements inserted. Specifically, the server executes logic to insert advertisements at the appropriate place for each sentence.
[0638] Step 6:
[0639] The server sends the reconstructed article to the user's device. As input, it receives the reconstructed article. As output, it sends the article to the user's device. Specifically, the server uses an HTTP request to display the reconstructed article in the user's browser.
[0640] Step 7:
[0641] The server analyzes the context of each individual scene in a video while the video accessed by the user is being played. As input, it receives video data and timestamps. As output, it returns contextual information for each scene. Specifically, the server uses a generative AI model to analyze each scene in the video.
[0642] Step 8:
[0643] The server selects advertisements relevant to the scenes based on the analysis results and generates instructions to display them at specific times in the video. It receives contextual information for each scene as input. It returns instructions to display the advertisements as output. Specifically, the server compares the analysis results for each scene with an advertisement database, selects appropriate advertisements, and sends instructions to the user's device for the timing of their display.
[0644] Step 9:
[0645] The device displays the reconstructed article in the user's browser and displays advertisements corresponding to each sentence at the appropriate time. As input, it receives the reconstructed article and advertisement display instructions received from the server. As output, it displays the article and advertisements in the user's browser. Specifically, the device controls the browser display using HTML and JavaScript and displays advertisements at the appropriate time.
[0646] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0647] This invention is a system that displays advertisements at optimal times based on not only user attribute information but also the context of the content itself and the user's emotional state in web page articles and video content. By adding an emotion engine, this system can recognize user emotions and further improve the accuracy and effectiveness of advertisement display.
[0648] 1. Get and split articles
[0649] When a user accesses a specific web page, the device (user's browser) sends the request to the server. The server receives the request and retrieves the article from the target web page. The server then divides the retrieved article into sentences. For example, it may divide the sentences into "Daily moderate exercise is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0650] 2. Context analysis and emotional state recognition
[0651] The server uses a generative AI model to analyze the context of each sentence. Context includes topic and emotional information. The server also uses an emotion engine to detect the user's emotional state. For example, the server can analyze facial expressions and tone of voice using a camera or microphone.
[0652] 3. Ad selection
[0653] Based on the analysis results, the server searches the advertisement database for advertisements related to each sentence and the user's emotional state. For example, if the sentence is related to "exercise" and the user is in a "positive" emotional state, an advertisement for a fitness gym will be selected. The analysis results of the emotion engine play an important role in advertisement selection.
[0654] 4. Restructure and submit your article
[0655] The server then incorporates the selected advertisements into the original article to generate a reconstructed article, which is then sent to the user's device. The reconstructed article also includes advertisements based on the user's emotional state.
[0656] 5. Display of articles and advertisements
[0657] The device displays the reconstructed article in the user's browser. The device displays relevant advertisements immediately after each sentence. For example, a fitness gym advertisement is displayed after the sentence "Daily moderate exercise is good for your physical and mental health," while a health food advertisement is displayed after the sentence "A balanced diet strengthens your immune system." As the user continues reading the article, they can naturally see highly relevant advertisements based on their emotional state and context.
[0658] Video analysis and ad insertion
[0659] When a user starts watching a particular video, the device sends the request to the server, which receives the request and prepares the video for analysis. The server uses a generative AI model to analyze the context of each scene in the video and an emotion engine to recognize the user's emotional state.
[0660] Based on the analysis results, the server searches the advertisement database for advertisements relevant to each scene and the user's emotional state. For example, if a cooking scene is displayed and the user finds it "interesting," an advertisement for cooking utensils will be selected. The server creates timing instructions including the selected advertisement and reconstructs the video to be displayed at a specific moment. The reconstructed video data is sent to the user's device. The device displays relevant advertisements at the specified timing for the video being played, allowing the user to naturally view advertisements based on the content of the video and the user's emotional state.
[0661] This invention is a system that can take into account the user's emotional state, increase the effectiveness of advertising, and improve the user experience by combining advertising display related to the content of web articles and entire videos with an emotion engine.
[0662] The processing flow will be explained below.
[0663] Step 1:
[0664] A user visits a specific web page.
[0665] Step 2:
[0666] The device (user's browser) sends the request to the server.
[0667] Step 3:
[0668] The server receives the request and retrieves the article from the target web page.
[0669] Step 4:
[0670] The server divides the text data of the acquired article into sentences.
[0671] Step 5:
[0672] The server feeds each sentence into a generative AI model to analyze the context, which includes topic and sentiment information.
[0673] Step 6:
[0674] The server uses an emotion engine to detect the user's emotional state by analyzing the user's facial expressions and tone of voice using the device's camera and microphone.
[0675] Step 7:
[0676] Based on the analysis results of the generative AI model and the results of the emotion engine, the server selects advertisements from the advertising database that are relevant to each sentence and the user's emotional state.
[0677] Step 8:
[0678] The server lists the selected ads and determines the best time and place to display them.
[0679] Step 9:
[0680] The server reconstructs the article so that it includes advertisements and transmits it to the user's terminal.
[0681] Step 10:
[0682] The terminal displays the received reconstructed article on the user's browser.
[0683] Step 11:
[0684] The device displays relevant ads immediately after each sentence, allowing users to naturally see relevant ads based on their context and emotional state as they continue reading the article.
[0685] Steps for Video
[0686] Step 1:
[0687] A user begins watching a particular video.
[0688] Step 2:
[0689] The terminal sends the request to the server.
[0690] Step 3:
[0691] The server receives the request and prepares the video for analysis.
[0692] Step 4:
[0693] The server uses a generative AI model to analyze the context of each scene in the video.
[0694] Step 5:
[0695] The server uses an emotion engine to set the user's emotional state by analyzing the user's emotional response to the video they are watching.
[0696] Step 6:
[0697] The server searches the advertising database for advertisements relevant to the scene based on the analysis results obtained from the generative AI model and emotion engine.
[0698] Step 7:
[0699] The server selects advertisements for cooking tools for the "cooking scene" and advertisements for sports equipment for the "exercise scene." The server adjusts the content and timing based on the user's emotional state.
[0700] Step 8:
[0701] The server creates timing instructions containing advertisements and reconfigures them to appear at specific moments in the video.
[0702] Step 9:
[0703] The server transmits the reconstructed video data to the user's terminal.
[0704] Step 10:
[0705] The device will display relevant ads at designated times while the video is playing, allowing users to see ads that are highly relevant to the content of the video in a natural way.
[0706] Example 2
[0707] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0708] Conventional advertising systems typically display ads based on user attribute information, but this method fails to fully consider each user's emotional state or the context of the content, resulting in ineffective ad display. Furthermore, even in video content, the timing and content of ad insertions often fail to match the user's emotions, resulting in a poor user experience. It is necessary to provide a system that can resolve these issues and display more effective and relevant ads to users.
[0709] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0710] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; means for acquiring emotional information using an emotion engine that recognizes the user's emotional state; means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; means for reconstructing the article by inserting the selected advertisements into the article at appropriate times and locations; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in the video using a generative AI model while the video is being played; and means for selecting advertisements related to the scenes based on the analysis results and displaying them at specific times in the video. This enables more effective and relevant advertisements to be displayed based on the user's emotional state and the context of the content.
[0711] "User" means an individual who accesses the Website or Application and uses or consumes its content.
[0712] A "device" is a device, such as a computer, smartphone, or tablet, that a user uses to connect to the Internet and view web pages and videos.
[0713] A "server" is a computer system that stores, manages, and transmits and receives data over a network, and is a device that provides content in response to user requests.
[0714] A "web page" is a unit of information that is published on the Internet and can be viewed by users through a browser; it is a digital page that consists of text, images, videos, etc.
[0715] An "article" is text-based information content posted on a web page, and may be in the form of news, blogs, commentaries, etc.
[0716] A "sentence" is a complete sentence in an article, the smallest meaningful unit of text.
[0717] A "generative AI model" is a machine learning model that uses artificial intelligence to generate, analyze, translate, and perform other tasks on text, and is based on natural language processing technology.
[0718] "Context" refers to the meaning or background that a particular part of a sentence or image has in relation to other parts or the whole.
[0719] "Analysis" is the process of breaking down data or information to understand its components and patterns.
[0720] An "emotion engine" is a technology that analyzes external information such as a user's facial expressions and tone of voice to determine the user's emotional state.
[0721] "Emotional state" refers to the type and intensity of emotion a user feels in a particular situation, and includes states such as "positive," "negative," and "interesting."
[0722] "Advertising" is promotional content that presents information about products and services and stimulates consumers' desire to purchase them.
[0723] A "list" refers to an ordered arrangement of related items.
[0724] "Timing" refers to the moment in time when a particular event or action occurs.
[0725] A "location" refers to a specific location within a web page or video.
[0726] "Reconstruction" means constructing the original content in a new form by inserting new elements or rearranging them.
[0727] "Video" is digital media content that combines a sequence of images and sounds to tell a visual and audio story.
[0728] A "scene" refers to a specific sequence of images and sounds in a video, and is a unit that depicts a specific situation or activity.
[0729] This invention is a system that displays advertisements at optimal times based on not only user attribute information but also the context of the content itself and the user's emotional state in web page articles and video content. By adding an emotion engine, this system can recognize user emotions and further improve the accuracy and effectiveness of advertisement display.
[0730] First, when a user accesses a specific web page, their device (the user's browser) sends the request to a server. For example, if a user accesses a news site to read an article about health, this request is sent. The server receives the request and retrieves the article from the target web page. Next, the server uses a generative AI model (for example, a model using natural language processing technology) to divide the retrieved article into sentences. Specifically, it divides the article into sentences such as "Daily moderate exercise is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0731] The server uses a generative AI model to analyze the context of each sentence. This context includes topic and emotional information. The server also uses an emotion engine (e.g., technology that analyzes facial expressions and vocal tones) to recognize the user's emotional state. Specifically, the server analyzes facial expressions and vocal tones captured in real time through the device's camera and microphone to determine whether the user is in a "positive," "negative," "interested," or other emotional state.
[0732] Next, the server searches the advertisement database for advertisements related to each sentence and the user's emotional state based on the analysis results. For example, if the sentence is related to "exercise" and the user is in a "positive" emotional state, advertisements for fitness gyms will be selected. The selected advertisements will be listed.
[0733] The server then incorporates the selected advertisements into the original article to generate a reconstructed article. For example, it might insert an advertisement for a fitness gym after the sentence, "Doing moderate exercise every day is good for your physical and mental health." The reconstructed article data is then sent to the terminal.
[0734] The device then displays the reconstructed article in the user's browser. Relevant advertisements are displayed immediately after each sentence, allowing users to naturally see highly relevant advertisements based on their emotional state and context as they read the article. For example, an advertisement for a health food product is displayed after the sentence, "A balanced diet boosts your immune system."
[0735] When a user starts watching a specific video, the device sends a request to the server. The server receives the request and prepares the video for analysis. The server uses a generative AI model to analyze the context of each scene in the video and an emotion engine to recognize the user's emotional state and suggest advertisements. The specific analysis takes into account the content displayed in each scene of the video and the user's reactions (facial expressions, tone of voice, etc.).
[0736] For example, if a user finds a cooking scene "interesting," an advertisement for cooking utensils will be selected. The server reconstructs the selected advertisement to be displayed at a specific moment in the video, and sends the reconstructed video data to the device. The device then displays the relevant advertisement at a specified timing during video playback. While watching the video, the user can naturally see advertisements based on their emotional state and the context of the video.
[0737] For example, the following prompt might be used: "Explain how to analyze the article content of a web page a user views and use an emotion engine to display relevant ads."
[0738] By implementing this system, it will be possible to display more effective and relevant advertisements that take into consideration the user's emotional state and the context of the content, which is expected to improve the effectiveness of advertisements and the user experience.
[0739] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0740] Step 1:
[0741] When a user accesses a specific web page, the device (the user's browser) sends the request to the server. The input is the URL of the web page being accessed, and the output is the article data for the web page. In concrete terms, when a user accesses the "health" category on a news site, the browser requests "https: / / example.com / health" from the server. The server analyzes the request and retrieves the article data from the specified URL.
[0742] Step 2:
[0743] The server splits the retrieved article data into sentences. The input is the article data of the web page, and the output is text data split into sentences. Specifically, the server uses the "Sentence Splitter API" to split the article text into sentences such as "Doing moderate exercise every day is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0744] Step 3:
[0745] The server uses a generative AI model (e.g., a model using natural language processing technology) to analyze the context of each sentence. The input is text data divided into sentences, and the output is the context data for each sentence. Specifically, the server sends each sentence to the generative AI model, which extracts the topics "health" and "exercise" and positive sentiment information from the sentence "exercise is good for your health."
[0746] Step 4:
[0747] The server recognizes the user's emotional state using an emotion engine (e.g., technology that analyzes facial expressions and voice tones). The input is real-time data acquired through the device's camera and microphone, and the output is the user's emotional state data. Specifically, the server analyzes the user's facial expressions and voice tone to determine whether the user's emotional state is "positive," "negative," "interested," etc.
[0748] Step 5:
[0749] The server searches for relevant advertisements from the advertisement database based on the analysis results (contextual data and emotional state data). The input is contextual data and emotional state data, and the output is a list of selected advertisements. Specifically, the server selects advertisements for fitness gyms if the sentence is related to "exercise" and the user is in a "positive" emotional state.
[0750] Step 6:
[0751] The server incorporates the selected advertisements into the original article and generates a reconstructed article. The input is text data divided into sentences and a list of advertisements, and the output is the reconstructed article data. Specifically, the server adds an advertisement for a fitness gym after the sentence, "Doing moderate exercise every day is good for your physical and mental health."
[0752] Step 7:
[0753] The server sends the reconstructed article data to the terminal. The input is the reconstructed article data, and the output is the data sent to the user's terminal. In concrete terms, the server sends the reconstructed article data to the terminal and it is displayed in the specified browser.
[0754] Step 8:
[0755] The terminal displays the received reconstructed article in the user's browser. The input is the reconstructed article data, and the output is the information displayed in the user's browser. In concrete terms, the terminal displays the reconstructed article in the browser, and an advertisement for a health food is displayed after the sentence, "A balanced diet boosts your immune system."
[0756] Step 9:
[0757] When a user starts watching a particular video, the device sends a request to the server. The input is the URL of the video being watched, and the output is the video data. In concrete terms, when a user requests "https: / / example.com / cooking" to watch a cooking tutorial video, the device sends this information to the server.
[0758] Step 10:
[0759] The server receives the request and prepares the target video for analysis. The input is the video data, and the output is data for each scene in the video. Specifically, the server divides the video into frames and identifies each scene.
[0760] Step 11:
[0761] The server uses a generative AI model to analyze the context of each scene in the video. The input is the video data for each scene, and the output is the contextual data for each scene. Specifically, the server sends each scene to the generative AI model, identifies the context, and extracts information such as "cooking scene" or "exercise scene."
[0762] Step 12:
[0763] The server uses an emotion engine to recognize the user's emotional state. The input is real-time data acquired through the device's camera and microphone, and the output is the user's emotional state data. Specifically, the server analyzes the user's facial expressions and tone of voice while watching a video to determine whether the user is in an "interesting" emotional state.
[0764] Step 13:
[0765] The server searches for relevant advertisements from the advertisement database based on the analysis results (scene context data and emotional state data). The input is the scene context data and emotional state data, and the output is a list of selected advertisements. Specifically, when a cooking scene is displayed and the user feels that it is "interesting," the server selects an advertisement for cooking utensils.
[0766] Step 14:
[0767] The server inserts the selected advertisements into specific scenes of the video and reconstructs it. The input is the video data for each scene and the advertisement list, and the output is the reconstructed video data. Specifically, the server inserts advertisements at the specified timing and reconstructs the video.
[0768] Step 15:
[0769] The server transmits the reconstructed video data to the terminal. The input is the reconstructed video data, and the output is the data to be transmitted to the user's terminal. In concrete terms, the server transmits the reconstructed video data to the terminal.
[0770] Step 16:
[0771] The device displays relevant advertisements at specified times during the video being played. The input is the reconstructed video data, and the output is a video with advertisements that is displayed to the user. Specifically, the device displays advertisements for cooking utensils during specific scenes in the video.
[0772] (Application example 2)
[0773] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0774] Conventional advertising systems for web pages and videos display ads based on user attribute information and the general context of the content, but this does not fully consider the user's emotional state or the specific context of the moment. As a result, ads are often ineffective or not perceived as natural by users. Furthermore, even when using smart glasses or head-mounted displays, there is a lack of technology that can analyze the user's emotional state in real time and display appropriate ads. This has led to a demand for improved user experience.
[0775] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0776] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; sentiment analysis means for recognizing an emotional state from the context of the acquired articles; means for selecting relevant advertisements based on the analysis results and the emotional state and creating a list of advertisements; means for reconstructing the articles by inserting the selected advertisements into the articles at appropriate times and places; means for transmitting the reconstructed articles to the user's terminal; and means for displaying the reconstructed articles received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times. This enables more effective and natural advertisement display based on the user's emotional state and the specific context of the content.
[0777] A "User" is an individual who accesses a particular web page or video content.
[0778] A "web page" is a unit of information that is published on the Internet and consists of text, images, videos, etc.
[0779] An "article" refers to a sentence or text content that is published on a web page.
[0780] A "sentence" is a unit of text that has an independent meaning.
[0781] A "generative AI model" is an artificial intelligence model used to perform natural language processing and contextual analysis.
[0782] "Emotion analysis means" is a technology for recognizing a user's emotional state from facial expressions, tone of voice, etc.
[0783] "Advertising" refers to messages or content that convey information about products or services and encourage desired behavior.
[0784] A "list" refers to an ordered arrangement of related items.
[0785] "Timing" refers to the point in time at which a particular event or action should occur.
[0786] "Reconstruction" refers to inserting advertisements into existing content and rearranging it into a new form.
[0787] A "terminal" is a device used by a user, such as a computer, smart glasses, or head-mounted display.
[0788] "Video" is a media format that expresses movement by playing a series of still images.
[0789] A "scene" refers to a part of a video that depicts a specific situation or situation.
[0790] "Smart glasses" are a wearable device in the form of glasses worn by the user that can display information and analyze emotions.
[0791] A "head-mounted display" is a display device worn on the head that provides visual information.
[0792] "Real time" means that things happen at the same speed as real time.
[0793] To implement this invention, a system including a server, a terminal, and a user is required. The system can display advertisements related to the content of articles and videos on a web page at optimal timing.
[0794] The server retrieves articles from web pages accessed by users. The retrieved articles are divided into sentences, and the context of each sentence is analyzed using a generative AI model. At the same time, the user's emotional state is recognized using a sentiment analysis method. Based on the analysis results and the user's emotional state, relevant ads are selected and a list of ads is created.
[0795] The selected advertisements are inserted into the article at the appropriate time and place, and the article is reconstructed. This reconstructed article is then sent to the user's device. The user's device displays the received reconstructed article, and the advertisements corresponding to each sentence are displayed at the appropriate time.
[0796] Furthermore, while the video is playing, the server analyzes the context of each scene in the video using a generative AI model. The user's emotional state is recognized from the scene in the video, and an advertisement related to the scene is selected based on the analysis results and the emotional state. The selected advertisement is displayed at a specific time in the video.
[0797] When using smart glasses or head-mounted displays, the user's emotional state is analyzed using a built-in camera and microphone, and advertisements are selected and displayed in real time based on the emotional state and content context.
[0798] The hardware used includes smart glasses or a head-mounted display, a built-in camera and microphone, and the software used includes a sentiment analysis engine, a generative AI model (e.g., GPT-4), a web server (e.g., Node.js), and an advertising database (e.g., MongoDB).
[0799] As a concrete example, when a user is reading an article on a web page that reads, "Doing moderate exercise every day is good for your physical and mental health," the generative AI model analyzes the context of this sentence, and if the emotion engine recognizes the user's "positive" emotional state, an advertisement for a "fitness gym" will be selected. The generative AI model is analyzed using the following prompt sentence:
[0800] Example prompt sentence:
[0801] Analyze the context and emotional state of the following sentences and select the appropriate ad.
[0802] Document: Regular exercise every day is good for your physical and mental health.
[0803] Emotion: Positive
[0804] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0805] Step 1:
[0806] The server receives a request for the web page accessed by the user and retrieves the article from the web page.
[0807] Input: User request (Web page URL)
[0808] Output: Text data of the retrieved article
[0809] Specific operation: The server analyzes the HTTP request, retrieves the web page content from the specified URL, extracts the text data of the article, and passes it to the next step.
[0810] Step 2:
[0811] The server divides the retrieved article into sentences.
[0812] Input: Article text data
[0813] Output: A list of text data split into sentences
[0814] How it works: The server uses a text analysis algorithm to split the article into sentences. You can use a Python natural language processing library, for example.
[0815] Step 3:
[0816] The server analyzes the context of each sentence using a generative AI model.
[0817] Input: A list of text data split into sentences
[0818] Output: Analysis result data on the context of each sentence
[0819] How it works: It calls an API to operate the generative AI model and analyzes the context of each sentence, generating structured data that includes the topic and sentiment information of the sentence.
[0820] Step 4:
[0821] The server utilizes emotion analysis means to recognize the user's emotional state.
[0822] Input: User's facial expression data (camera) or voice data (microphone)
[0823] Output: User's emotional state (e.g. positive, negative, interesting, etc.)
[0824] How it works: The server sends data from the camera and microphone to the emotion analysis engine, which analyzes the user's current emotional state. The emotion analysis engine identifies the user's emotional state based on their facial expressions and tone of voice.
[0825] Step 5:
[0826] The server selects relevant advertisements from an advertisement database based on the analysis results and the emotional state.
[0827] Input: Contextual information for each sentence and the user's emotional state
[0828] Output: A list of relevant ads
[0829] Specific operations: Query the advertisement database and select advertisements that match the context of the sentence and the user's emotional state. Retrieve information about the selected advertisements in list form.
[0830] Step 6:
[0831] The server reconstructs the article by inserting the selected advertisements at the appropriate time and place.
[0832] Input: Context information for each sentence, list of ads
[0833] Output: Reconstructed article
[0834] Specific operation: The server inserts advertisements into the original text data and restructures the article so that the advertisements are displayed immediately after the sentence or at natural times along the flow of the content.
[0835] Step 7:
[0836] The server sends the reconstructed article to the user's terminal.
[0837] Input: Reconstructed article
[0838] Output: Send article to user's device
[0839] What it does: Converts the reconstructed article into HTML or another appropriate format and sends it to the user's browser or device.
[0840] Step 8:
[0841] The terminal displays the received reconstructed article and displays advertisements corresponding to each sentence at appropriate times.
[0842] Input: Reconstructed article
[0843] Output: Articles and advertisements displayed in the browser
[0844] Specific operation: The user's device receives the reconstructed article and displays it in a web browser, with relevant advertisements displayed immediately after each sentence of the article.
[0845] Step 9:
[0846] As the video plays, the server analyzes the context of each scene in the video using a generative AI model.
[0847] Input: Video file
[0848] Output: Context analysis results of video scenes
[0849] How it works: The server analyzes the video data frame by frame and understands the context of each scene based on a generative AI model.
[0850] Step 10:
[0851] The server recognizes the emotional state from scenes in the video and selects advertisements related to the scenes based on the analysis results and the emotional state.
[0852] Input: Context analysis results of video scenes, user emotional state
[0853] Output: A list of ads related to the scene
[0854] Specific operation: The server uses an emotion analysis engine to analyze the user's emotional state in real time and selects advertisements from the advertisement database that correspond to the context of the video scene.
[0855] Step 11:
[0856] The server displays the selected advertisement at a specific time in the video.
[0857] Input: A list of ads related to the scene
[0858] Output: Video data with ads inserted
[0859] Specific operation: The server reconstructs the video data and places the selected advertisement at a specific time so that it is naturally incorporated.
[0860] Step 12:
[0861] The device displays relevant advertisements at designated times in relation to the video being played.
[0862] Input: Video data with ads inserted
[0863] Output: Videos and ads watched by users
[0864] Specific operation: The user's device plays the reconstructed video and displays the advertisement at the specified time.
[0865] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0866] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0867] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0868] [Third embodiment]
[0869] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0870] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0871] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0872] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0873] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0874] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0875] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0876] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0877] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0878] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0879] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0880] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0881] This invention is a system that displays advertisements for web page articles and video content at optimal timing based on not only the user's attribute information but also the context of the content itself. This system operates as follows between the server, terminal, and user.
[0882] 1. Get and split articles
[0883] The server retrieves the articles from the web pages that the user visits.
[0884] The server divides the retrieved article into sentences, for example, "Doing moderate exercise every day is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[0885] 2. Context Analysis
[0886] The server uses a generative AI model to analyze the context of each sentence.
[0887] The server obtains topic and sentiment information from the generative AI model. For example, it determines that the first sentence is a "topic about exercise" and the second sentence is a "topic about diet."
[0888] 3. Ad selection
[0889] Based on the analysis results, the server searches for relevant advertisements from an advertisement database.
[0890] The server selects advertisements for fitness gyms for sentences about "exercise" and advertisements for health foods for sentences about "diet."
[0891] 4. Restructuring the article
[0892] The server incorporates the selected advertisements into the original article to generate a reconstructed article.
[0893] The server sends this article to the user's terminal.
[0894] 5. Display of articles and advertisements
[0895] The terminal displays the received reconstructed article on the user's browser.
[0896] The device displays advertisements related to each sentence at the appropriate time.
[0897] For example, if a user is viewing an article about health, an ad for a fitness gym could be displayed immediately after the sentence, "Doing moderate exercise every day is good for your physical and mental health," and an ad for health foods could be displayed immediately after the sentence, "A balanced diet will boost your immune system." This allows users to see highly relevant ads in a natural way, increasing the effectiveness of the ads.
[0898] Video analysis and ad insertion
[0899] As the video plays, the server analyzes each individual scene in the video using a generative AI model.
[0900] The server selects the appropriate advertisement for each scene and generates instructions to insert the advertisement at a specific timing in the video.
[0901] For example, if a user is watching a cooking recipe video, an advertisement related to the content of the scene will be inserted, such as an advertisement for cooking utensils, immediately after the cooking scene ends. This allows users to see advertisements that are relevant to the content of the video, increasing the effectiveness of the advertisements.
[0902] This invention is a system that optimizes the display of advertisements to users by taking into account the relevance of the content of web articles and videos as a whole, thereby improving the user experience and enabling advertisers to expect better results.
[0903] The processing flow will be explained below.
[0904] Step 1:
[0905] A user visits a specific web page.
[0906] Step 2:
[0907] The device (user's browser) sends the request to the server.
[0908] Step 3:
[0909] The server receives the request and retrieves the article from the target web page.
[0910] Step 4:
[0911] The server divides the text data of the acquired article into sentences.
[0912] Step 5:
[0913] The server feeds each sentence into a generative AI model and analyzes the context, which includes information such as topic and sentiment.
[0914] Step 6:
[0915] Based on the analysis results obtained from the generative AI model, the server searches the advertising database for advertisements related to each sentence.
[0916] Step 7:
[0917] The server selects advertisements for fitness gyms for sentences about "exercise" and advertisements for health foods for sentences about "diet."
[0918] Step 8:
[0919] The server lists the selected ads and determines the best time and place to display them.
[0920] Step 9:
[0921] The server reconstructs the article so that it includes advertisements and transmits it to the user's terminal.
[0922] Step 10:
[0923] The terminal displays the retrieved reconstructed article on the user's browser.
[0924] Step 11:
[0925] The device displays relevant ads immediately after each sentence, allowing users to see relevant ads in a natural way as they continue reading the article.
[0926] Steps for Video
[0927] Step 1:
[0928] A user begins watching a particular video.
[0929] Step 2:
[0930] The terminal sends the request to the server.
[0931] Step 3:
[0932] The server receives the request and prepares the video for analysis.
[0933] Step 4:
[0934] The server uses a generative AI model to analyze the context of each scene in the video.
[0935] Step 5:
[0936] Based on the analysis results, the server searches the advertisement database for advertisements related to each scene.
[0937] Step 6:
[0938] The server selects advertisements for cooking utensils for the "cooking scene" and advertisements for sports equipment for the "exercise scene."
[0939] Step 7:
[0940] The server creates timing instructions containing advertisements and reconfigures them to appear at specific moments in the video.
[0941] Step 8:
[0942] The server transmits the reconstructed video data to the user's terminal.
[0943] Step 9:
[0944] The device will display relevant ads at designated times while the video is playing, allowing users to see ads that are highly relevant to the content of the video in a natural way.
[0945] Example 1
[0946] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0947] With conventional methods of displaying advertisements on web pages and video content, advertisements are often displayed uniformly based on user attribute information, and are not properly associated with the context or scene of the content, which results in a lack of user interest and low advertising effectiveness. Furthermore, if the timing of advertisement insertion is inappropriate, the user experience may be impaired.
[0948] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0949] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences; means for analyzing the context of each sentence using a generative AI model; means for selecting relevant advertisements based on the analysis results and creating an advertisement list; means for inserting the selected advertisements into the article at appropriate times and locations to reconstruct it; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in the video using a generative AI model while the video is being played; means for selecting advertisements related to the scenes based on the analysis results and displaying them at specific times in the video; and means for displaying the reconstructed article received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times. This makes it possible to display advertisements that interest the user based on the context of the content, improving the effectiveness of the advertisements and the user experience.
[0950] A "server" is a central computer system that manages the data of web pages and video content accessed by users and performs various processing.
[0951] A "terminal" is a device used by a user to display web pages, reconstructed articles, and video content.
[0952] A "user" is a person who accesses and views web pages and video content.
[0953] An "article" is the text content on a web page, the information that users view.
[0954] A "sentence" is the smallest text unit that makes up an article, and is a sentence that completes one idea.
[0955] "Context" refers to the context and background of a particular sentence, and is an element that is analyzed by generative AI models.
[0956] A "generative AI model" is an artificial intelligence model for natural language processing and text generation, and is used for contextual analysis of sentences and advertisement selection.
[0957] An "advertisement" is a visual or audio message used by a business or service to promote its products or information.
[0958] An "ad list" is a collection of multiple advertisements selected based on the analyzed context.
[0959] A "reconstructed article" is a newly created text content created by appropriately inserting selected advertisements into the original article.
[0960] A "video scene" refers to a specific scene or segment within video content, and is the unit used as the basis for selecting advertisements.
[0961] "Timing" refers to the right moment or time to display an ad, and is an important factor in ensuring a good user experience.
[0962] This invention is a system for optimizing the display of advertisements on web pages and video content, and operates mainly between a server, a terminal, and a user as follows.
[0963] Configuration and Processing Overview
[0964] Retrieving and splitting articles
[0965] The server retrieves articles from web pages accessed by users. The retrieved articles are saved in HTML or text format and broken down into sentences. Specifically, natural language processing libraries such as nltk and spaCy can be used with programming languages such as Python. For example, the following can be done using a script:
[0966] python
[0967] import spacy
[0968] nlp = spacy.load("en_core_web_sm")
[0969] text = "Moderate exercise every day is good for your physical and mental health. A balanced diet strengthens your immune system."
[0970] doc = nlp(text)
[0971] sentences = [sent.text for sent in doc.sents]
[0972] print(sentences) Output: ['Daily moderate exercise is good for your physical and mental health.', 'A balanced diet strengthens your immune system.']
[0973] Contextual Analysis
[0974] The server uses a generative AI model (specifically, OpenAI's GPT-4) to analyze the context of each sentence. This analysis extracts the topic and sentiment information of each sentence. For example, for the sentence "Doing moderate exercise every day is good for your physical and mental health," the server sends the following prompt to the generative AI model:
[0975] Sample prompt: "Analyze the context of the following sentence to extract topic and emotional information. Sentence: 'Doing moderate exercise every day is good for your physical and mental health.'"
[0976] Based on the response from the generative AI model, "exercise topics" and "positive sentiment" are obtained.
[0977] Ad selection
[0978] The server then searches for relevant advertisements from its advertising database based on the analysis results. The advertisements are selected based on specific topics and emotions. For example, for the context of "exercise," it would select advertisements for fitness gyms.
[0979] Generate and submit restructured articles
[0980] The server then embeds the selected ads into the original article, generating a reconstructed article by inserting the ads at the appropriate places and times. The reconstructed article is then saved in HTML format and sent to the user's device.
[0981] View articles and ads
[0982] The device then displays the received reconstructed article in the user's browser. Specifically, it uses an HTML rendering engine to display the reconstructed article and advertisements, allowing the user to view the advertisements in a natural way.
[0983] Application to video content
[0984] The server analyzes scenes in the video as they play and selects appropriate ads. It uses a generative AI model to analyze the context of the scene and determine the timing of ad insertion. For example, while playing a cooking recipe video, it can generate instructions such as "display an ad for cooking utensils after the cooking scene."
[0985] Sample prompt: "Select the appropriate ad based on the content of the following video scene. Scene description: 'A scene introducing a cooking recipe.'"
[0986] Specific examples
[0987] If a user is viewing an article about health, an ad for a fitness gym could be displayed immediately after the sentence, "Doing moderate exercise every day is good for your physical and mental health," and an ad for health foods could be displayed immediately after the sentence, "A balanced diet will boost your immune system." This allows the user to naturally see more relevant ads, improving the effectiveness of the ads.
[0988] In the case of video content, if a user is watching a cooking recipe video, advertisements that match the content of the scene will be inserted, such as an advertisement for cooking utensils, immediately after the cooking scene ends.
[0989] As described above, this system utilizes generative AI models to effectively display ads that match users' interests based on context, thereby improving user experience and advertising effectiveness.
[0990] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0991] Step 1:
[0992] Retrieving articles
[0993] The server retrieves HTML content from the web page accessed by the user. It extracts the article portion from the retrieved HTML content and stores it in text format. The input is the web page URL, and the output is the extracted article text. Specifically, it uses the requests or BeautifulSoup library to retrieve the required data from the web page.
[0994] Step 2:
[0995] Article division
[0996] The server splits the retrieved articles into sentences. The input is the extracted article text, and the output is a list of split sentences. Specifically, it splits sentences using natural language processing libraries such as nltk and spaCy.
[0997] Step 3:
[0998] Contextual Analysis
[0999] The server uses a generative AI model to analyze the context of each sentence. The input is a list of segmented sentences, and the output is a list of sentences with topic and emotional information added. Specifically, the server sends the following prompt to the generative AI model:
[1000] Sample prompt: "Analyze the context of the following sentence to extract topic and emotional information. Sentence: 'Doing moderate exercise every day is good for your physical and mental health.'"
[1001] Based on the response from the generative AI model, topic and sentiment information is assigned to each sentence.
[1002] Step 4:
[1003] Ad selection
[1004] The server selects relevant advertisements based on the analysis results and creates an advertisement list. The input is a list of sentences with topic and emotion information, and the output is a list of advertisements. Specifically, it searches an advertisement database and selects the most suitable advertisement. For example, it matches a fitness gym advertisement to a sentence with the topic "exercise."
[1005] Step 5:
[1006] Generate restructured articles
[1007] The server inserts the selected advertisements into the article at the appropriate time and place to generate a reconstructed article. The input is the advertisement list and the original article text, and the output is the reconstructed article text. Specifically, the server inserts advertisements after each sentence.
[1008] Step 6:
[1009] Submitting a restructured article
[1010] The server sends the reconstructed article to the user's device. The input is the reconstructed article text, and the output is a notification of completion of transmission to the user's device. Specifically, the data is sent using the HTTP protocol.
[1011] Step 7:
[1012] View articles and ads
[1013] The terminal displays the received reconstructed article in the user's browser. The input is the reconstructed article text, and the output is the web page displayed to the user. Specifically, it uses an HTML rendering engine to display the article and advertisements in the browser.
[1014] Step 8:
[1015] Video scene analysis
[1016] The server analyzes individual scenes in the video as it plays. The input is the video data, and the output is the analyzed scene information. Specifically, it uses a generative AI model to analyze the context of the scene.
[1017] Sample prompt: "Select the appropriate ad based on the content of the following video scene. Scene description: 'A scene introducing a cooking recipe.'"
[1018] Step 9:
[1019] Ad selection and insertion instruction generation
[1020] The server selects appropriate advertisements based on the analyzed scene information and generates instructions to insert the advertisements at specific times in the video. The input is the analyzed scene information and an advertisement database, and the output is insertion instructions. Specifically, the server selects the optimal advertisement for each scene and determines the appropriate timing.
[1021] The above is the processing flow of this system. By clearly indicating the specific operations and inputs / outputs performed at each step, the processing flow of the entire system becomes clear.
[1022] (Application example 1)
[1023] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1024] Conventional advertising display systems have difficulty displaying ads that are relevant to the context of the content being viewed by the user or to specific scenes at the appropriate time. Furthermore, there is a problem of reduced advertising effectiveness due to the lack of relevance between the content of articles or videos and the ads. Furthermore, this can lead to a poor user experience, making it difficult for advertisers to provide effective ads.
[1025] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1026] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; means for reconstructing the article by inserting the selected advertisements into the article at appropriate times and places; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in a video accessed by the user using a generative AI model while the video is being played; and means for selecting advertisements related to scenes in the video based on the analysis results and displaying them at specific times in the video. This makes it possible to display advertisements that are highly relevant to the content viewed by the user at appropriate times.
[1027] A "server" is a computer system that provides data and resources for joint use over a network.
[1028] "User" means any person, organization, or device that uses a computer system or network service.
[1029] A "web page" is an information page published on the Internet, written in a markup language such as HTML, and viewable using a browser.
[1030] "Article" refers to a single work or document published in a newspaper, magazine, web page, etc.
[1031] A "sentence" is the smallest unit of a sentence or piece of writing, a series of words that has complete meaning.
[1032] "Context" refers to the relationship between adjacent words or sentences, and the overall situation in which it is placed.
[1033] A "generative AI model" is an artificial intelligence model trained using neural networks on large datasets, and is capable of generating, understanding, and analyzing text.
[1034] "Advertising" refers to information created to introduce products and services and encourage consumers to purchase them.
[1035] "Video" refers to moving video content that is made up of a series of images and sounds.
[1036] A "scene" is a single continuous action or event in video content that does not span a specific time period.
[1037] A "list" is a data structure that lists and organizes multiple items.
[1038] As an embodiment of the present invention, a system that performs the following processing steps is constructed.
[1039] First, the server retrieves the article from the web page accessed by the user. The retrieved article is then divided into sentences. The context of each sentence is then analyzed using a generative AI model. For example, the natural language processing model "GPT-4" is used to extract the topic and sentiment information of each sentence.
[1040] After the context analysis is completed, the server selects relevant advertisements from the advertisement database based on the analysis results, lists the selected advertisements, and reconstructs the article based on them. The reconstructed article is then sent to the user's device.
[1041] Next, while the video being viewed by the user is being played, the server analyzes each scene in the video using a generative AI model. It selects the appropriate ad for each scene and generates instructions to display it at a specific moment in the video. These instructions are then sent to the user's device, and the ad is displayed at the appropriate time.
[1042] The specific configuration for realizing this system involves running the following program. First, the server retrieves articles from a webpage using the Python libraries "requests" and "BeautifulSoup," and then splits the articles into sentences using the "split_by_sentence" function. Next, contextual analysis is performed using the generative AI model "GPT-4." The analysis is performed by sending a prompt sentence using the "GPT-4" API and obtaining a response to it. The following prompt sentence is used as an example of this analysis.
[1043] "Analyze the context of this sentence and identify the topic and emotion: Regular exercise every day is good for your physical and mental health."
[1044] "Analyze the context of this sentence and identify the topic and emotion: A balanced diet boosts your immunity."
[1045] Based on the analysis results, the server selects the most suitable advertisement from the advertisement database and reconstructs the article, which is then sent to the user's device and displayed in the browser.
[1046] For video analysis, each scene in the video is identified and the context of each scene is analyzed using a generative AI model. For example, if a user is watching a cooking recipe video, an advertisement for cooking utensils will be displayed immediately after the cooking scene ends. For context analysis for each scene, a prompt in the form of "Analyze the context of this scene and identify related advertisements:" is used.
[1047] In this way, it is possible to build a system that dynamically displays highly relevant advertisements at the appropriate time based on the content the user is viewing, thereby improving the user experience and maximizing advertising effectiveness.
[1048] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1049] Step 1:
[1050] The server retrieves the article of the web page accessed by the user. As input, it receives the URL of the web page accessed by the user. As output, it returns the text data of the entire article of the web page. Specifically, the server uses the "requests" library to retrieve the HTML of the web page and "BeautifulSoup" to extract the text data.
[1051] Step 2:
[1052] The server splits the retrieved articles into sentences. As input, it receives the text data of the entire article. As output, it returns a list of text split into sentences. Specifically, the server uses Python's string manipulation functions to split the article using delimiters such as "."
[1053] Step 3:
[1054] The server analyzes the context of each sentence using a generative AI model. As input, it receives a text list divided into sentences. As output, it returns the context information (topic and sentiment information) of each sentence. Specifically, the server sends a prompt sentence to the API of the generative AI model "GPT-4" and obtains the analysis results.
[1055] Step 4:
[1056] The server selects relevant advertisements based on the results of the context analysis. As input, it receives the context information of each sentence. As output, it returns a list of relevant advertisements. Specifically, the server matches the context information with an advertisement database to select the most suitable advertisement.
[1057] Step 5:
[1058] The server reconstructs the article by inserting the selected advertisements at the appropriate time and place. As input, it receives the segmented sentences and a list of selected advertisements. As output, it returns the reconstructed article with the advertisements inserted. Specifically, the server executes logic to insert advertisements at the appropriate place for each sentence.
[1059] Step 6:
[1060] The server sends the reconstructed article to the user's device. As input, it receives the reconstructed article. As output, it sends the article to the user's device. Specifically, the server uses an HTTP request to display the reconstructed article in the user's browser.
[1061] Step 7:
[1062] The server analyzes the context of each individual scene in a video while the video accessed by the user is being played. As input, it receives video data and timestamps. As output, it returns contextual information for each scene. Specifically, the server uses a generative AI model to analyze each scene in the video.
[1063] Step 8:
[1064] The server selects advertisements relevant to the scenes based on the analysis results and generates instructions to display them at specific times in the video. It receives contextual information for each scene as input. It returns instructions to display the advertisements as output. Specifically, the server compares the analysis results for each scene with an advertisement database, selects appropriate advertisements, and sends instructions to the user's device for the timing of their display.
[1065] Step 9:
[1066] The device displays the reconstructed article in the user's browser and displays advertisements corresponding to each sentence at the appropriate time. As input, it receives the reconstructed article and advertisement display instructions received from the server. As output, it displays the article and advertisements in the user's browser. Specifically, the device controls the browser display using HTML and JavaScript and displays advertisements at the appropriate time.
[1067] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1068] This invention is a system that displays advertisements at optimal times based on not only user attribute information but also the context of the content itself and the user's emotional state in web page articles and video content. By adding an emotion engine, this system can recognize user emotions and further improve the accuracy and effectiveness of advertisement display.
[1069] 1. Get and split articles
[1070] When a user accesses a specific web page, the device (user's browser) sends the request to the server. The server receives the request and retrieves the article from the target web page. The server then divides the retrieved article into sentences. For example, it may divide the sentences into "Daily moderate exercise is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[1071] 2. Context analysis and emotional state recognition
[1072] The server uses a generative AI model to analyze the context of each sentence. Context includes topic and emotional information. The server also uses an emotion engine to detect the user's emotional state. For example, the server can analyze facial expressions and tone of voice using a camera or microphone.
[1073] 3. Ad selection
[1074] Based on the analysis results, the server searches the advertisement database for advertisements related to each sentence and the user's emotional state. For example, if the sentence is related to "exercise" and the user is in a "positive" emotional state, an advertisement for a fitness gym will be selected. The analysis results of the emotion engine play an important role in advertisement selection.
[1075] 4. Restructure and submit your article
[1076] The server then incorporates the selected advertisements into the original article to generate a reconstructed article, which is then sent to the user's device. The reconstructed article also includes advertisements based on the user's emotional state.
[1077] 5. Display of articles and advertisements
[1078] The device displays the reconstructed article in the user's browser. The device displays relevant advertisements immediately after each sentence. For example, a fitness gym advertisement is displayed after the sentence "Daily moderate exercise is good for your physical and mental health," while a health food advertisement is displayed after the sentence "A balanced diet strengthens your immune system." As the user continues reading the article, they can naturally see highly relevant advertisements based on their emotional state and context.
[1079] Video analysis and ad insertion
[1080] When a user starts watching a particular video, the device sends the request to the server, which receives the request and prepares the video for analysis. The server uses a generative AI model to analyze the context of each scene in the video and an emotion engine to recognize the user's emotional state.
[1081] Based on the analysis results, the server searches the advertisement database for advertisements relevant to each scene and the user's emotional state. For example, if a cooking scene is displayed and the user finds it "interesting," an advertisement for cooking utensils will be selected. The server creates timing instructions including the selected advertisement and reconstructs the video to be displayed at a specific moment. The reconstructed video data is sent to the user's device. The device displays relevant advertisements at the specified timing for the video being played, allowing the user to naturally view advertisements based on the content of the video and the user's emotional state.
[1082] This invention is a system that can take into account the user's emotional state, increase the effectiveness of advertising, and improve the user experience by combining advertising display related to the content of web articles and entire videos with an emotion engine.
[1083] The processing flow will be explained below.
[1084] Step 1:
[1085] A user visits a specific web page.
[1086] Step 2:
[1087] The device (user's browser) sends the request to the server.
[1088] Step 3:
[1089] The server receives the request and retrieves the article from the target web page.
[1090] Step 4:
[1091] The server divides the text data of the acquired article into sentences.
[1092] Step 5:
[1093] The server feeds each sentence into a generative AI model to analyze the context, which includes topic and sentiment information.
[1094] Step 6:
[1095] The server uses an emotion engine to detect the user's emotional state by analyzing the user's facial expressions and tone of voice using the device's camera and microphone.
[1096] Step 7:
[1097] Based on the analysis results of the generative AI model and the results of the emotion engine, the server selects advertisements from the advertising database that are relevant to each sentence and the user's emotional state.
[1098] Step 8:
[1099] The server lists the selected ads and determines the best time and place to display them.
[1100] Step 9:
[1101] The server reconstructs the article so that it includes advertisements and transmits it to the user's terminal.
[1102] Step 10:
[1103] The terminal displays the received reconstructed article on the user's browser.
[1104] Step 11:
[1105] The device displays relevant ads immediately after each sentence, allowing users to naturally see relevant ads based on their context and emotional state as they continue reading the article.
[1106] Steps for Video
[1107] Step 1:
[1108] A user begins watching a particular video.
[1109] Step 2:
[1110] The terminal sends the request to the server.
[1111] Step 3:
[1112] The server receives the request and prepares the video for analysis.
[1113] Step 4:
[1114] The server uses a generative AI model to analyze the context of each scene in the video.
[1115] Step 5:
[1116] The server uses an emotion engine to set the user's emotional state by analyzing the user's emotional response to the video they are watching.
[1117] Step 6:
[1118] The server searches the advertising database for advertisements relevant to the scene based on the analysis results obtained from the generative AI model and emotion engine.
[1119] Step 7:
[1120] The server selects advertisements for cooking tools for the "cooking scene" and advertisements for sports equipment for the "exercise scene." The server adjusts the content and timing based on the user's emotional state.
[1121] Step 8:
[1122] The server creates timing instructions containing advertisements and reconfigures them to appear at specific moments in the video.
[1123] Step 9:
[1124] The server transmits the reconstructed video data to the user's terminal.
[1125] Step 10:
[1126] The device will display relevant ads at designated times while the video is playing, allowing users to see ads that are highly relevant to the content of the video in a natural way.
[1127] Example 2
[1128] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1129] Conventional advertising systems typically display ads based on user attribute information, but this method fails to fully consider each user's emotional state or the context of the content, resulting in ineffective ad display. Furthermore, even in video content, the timing and content of ad insertions often fail to match the user's emotions, resulting in a poor user experience. It is necessary to provide a system that can resolve these issues and display more effective and relevant ads to users.
[1130] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1131] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; means for acquiring emotional information using an emotion engine that recognizes the user's emotional state; means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; means for reconstructing the article by inserting the selected advertisements into the article at appropriate times and locations; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in the video using a generative AI model while the video is being played; and means for selecting advertisements related to the scenes based on the analysis results and displaying them at specific times in the video. This enables more effective and relevant advertisements to be displayed based on the user's emotional state and the context of the content.
[1132] "User" means an individual who accesses the Website or Application and uses or consumes its content.
[1133] A "device" is a device, such as a computer, smartphone, or tablet, that a user uses to connect to the Internet and view web pages and videos.
[1134] A "server" is a computer system that stores, manages, and transmits and receives data over a network, and is a device that provides content in response to user requests.
[1135] A "web page" is a unit of information that is published on the Internet and can be viewed by users through a browser; it is a digital page that consists of text, images, videos, etc.
[1136] An "article" is text-based information content posted on a web page, and may be in the form of news, blogs, commentaries, etc.
[1137] A "sentence" is a complete sentence in an article, the smallest meaningful unit of text.
[1138] A "generative AI model" is a machine learning model that uses artificial intelligence to generate, analyze, translate, and perform other tasks on text, and is based on natural language processing technology.
[1139] "Context" refers to the meaning or background that a particular part of a sentence or image has in relation to other parts or the whole.
[1140] "Analysis" is the process of breaking down data or information to understand its components and patterns.
[1141] An "emotion engine" is a technology that analyzes external information such as a user's facial expressions and tone of voice to determine the user's emotional state.
[1142] "Emotional state" refers to the type and intensity of emotion a user feels in a particular situation, and includes states such as "positive," "negative," and "interesting."
[1143] "Advertising" is promotional content that presents information about products and services and stimulates consumers' desire to purchase them.
[1144] A "list" refers to an ordered arrangement of related items.
[1145] "Timing" refers to the moment in time when a particular event or action occurs.
[1146] A "location" refers to a specific location within a web page or video.
[1147] "Reconstruction" means constructing the original content in a new form by inserting new elements or rearranging them.
[1148] "Video" is digital media content that combines a sequence of images and sounds to tell a visual and audio story.
[1149] A "scene" refers to a specific sequence of images and sounds in a video, and is a unit that depicts a specific situation or activity.
[1150] This invention is a system that displays advertisements at optimal times based on not only user attribute information but also the context of the content itself and the user's emotional state in web page articles and video content. By adding an emotion engine, this system can recognize user emotions and further improve the accuracy and effectiveness of advertisement display.
[1151] First, when a user accesses a specific web page, their device (the user's browser) sends the request to a server. For example, if a user accesses a news site to read an article about health, this request is sent. The server receives the request and retrieves the article from the target web page. Next, the server uses a generative AI model (for example, a model using natural language processing technology) to divide the retrieved article into sentences. Specifically, it divides the article into sentences such as "Daily moderate exercise is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[1152] The server uses a generative AI model to analyze the context of each sentence. This context includes topic and emotional information. The server also uses an emotion engine (e.g., technology that analyzes facial expressions and vocal tones) to recognize the user's emotional state. Specifically, the server analyzes facial expressions and vocal tones captured in real time through the device's camera and microphone to determine whether the user is in a "positive," "negative," "interested," or other emotional state.
[1153] Next, the server searches the advertisement database for advertisements related to each sentence and the user's emotional state based on the analysis results. For example, if the sentence is related to "exercise" and the user is in a "positive" emotional state, advertisements for fitness gyms will be selected. The selected advertisements will be listed.
[1154] The server then incorporates the selected advertisements into the original article to generate a reconstructed article. For example, it might insert an advertisement for a fitness gym after the sentence, "Doing moderate exercise every day is good for your physical and mental health." The reconstructed article data is then sent to the terminal.
[1155] The device then displays the reconstructed article in the user's browser. Relevant advertisements are displayed immediately after each sentence, allowing users to naturally see highly relevant advertisements based on their emotional state and context as they read the article. For example, an advertisement for a health food product is displayed after the sentence, "A balanced diet boosts your immune system."
[1156] When a user starts watching a specific video, the device sends a request to the server. The server receives the request and prepares the video for analysis. The server uses a generative AI model to analyze the context of each scene in the video and an emotion engine to recognize the user's emotional state and suggest advertisements. The specific analysis takes into account the content displayed in each scene of the video and the user's reactions (facial expressions, tone of voice, etc.).
[1157] For example, if a user finds a cooking scene "interesting," an advertisement for cooking utensils will be selected. The server reconstructs the selected advertisement to be displayed at a specific moment in the video, and sends the reconstructed video data to the device. The device then displays the relevant advertisement at a specified timing during video playback. While watching the video, the user can naturally see advertisements based on their emotional state and the context of the video.
[1158] For example, the following prompt might be used: "Explain how to analyze the article content of a web page a user views and use an emotion engine to display relevant ads."
[1159] By implementing this system, it will be possible to display more effective and relevant advertisements that take into consideration the user's emotional state and the context of the content, which is expected to improve the effectiveness of advertisements and the user experience.
[1160] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1161] Step 1:
[1162] When a user accesses a specific web page, the device (the user's browser) sends the request to the server. The input is the URL of the web page being accessed, and the output is the article data for the web page. In concrete terms, when a user accesses the "health" category on a news site, the browser requests "https: / / example.com / health" from the server. The server analyzes the request and retrieves the article data from the specified URL.
[1163] Step 2:
[1164] The server splits the retrieved article data into sentences. The input is the article data of the web page, and the output is text data split into sentences. Specifically, the server uses the "Sentence Splitter API" to split the article text into sentences such as "Doing moderate exercise every day is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[1165] Step 3:
[1166] The server uses a generative AI model (e.g., a model using natural language processing technology) to analyze the context of each sentence. The input is text data divided into sentences, and the output is the context data for each sentence. Specifically, the server sends each sentence to the generative AI model, which extracts the topics "health" and "exercise" and positive sentiment information from the sentence "exercise is good for your health."
[1167] Step 4:
[1168] The server recognizes the user's emotional state using an emotion engine (e.g., technology that analyzes facial expressions and voice tones). The input is real-time data acquired through the device's camera and microphone, and the output is the user's emotional state data. Specifically, the server analyzes the user's facial expressions and voice tone to determine whether the user's emotional state is "positive," "negative," "interested," etc.
[1169] Step 5:
[1170] The server searches for relevant advertisements from the advertisement database based on the analysis results (contextual data and emotional state data). The input is contextual data and emotional state data, and the output is a list of selected advertisements. Specifically, the server selects advertisements for fitness gyms if the sentence is related to "exercise" and the user is in a "positive" emotional state.
[1171] Step 6:
[1172] The server incorporates the selected advertisements into the original article and generates a reconstructed article. The input is text data divided into sentences and a list of advertisements, and the output is the reconstructed article data. Specifically, the server adds an advertisement for a fitness gym after the sentence, "Doing moderate exercise every day is good for your physical and mental health."
[1173] Step 7:
[1174] The server sends the reconstructed article data to the terminal. The input is the reconstructed article data, and the output is the data sent to the user's terminal. In concrete terms, the server sends the reconstructed article data to the terminal and it is displayed in the specified browser.
[1175] Step 8:
[1176] The terminal displays the received reconstructed article in the user's browser. The input is the reconstructed article data, and the output is the information displayed in the user's browser. In concrete terms, the terminal displays the reconstructed article in the browser, and an advertisement for a health food is displayed after the sentence, "A balanced diet boosts your immune system."
[1177] Step 9:
[1178] When a user starts watching a particular video, the device sends a request to the server. The input is the URL of the video being watched, and the output is the video data. In concrete terms, when a user requests "https: / / example.com / cooking" to watch a cooking tutorial video, the device sends this information to the server.
[1179] Step 10:
[1180] The server receives the request and prepares the target video for analysis. The input is the video data, and the output is data for each scene in the video. Specifically, the server divides the video into frames and identifies each scene.
[1181] Step 11:
[1182] The server uses a generative AI model to analyze the context of each scene in the video. The input is the video data for each scene, and the output is the contextual data for each scene. Specifically, the server sends each scene to the generative AI model, identifies the context, and extracts information such as "cooking scene" or "exercise scene."
[1183] Step 12:
[1184] The server uses an emotion engine to recognize the user's emotional state. The input is real-time data acquired through the device's camera and microphone, and the output is the user's emotional state data. Specifically, the server analyzes the user's facial expressions and tone of voice while watching a video to determine whether the user is in an "interesting" emotional state.
[1185] Step 13:
[1186] The server searches for relevant advertisements from the advertisement database based on the analysis results (scene context data and emotional state data). The input is the scene context data and emotional state data, and the output is a list of selected advertisements. Specifically, when a cooking scene is displayed and the user feels that it is "interesting," the server selects an advertisement for cooking utensils.
[1187] Step 14:
[1188] The server inserts the selected advertisements into specific scenes of the video and reconstructs it. The input is the video data for each scene and the advertisement list, and the output is the reconstructed video data. Specifically, the server inserts advertisements at the specified timing and reconstructs the video.
[1189] Step 15:
[1190] The server transmits the reconstructed video data to the terminal. The input is the reconstructed video data, and the output is the data to be transmitted to the user's terminal. In concrete terms, the server transmits the reconstructed video data to the terminal.
[1191] Step 16:
[1192] The device displays relevant advertisements at specified times during the video being played. The input is the reconstructed video data, and the output is a video with advertisements that is displayed to the user. Specifically, the device displays advertisements for cooking utensils during specific scenes in the video.
[1193] (Application example 2)
[1194] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1195] Conventional advertising systems for web pages and videos display ads based on user attribute information and the general context of the content, but this does not fully consider the user's emotional state or the specific context of the moment. As a result, ads are often ineffective or not perceived as natural by users. Furthermore, even when using smart glasses or head-mounted displays, there is a lack of technology that can analyze the user's emotional state in real time and display appropriate ads. This has led to a demand for improved user experience.
[1196] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1197] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; sentiment analysis means for recognizing an emotional state from the context of the acquired articles; means for selecting relevant advertisements based on the analysis results and the emotional state and creating a list of advertisements; means for reconstructing the articles by inserting the selected advertisements into the articles at appropriate times and places; means for transmitting the reconstructed articles to the user's terminal; and means for displaying the reconstructed articles received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times. This enables more effective and natural advertisement display based on the user's emotional state and the specific context of the content.
[1198] A "User" is an individual who accesses a particular web page or video content.
[1199] A "web page" is a unit of information that is published on the Internet and consists of text, images, videos, etc.
[1200] An "article" refers to a sentence or text content that is published on a web page.
[1201] A "sentence" is a unit of text that has an independent meaning.
[1202] A "generative AI model" is an artificial intelligence model used to perform natural language processing and contextual analysis.
[1203] "Emotion analysis means" is a technology for recognizing a user's emotional state from facial expressions, tone of voice, etc.
[1204] "Advertising" refers to messages or content that convey information about products or services and encourage desired behavior.
[1205] A "list" refers to an ordered arrangement of related items.
[1206] "Timing" refers to the point in time at which a particular event or action should occur.
[1207] "Reconstruction" refers to inserting advertisements into existing content and rearranging it into a new form.
[1208] A "terminal" is a device used by a user, such as a computer, smart glasses, or head-mounted display.
[1209] "Video" is a media format that expresses movement by playing a series of still images.
[1210] A "scene" refers to a part of a video that depicts a specific situation or situation.
[1211] "Smart glasses" are a wearable device in the form of glasses worn by the user that can display information and analyze emotions.
[1212] A "head-mounted display" is a display device worn on the head that provides visual information.
[1213] "Real time" means that things happen at the same speed as real time.
[1214] To implement this invention, a system including a server, a terminal, and a user is required. The system can display advertisements related to the content of articles and videos on a web page at optimal timing.
[1215] The server retrieves articles from web pages accessed by users. The retrieved articles are divided into sentences, and the context of each sentence is analyzed using a generative AI model. At the same time, the user's emotional state is recognized using a sentiment analysis method. Based on the analysis results and the user's emotional state, relevant ads are selected and a list of ads is created.
[1216] The selected advertisements are inserted into the article at the appropriate time and place, and the article is reconstructed. This reconstructed article is then sent to the user's device. The user's device displays the received reconstructed article, and the advertisements corresponding to each sentence are displayed at the appropriate time.
[1217] Furthermore, while the video is playing, the server analyzes the context of each scene in the video using a generative AI model. The user's emotional state is recognized from the scene in the video, and an advertisement related to the scene is selected based on the analysis results and the emotional state. The selected advertisement is displayed at a specific time in the video.
[1218] When using smart glasses or head-mounted displays, the user's emotional state is analyzed using a built-in camera and microphone, and advertisements are selected and displayed in real time based on the emotional state and content context.
[1219] The hardware used includes smart glasses or a head-mounted display, a built-in camera and microphone, and the software used includes a sentiment analysis engine, a generative AI model (e.g., GPT-4), a web server (e.g., Node.js), and an advertising database (e.g., MongoDB).
[1220] As a concrete example, when a user is reading an article on a web page that reads, "Doing moderate exercise every day is good for your physical and mental health," the generative AI model analyzes the context of this sentence, and if the emotion engine recognizes the user's "positive" emotional state, an advertisement for a "fitness gym" will be selected. The generative AI model is analyzed using the following prompt sentence:
[1221] Example prompt sentence:
[1222] Analyze the context and emotional state of the following sentences and select the appropriate ad.
[1223] Document: Regular exercise every day is good for your physical and mental health.
[1224] Emotion: Positive
[1225] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1226] Step 1:
[1227] The server receives a request for the web page accessed by the user and retrieves the article from the web page.
[1228] Input: User request (Web page URL)
[1229] Output: Text data of the retrieved article
[1230] Specific operation: The server analyzes the HTTP request, retrieves the web page content from the specified URL, extracts the text data of the article, and passes it to the next step.
[1231] Step 2:
[1232] The server divides the retrieved article into sentences.
[1233] Input: Article text data
[1234] Output: A list of text data split into sentences
[1235] How it works: The server uses a text analysis algorithm to split the article into sentences. You can use a Python natural language processing library, for example.
[1236] Step 3:
[1237] The server analyzes the context of each sentence using a generative AI model.
[1238] Input: A list of text data split into sentences
[1239] Output: Analysis result data on the context of each sentence
[1240] How it works: It calls an API to operate the generative AI model and analyzes the context of each sentence, generating structured data that includes the topic and sentiment information of the sentence.
[1241] Step 4:
[1242] The server utilizes emotion analysis means to recognize the user's emotional state.
[1243] Input: User's facial expression data (camera) or voice data (microphone)
[1244] Output: User's emotional state (e.g. positive, negative, interesting, etc.)
[1245] How it works: The server sends data from the camera and microphone to the emotion analysis engine, which analyzes the user's current emotional state. The emotion analysis engine identifies the user's emotional state based on their facial expressions and tone of voice.
[1246] Step 5:
[1247] The server selects relevant advertisements from an advertisement database based on the analysis results and the emotional state.
[1248] Input: Contextual information for each sentence and the user's emotional state
[1249] Output: A list of relevant ads
[1250] Specific operations: Query the advertisement database and select advertisements that match the context of the sentence and the user's emotional state. Retrieve information about the selected advertisements in list form.
[1251] Step 6:
[1252] The server reconstructs the article by inserting the selected advertisements at the appropriate time and place.
[1253] Input: Context information for each sentence, list of ads
[1254] Output: Reconstructed article
[1255] Specific operation: The server inserts advertisements into the original text data and restructures the article so that the advertisements are displayed immediately after the sentence or at natural times along the flow of the content.
[1256] Step 7:
[1257] The server sends the reconstructed article to the user's terminal.
[1258] Input: Reconstructed article
[1259] Output: Send article to user's device
[1260] What it does: Converts the reconstructed article into HTML or another appropriate format and sends it to the user's browser or device.
[1261] Step 8:
[1262] The terminal displays the received reconstructed article and displays advertisements corresponding to each sentence at appropriate times.
[1263] Input: Reconstructed article
[1264] Output: Articles and advertisements displayed in the browser
[1265] Specific operation: The user's device receives the reconstructed article and displays it in a web browser, with relevant advertisements displayed immediately after each sentence of the article.
[1266] Step 9:
[1267] As the video plays, the server analyzes the context of each scene in the video using a generative AI model.
[1268] Input: Video file
[1269] Output: Context analysis results of video scenes
[1270] How it works: The server analyzes the video data frame by frame and understands the context of each scene based on a generative AI model.
[1271] Step 10:
[1272] The server recognizes the emotional state from scenes in the video and selects advertisements related to the scenes based on the analysis results and the emotional state.
[1273] Input: Context analysis results of video scenes, user emotional state
[1274] Output: A list of ads related to the scene
[1275] Specific operation: The server uses an emotion analysis engine to analyze the user's emotional state in real time and selects advertisements from the advertisement database that correspond to the context of the video scene.
[1276] Step 11:
[1277] The server displays the selected advertisement at a specific time in the video.
[1278] Input: A list of ads related to the scene
[1279] Output: Video data with ads inserted
[1280] Specific operation: The server reconstructs the video data and places the selected advertisement at a specific time so that it is naturally incorporated.
[1281] Step 12:
[1282] The device displays relevant advertisements at designated times in relation to the video being played.
[1283] Input: Video data with ads inserted
[1284] Output: Videos and ads watched by users
[1285] Specific operation: The user's device plays the reconstructed video and displays the advertisement at the specified time.
[1286] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1287] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1288] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1289] [Fourth embodiment]
[1290] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1291] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1292] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1293] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1294] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1295] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1296] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1297] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1298] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1299] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1300] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1301] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1302] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1303] This invention is a system that displays advertisements for web page articles and video content at optimal timing based on not only the user's attribute information but also the context of the content itself. This system operates as follows between the server, terminal, and user.
[1304] 1. Get and split articles
[1305] The server retrieves the articles from the web pages that the user visits.
[1306] The server divides the retrieved article into sentences, for example, "Doing moderate exercise every day is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[1307] 2. Context Analysis
[1308] The server uses a generative AI model to analyze the context of each sentence.
[1309] The server obtains topic and sentiment information from the generative AI model. For example, it determines that the first sentence is a "topic about exercise" and the second sentence is a "topic about diet."
[1310] 3. Ad selection
[1311] Based on the analysis results, the server searches for relevant advertisements from an advertisement database.
[1312] The server selects advertisements for fitness gyms for sentences about "exercise" and advertisements for health foods for sentences about "diet."
[1313] 4. Restructuring the article
[1314] The server incorporates the selected advertisements into the original article to generate a reconstructed article.
[1315] The server sends this article to the user's terminal.
[1316] 5. Display of articles and advertisements
[1317] The terminal displays the received reconstructed article on the user's browser.
[1318] The device displays advertisements related to each sentence at the appropriate time.
[1319] For example, if a user is viewing an article about health, an ad for a fitness gym could be displayed immediately after the sentence, "Doing moderate exercise every day is good for your physical and mental health," and an ad for health foods could be displayed immediately after the sentence, "A balanced diet will boost your immune system." This allows users to see highly relevant ads in a natural way, increasing the effectiveness of the ads.
[1320] Video analysis and ad insertion
[1321] As the video plays, the server analyzes each individual scene in the video using a generative AI model.
[1322] The server selects the appropriate advertisement for each scene and generates instructions to insert the advertisement at a specific timing in the video.
[1323] For example, if a user is watching a cooking recipe video, an advertisement related to the content of the scene will be inserted, such as an advertisement for cooking utensils, immediately after the cooking scene ends. This allows users to see advertisements that are relevant to the content of the video, increasing the effectiveness of the advertisements.
[1324] This invention is a system that optimizes the display of advertisements to users by taking into account the relevance of the content of web articles and videos as a whole, thereby improving the user experience and enabling advertisers to expect better results.
[1325] The processing flow will be explained below.
[1326] Step 1:
[1327] A user visits a specific web page.
[1328] Step 2:
[1329] The device (user's browser) sends the request to the server.
[1330] Step 3:
[1331] The server receives the request and retrieves the article from the target web page.
[1332] Step 4:
[1333] The server divides the text data of the acquired article into sentences.
[1334] Step 5:
[1335] The server feeds each sentence into a generative AI model and analyzes the context, which includes information such as topic and sentiment.
[1336] Step 6:
[1337] Based on the analysis results obtained from the generative AI model, the server searches the advertising database for advertisements related to each sentence.
[1338] Step 7:
[1339] The server selects advertisements for fitness gyms for sentences about "exercise" and advertisements for health foods for sentences about "diet."
[1340] Step 8:
[1341] The server lists the selected ads and determines the best time and place to display them.
[1342] Step 9:
[1343] The server reconstructs the article so that it includes advertisements and transmits it to the user's terminal.
[1344] Step 10:
[1345] The terminal displays the retrieved reconstructed article on the user's browser.
[1346] Step 11:
[1347] The device displays relevant ads immediately after each sentence, allowing users to see relevant ads in a natural way as they continue reading the article.
[1348] Steps for Video
[1349] Step 1:
[1350] A user begins watching a particular video.
[1351] Step 2:
[1352] The terminal sends the request to the server.
[1353] Step 3:
[1354] The server receives the request and prepares the video for analysis.
[1355] Step 4:
[1356] The server uses a generative AI model to analyze the context of each scene in the video.
[1357] Step 5:
[1358] Based on the analysis results, the server searches the advertisement database for advertisements related to each scene.
[1359] Step 6:
[1360] The server selects advertisements for cooking utensils for the "cooking scene" and advertisements for sports equipment for the "exercise scene."
[1361] Step 7:
[1362] The server creates timing instructions containing advertisements and reconfigures them to appear at specific moments in the video.
[1363] Step 8:
[1364] The server transmits the reconstructed video data to the user's terminal.
[1365] Step 9:
[1366] The device will display relevant ads at designated times while the video is playing, allowing users to see ads that are highly relevant to the content of the video in a natural way.
[1367] Example 1
[1368] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1369] With conventional methods of displaying advertisements on web pages and video content, advertisements are often displayed uniformly based on user attribute information, and are not properly associated with the context or scene of the content, which results in a lack of user interest and low advertising effectiveness. Furthermore, if the timing of advertisement insertion is inappropriate, the user experience may be impaired.
[1370] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1371] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences; means for analyzing the context of each sentence using a generative AI model; means for selecting relevant advertisements based on the analysis results and creating an advertisement list; means for inserting the selected advertisements into the article at appropriate times and locations to reconstruct it; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in the video using a generative AI model while the video is being played; means for selecting advertisements related to the scenes based on the analysis results and displaying them at specific times in the video; and means for displaying the reconstructed article received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times. This makes it possible to display advertisements that interest the user based on the context of the content, improving the effectiveness of the advertisements and the user experience.
[1372] A "server" is a central computer system that manages the data of web pages and video content accessed by users and performs various processing.
[1373] A "terminal" is a device used by a user to display web pages, reconstructed articles, and video content.
[1374] A "user" is a person who accesses and views web pages and video content.
[1375] An "article" is the text content on a web page, the information that users view.
[1376] A "sentence" is the smallest text unit that makes up an article, and is a sentence that completes one idea.
[1377] "Context" refers to the context and background of a particular sentence, and is an element that is analyzed by generative AI models.
[1378] A "generative AI model" is an artificial intelligence model for natural language processing and text generation, and is used for contextual analysis of sentences and advertisement selection.
[1379] An "advertisement" is a visual or audio message used by a business or service to promote its products or information.
[1380] An "ad list" is a collection of multiple advertisements selected based on the analyzed context.
[1381] A "reconstructed article" is a newly created text content created by appropriately inserting selected advertisements into the original article.
[1382] A "video scene" refers to a specific scene or segment within video content, and is the unit used as the basis for selecting advertisements.
[1383] "Timing" refers to the right moment or time to display an ad, and is an important factor in ensuring a good user experience.
[1384] This invention is a system for optimizing the display of advertisements on web pages and video content, and operates mainly between a server, a terminal, and a user as follows.
[1385] Configuration and Processing Overview
[1386] Retrieving and splitting articles
[1387] The server retrieves articles from web pages accessed by users. The retrieved articles are saved in HTML or text format and broken down into sentences. Specifically, natural language processing libraries such as nltk and spaCy can be used with programming languages such as Python. For example, the following can be done using a script:
[1388] python
[1389] import spacy
[1390] nlp = spacy.load("en_core_web_sm")
[1391] text = "Moderate exercise every day is good for your physical and mental health. A balanced diet strengthens your immune system."
[1392] doc = nlp(text)
[1393] sentences = [sent.text for sent in doc.sents]
[1394] print(sentences) Output: ['Daily moderate exercise is good for your physical and mental health.', 'A balanced diet strengthens your immune system.']
[1395] Contextual Analysis
[1396] The server uses a generative AI model (specifically, OpenAI's GPT-4) to analyze the context of each sentence. This analysis extracts the topic and sentiment information of each sentence. For example, for the sentence "Doing moderate exercise every day is good for your physical and mental health," the server sends the following prompt to the generative AI model:
[1397] Sample prompt: "Analyze the context of the following sentence to extract topic and emotional information. Sentence: 'Doing moderate exercise every day is good for your physical and mental health.'"
[1398] Based on the response from the generative AI model, "exercise topics" and "positive sentiment" are obtained.
[1399] Ad selection
[1400] The server then searches for relevant advertisements from its advertising database based on the analysis results. The advertisements are selected based on specific topics and emotions. For example, for the context of "exercise," it would select advertisements for fitness gyms.
[1401] Generate and submit restructured articles
[1402] The server then embeds the selected ads into the original article, generating a reconstructed article by inserting the ads at the appropriate places and times. The reconstructed article is then saved in HTML format and sent to the user's device.
[1403] View articles and ads
[1404] The device then displays the received reconstructed article in the user's browser. Specifically, it uses an HTML rendering engine to display the reconstructed article and advertisements, allowing the user to view the advertisements in a natural way.
[1405] Application to video content
[1406] The server analyzes scenes in the video as they play and selects appropriate ads. It uses a generative AI model to analyze the context of the scene and determine the timing of ad insertion. For example, while playing a cooking recipe video, it can generate instructions such as "display an ad for cooking utensils after the cooking scene."
[1407] Sample prompt: "Select the appropriate ad based on the content of the following video scene. Scene description: 'A scene introducing a cooking recipe.'"
[1408] Specific examples
[1409] If a user is viewing an article about health, an ad for a fitness gym could be displayed immediately after the sentence, "Doing moderate exercise every day is good for your physical and mental health," and an ad for health foods could be displayed immediately after the sentence, "A balanced diet will boost your immune system." This allows the user to naturally see more relevant ads, improving the effectiveness of the ads.
[1410] In the case of video content, if a user is watching a cooking recipe video, advertisements that match the content of the scene will be inserted, such as an advertisement for cooking utensils, immediately after the cooking scene ends.
[1411] As described above, this system utilizes generative AI models to effectively display ads that match users' interests based on context, thereby improving user experience and advertising effectiveness.
[1412] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1413] Step 1:
[1414] Retrieving articles
[1415] The server retrieves HTML content from the web page accessed by the user. It extracts the article portion from the retrieved HTML content and stores it in text format. The input is the web page URL, and the output is the extracted article text. Specifically, it uses the requests or BeautifulSoup library to retrieve the required data from the web page.
[1416] Step 2:
[1417] Article division
[1418] The server splits the retrieved articles into sentences. The input is the extracted article text, and the output is a list of split sentences. Specifically, it splits sentences using natural language processing libraries such as nltk and spaCy.
[1419] Step 3:
[1420] Contextual Analysis
[1421] The server uses a generative AI model to analyze the context of each sentence. The input is a list of segmented sentences, and the output is a list of sentences with topic and emotional information added. Specifically, the server sends the following prompt to the generative AI model:
[1422] Sample prompt: "Analyze the context of the following sentence to extract topic and emotional information. Sentence: 'Doing moderate exercise every day is good for your physical and mental health.'"
[1423] Based on the response from the generative AI model, topic and sentiment information is assigned to each sentence.
[1424] Step 4:
[1425] Ad selection
[1426] The server selects relevant advertisements based on the analysis results and creates an advertisement list. The input is a list of sentences with topic and emotion information, and the output is a list of advertisements. Specifically, it searches an advertisement database and selects the most suitable advertisement. For example, it matches a fitness gym advertisement to a sentence with the topic "exercise."
[1427] Step 5:
[1428] Generate restructured articles
[1429] The server inserts the selected advertisements into the article at the appropriate time and place to generate a reconstructed article. The input is the advertisement list and the original article text, and the output is the reconstructed article text. Specifically, the server inserts advertisements after each sentence.
[1430] Step 6:
[1431] Submitting a restructured article
[1432] The server sends the reconstructed article to the user's device. The input is the reconstructed article text, and the output is a notification of completion of transmission to the user's device. Specifically, the data is sent using the HTTP protocol.
[1433] Step 7:
[1434] View articles and ads
[1435] The terminal displays the received reconstructed article in the user's browser. The input is the reconstructed article text, and the output is the web page displayed to the user. Specifically, it uses an HTML rendering engine to display the article and advertisements in the browser.
[1436] Step 8:
[1437] Video scene analysis
[1438] The server analyzes individual scenes in the video as it plays. The input is the video data, and the output is the analyzed scene information. Specifically, it uses a generative AI model to analyze the context of the scene.
[1439] Sample prompt: "Select the appropriate ad based on the content of the following video scene. Scene description: 'A scene introducing a cooking recipe.'"
[1440] Step 9:
[1441] Ad selection and insertion instruction generation
[1442] The server selects appropriate advertisements based on the analyzed scene information and generates instructions to insert the advertisements at specific times in the video. The input is the analyzed scene information and an advertisement database, and the output is insertion instructions. Specifically, the server selects the optimal advertisement for each scene and determines the appropriate timing.
[1443] The above is the processing flow of this system. By clearly indicating the specific operations and inputs / outputs performed at each step, the processing flow of the entire system becomes clear.
[1444] (Application example 1)
[1445] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1446] Conventional advertising display systems have difficulty displaying ads that are relevant to the context of the content being viewed by the user or to specific scenes at the appropriate time. Furthermore, there is a problem of reduced advertising effectiveness due to the lack of relevance between the content of articles or videos and the ads. Furthermore, this can lead to a poor user experience, making it difficult for advertisers to provide effective ads.
[1447] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1448] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; means for reconstructing the article by inserting the selected advertisements into the article at appropriate times and places; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in a video accessed by the user using a generative AI model while the video is being played; and means for selecting advertisements related to scenes in the video based on the analysis results and displaying them at specific times in the video. This makes it possible to display advertisements that are highly relevant to the content viewed by the user at appropriate times.
[1449] A "server" is a computer system that provides data and resources for joint use over a network.
[1450] "User" means any person, organization, or device that uses a computer system or network service.
[1451] A "web page" is an information page published on the Internet, written in a markup language such as HTML, and viewable using a browser.
[1452] "Article" refers to a single work or document published in a newspaper, magazine, web page, etc.
[1453] A "sentence" is the smallest unit of a sentence or piece of writing, a series of words that has complete meaning.
[1454] "Context" refers to the relationship between adjacent words or sentences, and the overall situation in which it is placed.
[1455] A "generative AI model" is an artificial intelligence model trained using neural networks on large datasets, and is capable of generating, understanding, and analyzing text.
[1456] "Advertising" refers to information created to introduce products and services and encourage consumers to purchase them.
[1457] "Video" refers to moving video content that is made up of a series of images and sounds.
[1458] A "scene" is a single continuous action or event in video content that does not span a specific time period.
[1459] A "list" is a data structure that lists and organizes multiple items.
[1460] As an embodiment of the present invention, a system that performs the following processing steps is constructed.
[1461] First, the server retrieves the article from the web page accessed by the user. The retrieved article is then divided into sentences. The context of each sentence is then analyzed using a generative AI model. For example, the natural language processing model "GPT-4" is used to extract the topic and sentiment information of each sentence.
[1462] After the context analysis is completed, the server selects relevant advertisements from the advertisement database based on the analysis results, lists the selected advertisements, and reconstructs the article based on them. The reconstructed article is then sent to the user's device.
[1463] Next, while the video being viewed by the user is being played, the server analyzes each scene in the video using a generative AI model. It selects the appropriate ad for each scene and generates instructions to display it at a specific moment in the video. These instructions are then sent to the user's device, and the ad is displayed at the appropriate time.
[1464] The specific configuration for realizing this system involves running the following program. First, the server retrieves articles from a webpage using the Python libraries "requests" and "BeautifulSoup," and then splits the articles into sentences using the "split_by_sentence" function. Next, contextual analysis is performed using the generative AI model "GPT-4." The analysis is performed by sending a prompt sentence using the "GPT-4" API and obtaining a response to it. The following prompt sentence is used as an example of this analysis.
[1465] "Analyze the context of this sentence and identify the topic and emotion: Regular exercise every day is good for your physical and mental health."
[1466] "Analyze the context of this sentence and identify the topic and emotion: A balanced diet boosts your immunity."
[1467] Based on the analysis results, the server selects the most suitable advertisement from the advertisement database and reconstructs the article, which is then sent to the user's device and displayed in the browser.
[1468] For video analysis, each scene in the video is identified and the context of each scene is analyzed using a generative AI model. For example, if a user is watching a cooking recipe video, an advertisement for cooking utensils will be displayed immediately after the cooking scene ends. For context analysis for each scene, a prompt in the form of "Analyze the context of this scene and identify related advertisements:" is used.
[1469] In this way, it is possible to build a system that dynamically displays highly relevant advertisements at the appropriate time based on the content the user is viewing, thereby improving the user experience and maximizing advertising effectiveness.
[1470] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1471] Step 1:
[1472] The server retrieves the article of the web page accessed by the user. As input, it receives the URL of the web page accessed by the user. As output, it returns the text data of the entire article of the web page. Specifically, the server uses the "requests" library to retrieve the HTML of the web page and "BeautifulSoup" to extract the text data.
[1473] Step 2:
[1474] The server splits the retrieved articles into sentences. As input, it receives the text data of the entire article. As output, it returns a list of text split into sentences. Specifically, the server uses Python's string manipulation functions to split the article using delimiters such as "."
[1475] Step 3:
[1476] The server analyzes the context of each sentence using a generative AI model. As input, it receives a text list divided into sentences. As output, it returns the context information (topic and sentiment information) of each sentence. Specifically, the server sends a prompt sentence to the API of the generative AI model "GPT-4" and obtains the analysis results.
[1477] Step 4:
[1478] The server selects relevant advertisements based on the results of the context analysis. As input, it receives the context information of each sentence. As output, it returns a list of relevant advertisements. Specifically, the server matches the context information with an advertisement database to select the most suitable advertisement.
[1479] Step 5:
[1480] The server reconstructs the article by inserting the selected advertisements at the appropriate time and place. As input, it receives the segmented sentences and a list of selected advertisements. As output, it returns the reconstructed article with the advertisements inserted. Specifically, the server executes logic to insert advertisements at the appropriate place for each sentence.
[1481] Step 6:
[1482] The server sends the reconstructed article to the user's device. As input, it receives the reconstructed article. As output, it sends the article to the user's device. Specifically, the server uses an HTTP request to display the reconstructed article in the user's browser.
[1483] Step 7:
[1484] The server analyzes the context of each individual scene in a video while the video accessed by the user is being played. As input, it receives video data and timestamps. As output, it returns contextual information for each scene. Specifically, the server uses a generative AI model to analyze each scene in the video.
[1485] Step 8:
[1486] The server selects advertisements relevant to the scenes based on the analysis results and generates instructions to display them at specific times in the video. It receives contextual information for each scene as input. It returns instructions to display the advertisements as output. Specifically, the server compares the analysis results for each scene with an advertisement database, selects appropriate advertisements, and sends instructions to the user's device for the timing of their display.
[1487] Step 9:
[1488] The device displays the reconstructed article in the user's browser and displays advertisements corresponding to each sentence at the appropriate time. As input, it receives the reconstructed article and advertisement display instructions received from the server. As output, it displays the article and advertisements in the user's browser. Specifically, the device controls the browser display using HTML and JavaScript and displays advertisements at the appropriate time.
[1489] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1490] This invention is a system that displays advertisements at optimal times based on not only user attribute information but also the context of the content itself and the user's emotional state in web page articles and video content. By adding an emotion engine, this system can recognize user emotions and further improve the accuracy and effectiveness of advertisement display.
[1491] 1. Get and split articles
[1492] When a user accesses a specific web page, the device (user's browser) sends the request to the server. The server receives the request and retrieves the article from the target web page. The server then divides the retrieved article into sentences. For example, it may divide the sentences into "Daily moderate exercise is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[1493] 2. Context analysis and emotional state recognition
[1494] The server uses a generative AI model to analyze the context of each sentence. Context includes topic and emotional information. The server also uses an emotion engine to detect the user's emotional state. For example, the server can analyze facial expressions and tone of voice using a camera or microphone.
[1495] 3. Ad selection
[1496] Based on the analysis results, the server searches the advertisement database for advertisements related to each sentence and the user's emotional state. For example, if the sentence is related to "exercise" and the user is in a "positive" emotional state, an advertisement for a fitness gym will be selected. The analysis results of the emotion engine play an important role in advertisement selection.
[1497] 4. Restructure and submit your article
[1498] The server then incorporates the selected advertisements into the original article to generate a reconstructed article, which is then sent to the user's device. The reconstructed article also includes advertisements based on the user's emotional state.
[1499] 5. Display of articles and advertisements
[1500] The device displays the reconstructed article in the user's browser. The device displays relevant advertisements immediately after each sentence. For example, a fitness gym advertisement is displayed after the sentence "Daily moderate exercise is good for your physical and mental health," while a health food advertisement is displayed after the sentence "A balanced diet strengthens your immune system." As the user continues reading the article, they can naturally see highly relevant advertisements based on their emotional state and context.
[1501] Video analysis and ad insertion
[1502] When a user starts watching a particular video, the device sends the request to the server, which receives the request and prepares the video for analysis. The server uses a generative AI model to analyze the context of each scene in the video and an emotion engine to recognize the user's emotional state.
[1503] Based on the analysis results, the server searches the advertisement database for advertisements relevant to each scene and the user's emotional state. For example, if a cooking scene is displayed and the user finds it "interesting," an advertisement for cooking utensils will be selected. The server creates timing instructions including the selected advertisement and reconstructs the video to be displayed at a specific moment. The reconstructed video data is sent to the user's device. The device displays relevant advertisements at the specified timing for the video being played, allowing the user to naturally view advertisements based on the content of the video and the user's emotional state.
[1504] This invention is a system that can take into account the user's emotional state, increase the effectiveness of advertising, and improve the user experience by combining advertising display related to the content of web articles and entire videos with an emotion engine.
[1505] The processing flow will be explained below.
[1506] Step 1:
[1507] A user visits a specific web page.
[1508] Step 2:
[1509] The device (user's browser) sends the request to the server.
[1510] Step 3:
[1511] The server receives the request and retrieves the article from the target web page.
[1512] Step 4:
[1513] The server divides the text data of the acquired article into sentences.
[1514] Step 5:
[1515] The server feeds each sentence into a generative AI model to analyze the context, which includes topic and sentiment information.
[1516] Step 6:
[1517] The server uses an emotion engine to detect the user's emotional state by analyzing the user's facial expressions and tone of voice using the device's camera and microphone.
[1518] Step 7:
[1519] Based on the analysis results of the generative AI model and the results of the emotion engine, the server selects advertisements from the advertising database that are relevant to each sentence and the user's emotional state.
[1520] Step 8:
[1521] The server lists the selected ads and determines the best time and place to display them.
[1522] Step 9:
[1523] The server reconstructs the article so that it includes advertisements and transmits it to the user's terminal.
[1524] Step 10:
[1525] The terminal displays the received reconstructed article on the user's browser.
[1526] Step 11:
[1527] The device displays relevant ads immediately after each sentence, allowing users to naturally see relevant ads based on their context and emotional state as they continue reading the article.
[1528] Steps for Video
[1529] Step 1:
[1530] A user begins watching a particular video.
[1531] Step 2:
[1532] The terminal sends the request to the server.
[1533] Step 3:
[1534] The server receives the request and prepares the video for analysis.
[1535] Step 4:
[1536] The server uses a generative AI model to analyze the context of each scene in the video.
[1537] Step 5:
[1538] The server uses an emotion engine to set the user's emotional state by analyzing the user's emotional response to the video they are watching.
[1539] Step 6:
[1540] The server searches the advertising database for advertisements relevant to the scene based on the analysis results obtained from the generative AI model and emotion engine.
[1541] Step 7:
[1542] The server selects advertisements for cooking tools for the "cooking scene" and advertisements for sports equipment for the "exercise scene." The server adjusts the content and timing based on the user's emotional state.
[1543] Step 8:
[1544] The server creates timing instructions containing advertisements and reconfigures them to appear at specific moments in the video.
[1545] Step 9:
[1546] The server transmits the reconstructed video data to the user's terminal.
[1547] Step 10:
[1548] The device will display relevant ads at designated times while the video is playing, allowing users to see ads that are highly relevant to the content of the video in a natural way.
[1549] Example 2
[1550] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1551] Conventional advertising systems typically display ads based on user attribute information, but this method fails to fully consider each user's emotional state or the context of the content, resulting in ineffective ad display. Furthermore, even in video content, the timing and content of ad insertions often fail to match the user's emotions, resulting in a poor user experience. It is necessary to provide a system that can resolve these issues and display more effective and relevant ads to users.
[1552] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1553] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; means for acquiring emotional information using an emotion engine that recognizes the user's emotional state; means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; means for reconstructing the article by inserting the selected advertisements into the article at appropriate times and locations; means for transmitting the reconstructed article to the user's terminal; means for analyzing the context of individual scenes in the video using a generative AI model while the video is being played; and means for selecting advertisements related to the scenes based on the analysis results and displaying them at specific times in the video. This enables more effective and relevant advertisements to be displayed based on the user's emotional state and the context of the content.
[1554] "User" means an individual who accesses the Website or Application and uses or consumes its content.
[1555] A "device" is a device, such as a computer, smartphone, or tablet, that a user uses to connect to the Internet and view web pages and videos.
[1556] A "server" is a computer system that stores, manages, and transmits and receives data over a network, and is a device that provides content in response to user requests.
[1557] A "web page" is a unit of information that is published on the Internet and can be viewed by users through a browser; it is a digital page that consists of text, images, videos, etc.
[1558] An "article" is text-based information content posted on a web page, and may be in the form of news, blogs, commentaries, etc.
[1559] A "sentence" is a complete sentence in an article, the smallest meaningful unit of text.
[1560] A "generative AI model" is a machine learning model that uses artificial intelligence to generate, analyze, translate, and perform other tasks on text, and is based on natural language processing technology.
[1561] "Context" refers to the meaning or background that a particular part of a sentence or image has in relation to other parts or the whole.
[1562] "Analysis" is the process of breaking down data or information to understand its components and patterns.
[1563] An "emotion engine" is a technology that analyzes external information such as a user's facial expressions and tone of voice to determine the user's emotional state.
[1564] "Emotional state" refers to the type and intensity of emotion a user feels in a particular situation, and includes states such as "positive," "negative," and "interesting."
[1565] "Advertising" is promotional content that presents information about products and services and stimulates consumers' desire to purchase them.
[1566] A "list" refers to an ordered arrangement of related items.
[1567] "Timing" refers to the moment in time when a particular event or action occurs.
[1568] A "location" refers to a specific location within a web page or video.
[1569] "Reconstruction" means constructing the original content in a new form by inserting new elements or rearranging them.
[1570] "Video" is digital media content that combines a sequence of images and sounds to tell a visual and audio story.
[1571] A "scene" refers to a specific sequence of images and sounds in a video, and is a unit that depicts a specific situation or activity.
[1572] This invention is a system that displays advertisements at optimal times based on not only user attribute information but also the context of the content itself and the user's emotional state in web page articles and video content. By adding an emotion engine, this system can recognize user emotions and further improve the accuracy and effectiveness of advertisement display.
[1573] First, when a user accesses a specific web page, their device (the user's browser) sends the request to a server. For example, if a user accesses a news site to read an article about health, this request is sent. The server receives the request and retrieves the article from the target web page. Next, the server uses a generative AI model (for example, a model using natural language processing technology) to divide the retrieved article into sentences. Specifically, it divides the article into sentences such as "Daily moderate exercise is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[1574] The server uses a generative AI model to analyze the context of each sentence. This context includes topic and emotional information. The server also uses an emotion engine (e.g., technology that analyzes facial expressions and vocal tones) to recognize the user's emotional state. Specifically, the server analyzes facial expressions and vocal tones captured in real time through the device's camera and microphone to determine whether the user is in a "positive," "negative," "interested," or other emotional state.
[1575] Next, the server searches the advertisement database for advertisements related to each sentence and the user's emotional state based on the analysis results. For example, if the sentence is related to "exercise" and the user is in a "positive" emotional state, advertisements for fitness gyms will be selected. The selected advertisements will be listed.
[1576] The server then incorporates the selected advertisements into the original article to generate a reconstructed article. For example, it might insert an advertisement for a fitness gym after the sentence, "Doing moderate exercise every day is good for your physical and mental health." The reconstructed article data is then sent to the terminal.
[1577] The device then displays the reconstructed article in the user's browser. Relevant advertisements are displayed immediately after each sentence, allowing users to naturally see highly relevant advertisements based on their emotional state and context as they read the article. For example, an advertisement for a health food product is displayed after the sentence, "A balanced diet boosts your immune system."
[1578] When a user starts watching a specific video, the device sends a request to the server. The server receives the request and prepares the video for analysis. The server uses a generative AI model to analyze the context of each scene in the video and an emotion engine to recognize the user's emotional state and suggest advertisements. The specific analysis takes into account the content displayed in each scene of the video and the user's reactions (facial expressions, tone of voice, etc.).
[1579] For example, if a user finds a cooking scene "interesting," an advertisement for cooking utensils will be selected. The server reconstructs the selected advertisement to be displayed at a specific moment in the video, and sends the reconstructed video data to the device. The device then displays the relevant advertisement at a specified timing during video playback. While watching the video, the user can naturally see advertisements based on their emotional state and the context of the video.
[1580] For example, the following prompt might be used: "Explain how to analyze the article content of a web page a user views and use an emotion engine to display relevant ads."
[1581] By implementing this system, it will be possible to display more effective and relevant advertisements that take into consideration the user's emotional state and the context of the content, which is expected to improve the effectiveness of advertisements and the user experience.
[1582] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1583] Step 1:
[1584] When a user accesses a specific web page, the device (the user's browser) sends the request to the server. The input is the URL of the web page being accessed, and the output is the article data for the web page. In concrete terms, when a user accesses the "health" category on a news site, the browser requests "https: / / example.com / health" from the server. The server analyzes the request and retrieves the article data from the specified URL.
[1585] Step 2:
[1586] The server splits the retrieved article data into sentences. The input is the article data of the web page, and the output is text data split into sentences. Specifically, the server uses the "Sentence Splitter API" to split the article text into sentences such as "Doing moderate exercise every day is good for your physical and mental health" and "A balanced diet strengthens your immune system."
[1587] Step 3:
[1588] The server uses a generative AI model (e.g., a model using natural language processing technology) to analyze the context of each sentence. The input is text data divided into sentences, and the output is the context data for each sentence. Specifically, the server sends each sentence to the generative AI model, which extracts the topics "health" and "exercise" and positive sentiment information from the sentence "exercise is good for your health."
[1589] Step 4:
[1590] The server recognizes the user's emotional state using an emotion engine (e.g., technology that analyzes facial expressions and voice tones). The input is real-time data acquired through the device's camera and microphone, and the output is the user's emotional state data. Specifically, the server analyzes the user's facial expressions and voice tone to determine whether the user's emotional state is "positive," "negative," "interested," etc.
[1591] Step 5:
[1592] The server searches for relevant advertisements from the advertisement database based on the analysis results (contextual data and emotional state data). The input is contextual data and emotional state data, and the output is a list of selected advertisements. Specifically, the server selects advertisements for fitness gyms if the sentence is related to "exercise" and the user is in a "positive" emotional state.
[1593] Step 6:
[1594] The server incorporates the selected advertisements into the original article and generates a reconstructed article. The input is text data divided into sentences and a list of advertisements, and the output is the reconstructed article data. Specifically, the server adds an advertisement for a fitness gym after the sentence, "Doing moderate exercise every day is good for your physical and mental health."
[1595] Step 7:
[1596] The server sends the reconstructed article data to the terminal. The input is the reconstructed article data, and the output is the data sent to the user's terminal. In concrete terms, the server sends the reconstructed article data to the terminal and it is displayed in the specified browser.
[1597] Step 8:
[1598] The terminal displays the received reconstructed article in the user's browser. The input is the reconstructed article data, and the output is the information displayed in the user's browser. In concrete terms, the terminal displays the reconstructed article in the browser, and an advertisement for a health food is displayed after the sentence, "A balanced diet boosts your immune system."
[1599] Step 9:
[1600] When a user starts watching a particular video, the device sends a request to the server. The input is the URL of the video being watched, and the output is the video data. In concrete terms, when a user requests "https: / / example.com / cooking" to watch a cooking tutorial video, the device sends this information to the server.
[1601] Step 10:
[1602] The server receives the request and prepares the target video for analysis. The input is the video data, and the output is data for each scene in the video. Specifically, the server divides the video into frames and identifies each scene.
[1603] Step 11:
[1604] The server uses a generative AI model to analyze the context of each scene in the video. The input is the video data for each scene, and the output is the contextual data for each scene. Specifically, the server sends each scene to the generative AI model, identifies the context, and extracts information such as "cooking scene" or "exercise scene."
[1605] Step 12:
[1606] The server uses an emotion engine to recognize the user's emotional state. The input is real-time data acquired through the device's camera and microphone, and the output is the user's emotional state data. Specifically, the server analyzes the user's facial expressions and tone of voice while watching a video to determine whether the user is in an "interesting" emotional state.
[1607] Step 13:
[1608] The server searches for relevant advertisements from the advertisement database based on the analysis results (scene context data and emotional state data). The input is the scene context data and emotional state data, and the output is a list of selected advertisements. Specifically, when a cooking scene is displayed and the user feels that it is "interesting," the server selects an advertisement for cooking utensils.
[1609] Step 14:
[1610] The server inserts the selected advertisements into specific scenes of the video and reconstructs it. The input is the video data for each scene and the advertisement list, and the output is the reconstructed video data. Specifically, the server inserts advertisements at the specified timing and reconstructs the video.
[1611] Step 15:
[1612] The server transmits the reconstructed video data to the terminal. The input is the reconstructed video data, and the output is the data to be transmitted to the user's terminal. In concrete terms, the server transmits the reconstructed video data to the terminal.
[1613] Step 16:
[1614] The device displays relevant advertisements at specified times during the video being played. The input is the reconstructed video data, and the output is a video with advertisements that is displayed to the user. Specifically, the device displays advertisements for cooking utensils during specific scenes in the video.
[1615] (Application example 2)
[1616] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1617] Conventional advertising systems for web pages and videos display ads based on user attribute information and the general context of the content, but this does not fully consider the user's emotional state or the specific context of the moment. As a result, ads are often ineffective or not perceived as natural by users. Furthermore, even when using smart glasses or head-mounted displays, there is a lack of technology that can analyze the user's emotional state in real time and display appropriate ads. This has led to a demand for improved user experience.
[1618] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1619] In this invention, the server includes: means for acquiring articles from web pages accessed by a user; means for dividing the acquired articles into sentences and analyzing the context of each sentence using a generative AI model; sentiment analysis means for recognizing an emotional state from the context of the acquired articles; means for selecting relevant advertisements based on the analysis results and the emotional state and creating a list of advertisements; means for reconstructing the articles by inserting the selected advertisements into the articles at appropriate times and places; means for transmitting the reconstructed articles to the user's terminal; and means for displaying the reconstructed articles received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times. This enables more effective and natural advertisement display based on the user's emotional state and the specific context of the content.
[1620] A "User" is an individual who accesses a particular web page or video content.
[1621] A "web page" is a unit of information that is published on the Internet and consists of text, images, videos, etc.
[1622] An "article" refers to a sentence or text content that is published on a web page.
[1623] A "sentence" is a unit of text that has an independent meaning.
[1624] A "generative AI model" is an artificial intelligence model used to perform natural language processing and contextual analysis.
[1625] "Emotion analysis means" is a technology for recognizing a user's emotional state from facial expressions, tone of voice, etc.
[1626] "Advertising" refers to messages or content that convey information about products or services and encourage desired behavior.
[1627] A "list" refers to an ordered arrangement of related items.
[1628] "Timing" refers to the point in time at which a particular event or action should occur.
[1629] "Reconstruction" refers to inserting advertisements into existing content and rearranging it into a new form.
[1630] A "terminal" is a device used by a user, such as a computer, smart glasses, or head-mounted display.
[1631] "Video" is a media format that expresses movement by playing a series of still images.
[1632] A "scene" refers to a part of a video that depicts a specific situation or situation.
[1633] "Smart glasses" are a wearable device in the form of glasses worn by the user that can display information and analyze emotions.
[1634] A "head-mounted display" is a display device worn on the head that provides visual information.
[1635] "Real time" means that things happen at the same speed as real time.
[1636] To implement this invention, a system including a server, a terminal, and a user is required. The system can display advertisements related to the content of articles and videos on a web page at optimal timing.
[1637] The server retrieves articles from web pages accessed by users. The retrieved articles are divided into sentences, and the context of each sentence is analyzed using a generative AI model. At the same time, the user's emotional state is recognized using a sentiment analysis method. Based on the analysis results and the user's emotional state, relevant ads are selected and a list of ads is created.
[1638] The selected advertisements are inserted into the article at the appropriate time and place, and the article is reconstructed. This reconstructed article is then sent to the user's device. The user's device displays the received reconstructed article, and the advertisements corresponding to each sentence are displayed at the appropriate time.
[1639] Furthermore, while the video is playing, the server analyzes the context of each scene in the video using a generative AI model. The user's emotional state is recognized from the scene in the video, and an advertisement related to the scene is selected based on the analysis results and the emotional state. The selected advertisement is displayed at a specific time in the video.
[1640] When using smart glasses or head-mounted displays, the user's emotional state is analyzed using a built-in camera and microphone, and advertisements are selected and displayed in real time based on the emotional state and content context.
[1641] The hardware used includes smart glasses or a head-mounted display, a built-in camera and microphone, and the software used includes a sentiment analysis engine, a generative AI model (e.g., GPT-4), a web server (e.g., Node.js), and an advertising database (e.g., MongoDB).
[1642] As a concrete example, when a user is reading an article on a web page that reads, "Doing moderate exercise every day is good for your physical and mental health," the generative AI model analyzes the context of this sentence, and if the emotion engine recognizes the user's "positive" emotional state, an advertisement for a "fitness gym" will be selected. The generative AI model is analyzed using the following prompt sentence:
[1643] Example prompt sentence:
[1644] Analyze the context and emotional state of the following sentences and select the appropriate ad.
[1645] Document: Regular exercise every day is good for your physical and mental health.
[1646] Emotion: Positive
[1647] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1648] Step 1:
[1649] The server receives a request for the web page accessed by the user and retrieves the article from the web page.
[1650] Input: User request (Web page URL)
[1651] Output: Text data of the retrieved article
[1652] Specific operation: The server analyzes the HTTP request, retrieves the web page content from the specified URL, extracts the text data of the article, and passes it to the next step.
[1653] Step 2:
[1654] The server divides the retrieved article into sentences.
[1655] Input: Article text data
[1656] Output: A list of text data split into sentences
[1657] How it works: The server uses a text analysis algorithm to split the article into sentences. You can use a Python natural language processing library, for example.
[1658] Step 3:
[1659] The server analyzes the context of each sentence using a generative AI model.
[1660] Input: A list of text data split into sentences
[1661] Output: Analysis result data on the context of each sentence
[1662] How it works: It calls an API to operate the generative AI model and analyzes the context of each sentence, generating structured data that includes the topic and sentiment information of the sentence.
[1663] Step 4:
[1664] The server utilizes emotion analysis means to recognize the user's emotional state.
[1665] Input: User's facial expression data (camera) or voice data (microphone)
[1666] Output: User's emotional state (e.g. positive, negative, interesting, etc.)
[1667] How it works: The server sends data from the camera and microphone to the emotion analysis engine, which analyzes the user's current emotional state. The emotion analysis engine identifies the user's emotional state based on their facial expressions and tone of voice.
[1668] Step 5:
[1669] The server selects relevant advertisements from an advertisement database based on the analysis results and the emotional state.
[1670] Input: Contextual information for each sentence and the user's emotional state
[1671] Output: A list of relevant ads
[1672] Specific operations: Query the advertisement database and select advertisements that match the context of the sentence and the user's emotional state. Retrieve information about the selected advertisements in list form.
[1673] Step 6:
[1674] The server reconstructs the article by inserting the selected advertisements at the appropriate time and place.
[1675] Input: Context information for each sentence, list of ads
[1676] Output: Reconstructed article
[1677] Specific operation: The server inserts advertisements into the original text data and restructures the article so that the advertisements are displayed immediately after the sentence or at natural times along the flow of the content.
[1678] Step 7:
[1679] The server sends the reconstructed article to the user's terminal.
[1680] Input: Reconstructed article
[1681] Output: Send article to user's device
[1682] What it does: Converts the reconstructed article into HTML or another appropriate format and sends it to the user's browser or device.
[1683] Step 8:
[1684] The terminal displays the received reconstructed article and displays advertisements corresponding to each sentence at appropriate times.
[1685] Input: Reconstructed article
[1686] Output: Articles and advertisements displayed in the browser
[1687] Specific operation: The user's device receives the reconstructed article and displays it in a web browser, with relevant advertisements displayed immediately after each sentence of the article.
[1688] Step 9:
[1689] As the video plays, the server analyzes the context of each scene in the video using a generative AI model.
[1690] Input: Video file
[1691] Output: Context analysis results of video scenes
[1692] How it works: The server analyzes the video data frame by frame and understands the context of each scene based on a generative AI model.
[1693] Step 10:
[1694] The server recognizes the emotional state from scenes in the video and selects advertisements related to the scenes based on the analysis results and the emotional state.
[1695] Input: Context analysis results of video scenes, user emotional state
[1696] Output: A list of ads related to the scene
[1697] Specific operation: The server uses an emotion analysis engine to analyze the user's emotional state in real time and selects advertisements from the advertisement database that correspond to the context of the video scene.
[1698] Step 11:
[1699] The server displays the selected advertisement at a specific time in the video.
[1700] Input: A list of ads related to the scene
[1701] Output: Video data with ads inserted
[1702] Specific operation: The server reconstructs the video data and places the selected advertisement at a specific time so that it is naturally incorporated.
[1703] Step 12:
[1704] The device displays relevant advertisements at designated times in relation to the video being played.
[1705] Input: Video data with ads inserted
[1706] Output: Videos and ads watched by users
[1707] Specific operation: The user's device plays the reconstructed video and displays the advertisement at the specified time.
[1708] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1709] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1710] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1711] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1712] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1713] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1714] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1715] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1716] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1717] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1718] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1719] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1720] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1721] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1722] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1723] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1724] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1725] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1726] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1727] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1728] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1729] The following is further disclosed regarding the above embodiment.
[1730] (Claim 1)
[1731] A means of retrieving articles from web pages visited by users;
[1732] A means of dividing the retrieved articles into sentences and analyzing the context of each sentence using a generative AI model;
[1733] A means for selecting relevant advertisements based on the analysis results and creating a list of advertisements;
[1734] A means of inserting and reconstructing selected advertisements into articles at the appropriate time and place;
[1735] A means for transmitting the reconstructed article to a user's terminal;
[1736] A system including:
[1737] (Claim 2)
[1738] A means for analyzing the context of individual scenes within a video using a generative AI model while the video is playing; and
[1739] A method for selecting advertisements relevant to the scene based on the analysis results and displaying them at specific times in the video;
[1740] The system of claim 1 further comprising:
[1741] (Claim 3)
[1742] 2. The system according to claim 1, further comprising means for displaying the reconstructed article received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times.
[1743] "Example 1"
[1744] (Claim 1)
[1745] A means of retrieving articles from web pages visited by users;
[1746] A means for dividing the acquired articles into sentence units;
[1747] A means of analyzing the context of each sentence using a generative AI model;
[1748] A means for selecting relevant advertisements based on the analysis results and creating an advertisement list;
[1749] A means of inserting and reconstructing selected advertisements into articles at the appropriate time and place;
[1750] A means for transmitting the reconstructed article to a user's terminal;
[1751] A system including:
[1752] (Claim 2)
[1753] A means for analyzing the context of individual scenes within a video using a generative AI model while the video is playing; and
[1754] A method for selecting advertisements relevant to the scene based on the analysis results and displaying them at specific times in the video;
[1755] The system of claim ...
Claims
1. A means of retrieving articles from web pages visited by users; A means of dividing the retrieved articles into sentences and analyzing the context of each sentence using a generative AI model; A means for selecting relevant advertisements based on the analysis results and creating a list of advertisements; A means of inserting and reconstructing selected advertisements into articles at the appropriate time and place; A means for transmitting the reconstructed article to a user's terminal; A system including:
2. A means for analyzing the context of individual scenes within a video using a generative AI model while the video is playing; and A method for selecting advertisements relevant to the scene based on the analysis results and displaying them at specific times in the video; The system of claim 1 further comprising:
3. 2. The system according to claim 1, further comprising means for displaying the reconstructed article received by the user's terminal and displaying advertisements corresponding to each sentence at appropriate times.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A