system
By analyzing video content to integrate ads naturally, the system addresses the disruption caused by traditional ad insertion, ensuring a seamless and satisfying user experience while enhancing ad effectiveness.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-06
AI Technical Summary
The increasing number of advertisements in video streaming services disrupts the user experience, leading to discomfort and reduced satisfaction, necessitating a solution that integrates ads seamlessly without interrupting the viewing experience.
A system that analyzes video content to identify suitable ad insertion points, selects appropriate advertising materials, adjusts color tone, lighting, and resolution using generative AI to naturally integrate ads, and uploads a new video file for seamless playback.
Enables users to watch videos without interruptions, providing a comfortable and immersive experience while maximizing advertising effectiveness.
Smart Images

Figure 2026037956000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The number of advertisements in video streaming services is increasing every year, and users are becoming increasingly reluctant to watch them due to the time wasted and the forced nature of the content. This leads to problems such as an inability to enjoy a comfortable viewing experience and a decrease in satisfaction. However, given the current business model of video streaming, advertisement delivery is essential for generating revenue. Therefore, it is necessary to improve the user experience of video streaming services while eliminating the forced nature of advertisements. [Means for solving the problem]
[0005] The present invention provides a means for receiving video content and identifying appropriate points for inserting advertisements through video analysis. It also provides a means for selecting appropriate advertising materials from an advertising database based on the identified points, and analyzing and adjusting the color tone, lighting, and resolution of the video to naturally integrate the advertisements into the video. This allows a new video file containing the generated advertisements to be generated and uploaded to a distribution server. This allows for seamless advertisement insertion, allowing users to watch the video without experiencing interruptions due to the advertisements.
[0006] "Video content" is media containing video and audio that is distributed digitally.
[0007] "Video analytics" is a technology that extracts information from digital images and videos and identifies specific patterns and scenes.
[0008] "Advertising materials" refers to digital data such as video, images, text, and audio that constitute the content of an advertisement.
[0009] "Generative AI" refers to algorithms that use artificial intelligence techniques to automatically generate, edit, or synthesize media content for a specific purpose.
[0010] A "distribution server" is a server device used to distribute digital content to users via the Internet.
[0011] "User terminal" refers to the device used by a user to watch video content, and primarily refers to a computer, smartphone, tablet, etc.
[0012] "Seamless" means a consistent state without any physical, temporal, or visual interruptions. [Brief explanation of the drawings]
[0013] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] Explaining program processing in natural language
[0035] server
[0036] 1. Receiving and analyzing the video
[0037] The server receives the video content and analyzes each frame of the video, using video analysis algorithms to identify specific scenes or objects within the video (e.g., building walls, billboards, sides of cars, etc.) that are deemed suitable points for inserting advertisements.
[0038] 2. Selection of advertising materials
[0039] Based on the analysis results, the server selects the appropriate advertising material from its advertising database. This selection is based on an algorithm that selects the most suitable advertisement for the video content and target audience. For example, if a fashion brand advertisement is determined to be the most suitable for an urban scene, that advertisement will be selected.
[0040] 3. Content Coordination and Synthesis
[0041] The server uses generative AI to composite the selected advertising material into the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the advertising material is naturally integrated into the video. For example, a fashion brand advertisement can be composited onto a building sign, making it appear as if it were part of the original footage.
[0042] 4. Create and upload a new video file
[0043] A new video file containing the generated advertisement is generated and uploaded to a distribution server, allowing users to seamlessly watch the video with the embedded advertisement through a distribution platform.
[0044] Terminal
[0045] 1. Requesting and Receiving Videos
[0046] When a user selects a video they want to watch, the device sends a request to the distribution server, which then receives a new video file with an embedded advertisement.
[0047] 2. Play the video
[0048] The device then plays the received video file. During playback, the device buffers and decodes the video, providing a seamless viewing experience. There is no gap between the ads and the original video content.
[0049] User
[0050] 1. Select a video
[0051] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[0052] 2. Watching videos
[0053] A user watches a video playing on their device and notices ads that appear naturally within the video, but do not interrupt the normal viewing experience. For example, a user may see billboard ads or background ads that are naturally embedded in a movie.
[0054] 3. Continue watching
[0055] Users can enjoy videos continuously without interruptions caused by advertisements. They can immerse themselves in the content without feeling stressed by advertisements.
[0056] Specific examples
[0057] Server Processing
[0058] 1. The server receives the movie footage and performs video analysis, for example, identifying city scenes with billboards suitable for inserting advertisements.
[0059] 2. The server determines that an advertisement for a fashion brand would be ideal for this scene.
[0060] 3. The generative AI synthesizes a fashion brand's advertising video onto a specific billboard, integrating it into the scene.
[0061] 4. Upload the completed composite video file to the distribution server.
[0062] Terminal handling
[0063] 1. The user device accesses the distribution platform, selects the movie, and starts playback.
[0064] 2. City scenes from the film play, with fashion brand advertisements displayed on appropriate billboards, making them feel like part of a regular billboard.
[0065] 3. The user device provides a seamless viewing experience with no sense of incongruity between the advertisements and the movie.
[0066] User Experience
[0067] 1. The user begins watching a movie and notices a billboard advertisement that is displayed naturally within the city scene.
[0068] 2. The movie is not interrupted and users can watch ads while still being immersed in the story.
[0069] 3. This allows users to enjoy movies without being stressed by advertisements.
[0070] The processing flow will be explained below.
[0071] Server Processing
[0072] Step 1:
[0073] The server receives the video content and performs video decoding to convert each frame into an analyzable format.
[0074] Step 2:
[0075] The server runs video analysis algorithms to identify important scenes and objects in the video, such as building walls, billboards, or the sides of cars, and finds suitable spots for inserting ads.
[0076] Step 3:
[0077] Based on the analysis results, the server lists the identified points and plans to insert advertisements appropriate for those points.
[0078] Step 4:
[0079] The server selects appropriate advertising material from the advertising database based on the video content and target audience, and stores the selected advertising material in temporary memory.
[0080] Step 5:
[0081] The server uses generative AI to composite the selected ad material into specific points in the video, which involves analyzing the color, lighting, and resolution of the video and adjusting them to make the ad appear natural.
[0082] Step 6:
[0083] The server applies a compositing process to generate a new video file with ads embedded naturally and unobtrusively into the original content.
[0084] Step 7:
[0085] The server uploads the newly generated video file to the distribution server, making it available for viewing by users.
[0086] Terminal handling
[0087] Step 1:
[0088] The user terminal accesses the distribution platform and selects the video that the user wants to watch.
[0089] Step 2:
[0090] The terminal transmits a request for the selected video and receives a new video file with an embedded advertisement from the distribution server.
[0091] Step 3:
[0092] The device decodes and buffers the received video file for playback.
[0093] Step 4:
[0094] The device plays the video and processes it so that the advertisements in the video and the original content are displayed seamlessly.
[0095] User Action
[0096] Step 1:
[0097] Users select the video they want to watch on the video streaming platform and press the play button.
[0098] Step 2:
[0099] Users watch videos played on their device and are presented with ads that are naturally embedded within the video but do not interrupt the normal viewing experience.
[0100] Step 3:
[0101] Users can enjoy videos continuously without interruption due to advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[0102] Example 1
[0103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0104] The challenge is to seamlessly integrate advertisements into video content without causing a sense of incongruity to users. Conventional methods often result in videos with advertisements inserted that look unnatural, detracting from the user's viewing experience.
[0105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0106] In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for selecting appropriate advertising materials from an advertising database based on the identified points, means for analyzing and adjusting the color tone, lighting, and resolution of the video using a generative AI model to naturally combine the advertising materials into the video, means for generating a new video file including the generated advertisements and uploading it to a distribution server, means for a user to select a video they wish to watch on a video distribution platform, and means for a user terminal to send a request to the distribution server, receive the generated video file, and seamlessly play the video so that there is no sense of incongruity between the advertisements and the original video content. This allows users to enjoy a seamless and natural viewing experience while comfortably watching the advertisements.
[0107] "Video content" is a series of digital data including video and audio, and is a medium for users to view.
[0108] A "scene" refers to a specific frame or group of frames in video content, which represents a specific situation or background.
[0109] "Advertisement" means media content inserted into a video for the purpose of promoting a product or service.
[0110] A "point" refers to a specific position or timing within a scene that is suitable for inserting an advertisement.
[0111] "Means" refers to a method or mechanism for achieving a specific function or operation.
[0112] The "advertising database" is a database in which advertising materials are stored, and includes various advertising contents.
[0113] "Advertising materials" refers to specific media data used as advertising, including images, videos, text, etc.
[0114] "Analysis" is the process of examining data in detail and extracting specific information.
[0115] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to generate new data.
[0116] "Color tone" refers to the adjustment and balance of colors in videos and images.
[0117] "Lighting" refers to the intensity and direction of light in a scene, and is a factor in creating visual atmosphere.
[0118] "Resolution" is an indicator of how clearly the details of videos and images can be displayed.
[0119] A "distribution server" is a server for distributing videos and data to user terminals.
[0120] A "user terminal" is a device used by a user, including a PC, smartphone, tablet, etc.
[0121] A "request" is a request message sent from a user terminal to a server to obtain specific data.
[0122] "Prompt" refers to input text used to prompt a generative AI model to generate a particular result.
[0123] "Seamless" refers to a state without joints and indicates a natural, flowing continuity.
[0124] This invention provides a system for inserting advertisements into video content in a natural way, using a method for analyzing scenes and objects in the video, selecting the most suitable advertisements, and synthesizing them in a natural way. The following describes how to specifically implement this invention.
[0125] The server receives video content from users or content providers via HTTP requests or FTP. The received video content is broken down into frames using video analysis algorithms such as the OpenCV library or TENSORFLOW (registered trademark), and objects and scenes within each frame are analyzed. For example, distinctive objects such as building walls, signs, and the sides of cars are identified.
[0126] The server then selects appropriate advertising materials from an advertising database based on the analysis results. During this process, it uses recommendation algorithms such as the Google® AdSense API to select the most suitable advertisements based on the video content and target audience. For example, it may determine that a fashion brand advertisement is best suited for an urban scene.
[0127] Next, the server uses a generative AI model (such as DALL-E) to composite the selected advertising material into the video. Specifically, the server analyzes the color, lighting, and resolution of the video and adjusts it so that the advertising material is naturally integrated into the frame. At this time, the server uses prompts to the generative AI model to generate the desired results. For example, the server might use the prompt, "Generate a fashion brand advertisement on a billboard in an urban scene and display it naturally."
[0128] The completed video file is generated by the server, encoded, and uploaded to the distribution server via an HTTP POST request, allowing users to seamlessly watch the video with embedded ads through the video distribution platform.
[0129] When a user selects a video they want to watch on a video streaming platform and presses the play button, the device sends an HTTP GET request to the streaming server and receives the generated video file. The device decodes the video file and performs appropriate buffering during playback, providing the user with a seamless viewing experience. The user will notice advertisements embedded naturally within the original video, but the viewing experience is not interrupted.
[0130] As a specific example, the following prompt sentence is used:
[0131] "Generating and displaying fashion brand advertisements naturally on billboards in urban scenes"
[0132] "Naturally incorporate car ads into the background of movie scenes"
[0133] Users can enjoy the video without feeling uncomfortable by watching videos with these advertisements inserted naturally. In this way, the invention improves the user's viewing experience and maximizes the effectiveness of advertising.
[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0135] Step 1:
[0136] Video reception and analysis
[0137] The server receives video files from users or content providers via HTTP requests or FTP. It uses the OpenCV library to break down the received video files into frames and analyze the objects and scenes contained in each frame. For example, it identifies color, lighting, and specific objects (buildings, signs, cars, etc.). The input video file is in digital format, and the analysis results for each frame are obtained as output.
[0138] Step 2:
[0139] Selection of advertising materials
[0140] The server selects appropriate advertising materials from an advertising database based on the analysis results. Recommendation algorithms such as the Google AdSense API are used to select the most suitable advertisements for the video content, scene, and target audience. For example, it may determine that a fashion brand advertisement is most suitable for an urban scene. The analysis results and scene information are used as input, and the selected advertising materials are generated as output.
[0141] Step 3:
[0142] Content adjustment and synthesis with generative AI models
[0143] The server uses a generative AI model to naturally incorporate the selected advertising material into the video. A prompt is created and input to the generative AI model (e.g., DALL-E). The prompt is used as an example: "Generate a fashion brand advertisement on a billboard in an urban scene and display it naturally." The generative AI model generates the required advertising images and uses OpenCV to composite the advertising material into the video frame. The prompt and advertising material are used as input, and a frame with the advertisement composited is obtained as output.
[0144] Step 4:
[0145] Creating and uploading a new video file
[0146] The server generates a new video file from all composited frames, encodes and compresses it, and saves it in the specified format (e.g., MP4 or MKV). It then uploads the resulting video file to the distribution server via an HTTP POST request. It uses each composited frame as input and generates a new video file as output, which is then uploaded to the distribution server.
[0147] Step 5:
[0148] Requesting and Receiving Videos
[0149] A user selects a video they want to watch on a video distribution platform and presses the play button. This causes the user device to send an HTTP GET request to the distribution server. A new video file with an embedded advertisement is sent from the distribution server to the user device as an HTTP response. The user's request information is used as input, and the new video file is provided to the user device as output.
[0150] Step 6:
[0151] Play video
[0152] The user device decodes the received video file and plays it through the appropriate player. The device buffers during playback to ensure a seamless viewing experience for the user. Advertisements are displayed seamlessly between the original video content. The new video file is used as input, and a seamlessly played video is provided as output.
[0153] The above steps realize a system that inserts advertisements naturally into video content, providing users with a natural viewing experience.
[0154] (Application example 1)
[0155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0156] Conventional methods for inserting advertisements into video content often interrupt the viewer's experience, resulting in limited advertising effectiveness. Furthermore, it is difficult to integrate advertisements in a visually natural way, often causing viewers to feel uncomfortable. The present invention aims to provide a method for seamlessly and naturally inserting advertisements into video content without interrupting the viewer's experience.
[0157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0158] In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for selecting appropriate advertising materials from an advertising database, means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine the advertising materials with the video, means for seamlessly receiving and playing the generated video file at the user terminal, and means for analyzing the effectiveness of the advertisements and generating a report of the results. This enables visually natural and seamless advertisement insertion, making it possible to provide video content without interrupting the viewer's experience while increasing the effectiveness of the advertisements.
[0159] "Video Content" means a digital video file containing visual and audio information.
[0160] A "scene" refers to a series of frames in video content, and signifies a sequence of images at a specific location or time axis.
[0161] An "advertising database" is a digital repository for storing advertising materials, and is a system that stores a wide variety of advertising data.
[0162] "Advertising Materials" refers to digital content such as images, videos, and text used as advertising.
[0163] "Generative AI" refers to algorithms or models that use artificial intelligence techniques to generate new data or content.
[0164] A "distribution server" is a centralized computer system that provides digital content to user terminals over a network.
[0165] "User terminal" refers to a device, such as a smartphone, tablet, or PC, that allows a user to access digital content via the Internet.
[0166] "Seamless" refers to a state in which the continuity of operations and actions is uninterrupted, and is carried out smoothly without any sense of discomfort.
[0167] "Analysis" refers to the methods and processes used to analyze data and information and understand its structure and meaning.
[0168] "Synthesis" refers to the process of combining different digital content to create new data or images.
[0169] "Effectiveness analysis" refers to statistical or quantitative research or analysis conducted to evaluate the effectiveness or impact of advertising.
[0170] "Report" means a document that summarizes the results of research and analysis and organizes them visually and written.
[0171] The system embodying the present invention performs a series of processes for inserting advertisements into video content in a natural way. This system is mainly composed of a server and a user terminal.
[0172] Server Processing
[0173] Video reception and analysis
[0174] The server first receives the video content sent from the user device, then analyzes each frame of the video using a video analysis algorithm (e.g., OpenCV) to identify scenes and objects suitable for inserting advertisements, thereby determining the advertisement insertion points.
[0175] Selection of advertising materials
[0176] The server then selects appropriate advertising material from its advertising database based on the analysis results, using algorithms to select ads that fit the video's content and target audience. For example, fashion brand ads are often selected for city scenes.
[0177] Content Alignment and Composition
[0178] The server uses a generative AI model (e.g., moviepy) to seamlessly blend the selected ad material into the video. It analyzes the color, lighting, and resolution of the video and adjusts it so that the ad material appears as a natural part of the video. This step provides a seamless visual experience for the viewer.
[0179] Creating and uploading a new video file
[0180] A new video file containing the advertisement is generated and uploaded to the distribution server, which then provides the new video file in response to a request from the user terminal.
[0181] User terminal processing
[0182] Requesting and Receiving Videos
[0183] When the user selects a video that he or she wants to watch, the user terminal sends a request to the distribution server and receives a new video file that includes advertisements.
[0184] Play video
[0185] The device then plays the received video file, buffering and decoding the video as it plays, providing a seamless viewing experience with no sense of incongruity between the ad and the original video content.
[0186] User Experience
[0187] Users select a video they want to watch on a video streaming platform and start playing it. While the video is playing, they may notice advertisements inserted naturally, but they do not interrupt the viewing experience. For example, they may see a billboard advertisement displayed naturally within a movie scene. This allows users to enjoy the video without feeling stressed by advertisements.
[0188] Specific examples
[0189] An example of a prompt sentence is as follows:
[0190] Video Analysis Prompt
[0191] Identify suitable scenes in the input video for ad insertion and composite the latest ads appropriately into those locations.
[0192] Advertisements used:
[0193] The latest advertising video.
[0194] Expected Results:
[0195] A new video file is generated in which the advertisements are seamlessly integrated into the original footage.
[0196] As described above, the system of the present invention provides a series of processes for inserting advertisements into video content in a natural way, thereby providing users with a seamless viewing experience.
[0197] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0198] Step 1:
[0199] Video reception and analysis
[0200] The server receives video content sent from the user's device. It analyzes the received video frame by frame using a video analysis algorithm (e.g., OpenCV) to identify scenes and objects suitable for inserting advertisements. The input to this step is the received video content, and the output is information identifying the ad insertion points. Specifically, it determines where advertisements can be inserted naturally based on the objects and background information recognized through frame-by-frame analysis.
[0201] Step 2:
[0202] Selection of advertising materials
[0203] The server selects appropriate advertising material from the advertising database based on the ad insertion point identified in step 1. This selection uses an algorithm to select advertisements that match the video content information and target audience information. The input to this step is the identification information of the ad insertion point, and the output is the selected advertising material. Specifically, the server searches the advertising database for the most suitable advertisement and extracts the corresponding advertising material.
[0204] Step 3:
[0205] Content Alignment and Composition
[0206] The server uses a generative AI model (e.g., moviepy) to composite the selected ad material into the video in a natural way. At this time, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the ad material blends in with the original footage. The input for this step is the video content and the selected ad material, and the output is a new video file with the ad composited into it. Specifically, the color tone and transparency of the ad material are adjusted based on the color tone and light direction of each frame of the video, resulting in a seamless composite.
[0207] Step 4:
[0208] Creating and uploading a new video file
[0209] The server generates a new video file containing the composited advertisement and uploads the video file to the distribution server. The input of this step is the new video file with the composited advertisement, and the output is the video file uploaded to the distribution server. Specifically, the generated video file is encoded into a specified format and transferred to the distribution server.
[0210] Step 5:
[0211] Requesting and Receiving Videos
[0212] When the user selects a video they want to watch, the user terminal sends a request to the distribution server and receives a new video file including advertisements. The input to this step is the user request, and the output is the video file received from the distribution server. Specifically, an HTTP request is sent based on the user's instructions, and the new video file is received as a response.
[0213] Step 6:
[0214] Play video
[0215] The user device plays the received video file. During video playback, buffering and decoding are performed, maintaining a seamless connection between the advertisement and the original video content. The input for this step is the received video file, and the output is a seamlessly played video. Specifically, the video player function is used to buffer and decode the video using a codec, ensuring smooth playback.
[0216] Step 7:
[0217] Providing a viewing experience
[0218] A user selects a video they want to watch on a video streaming platform and starts playback. While the video is playing, they may notice advertisements inserted naturally, but the viewing experience is not interrupted. Specifically, they can enjoy the video continuously while viewing billboard advertisements that are naturally displayed within movie scenes. The input for this step is a video file received from the streaming server, and the output is a video that plays continuously without interruption.
[0219] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0220] Explaining program processing in natural language
[0221] server
[0222] 1. Receiving and analyzing the video
[0223] The server receives the video content, performs video decoding to convert each frame into an analyzable format, and runs video analysis algorithms to identify important scenes and objects in the video (e.g., building walls, billboards, sides of cars, etc.) to find suitable points for inserting advertisements.
[0224] 2. Acquiring Emotion Data
[0225] The server uses an emotion engine to obtain the user's emotions in real time while watching a video, which includes technology to determine the user's emotional state by analyzing their facial expressions, tone of voice, and eye movements.
[0226] 3. Selection of advertising materials
[0227] The server selects appropriate advertising materials from the advertising database based on the content of the video, the target audience, and the acquired user emotion data. For example, if the user is enjoying watching the video, entertainment-related advertisements may be deemed appropriate.
[0228] 4. Content Coordination and Synthesis
[0229] The server uses generative AI to composite the selected advertising material into specific points in the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it to make the advertisement appear natural. For example, a fashion brand advertisement can be composited onto a building sign, making it appear as if it were part of the original footage.
[0230] 5. Create and upload a new video file
[0231] A new video file containing the generated advertisement is generated and uploaded to a distribution server, allowing users to seamlessly watch the video with the embedded advertisement through a distribution platform.
[0232] Terminal
[0233] 1. Requesting and Receiving Videos
[0234] When the user selects a video they want to watch, the device sends a request to the distribution server, which then receives a new video file with an embedded advertisement.
[0235] 2. Play the video
[0236] The device decodes and buffers the received video file for playback.
[0237] 3. Seamless Ad Display
[0238] The device plays the video and performs playback processing so that the advertisements in the video and the original content are displayed seamlessly, providing the user with a seamless viewing experience.
[0239] User
[0240] 1. Select a video
[0241] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[0242] 2. Watching videos
[0243] A user watches a video playing on their device and is presented with ads that are naturally embedded within the video, without interrupting the normal viewing experience. For example, a user might watch a movie with a billboard ad or background ad that is naturally embedded within the video.
[0244] 3. Emotional Feedback
[0245] Real-time user emotional data is collected and used to select and adjust advertising materials, which may even dynamically change based on the user's emotional state while they are watching the video.
[0246] 4. Continue watching
[0247] Users can enjoy videos continuously without feeling interrupted by advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[0248] Specific examples
[0249] Server Processing
[0250] 1. The server receives the movie footage and performs video analysis, for example, identifying city scenes with billboards suitable for inserting advertisements.
[0251] 2. The server uses an emotion engine to obtain the user's emotion in real time while watching. For example, if the user is enjoying the video, the emotion is judged to be positive.
[0252] 3. Based on the positive emotional state, the server determines that an entertainment advertisement (e.g., a trailer for a new movie) would be a good fit for this scene.
[0253] 4. The generative AI composites a movie trailer onto a specific billboard, integrating it into the scene.
[0254] 5. Upload the completed composite video file to the distribution server.
[0255] Terminal handling
[0256] 1. The user device accesses the distribution platform, selects the movie, and starts playback.
[0257] 2. City scenes from the film play, with entertainment ads on appropriate billboards, making them feel like part of the regular billboard scene.
[0258] 3. The user device provides a seamless viewing experience with no sense of incongruity between the advertisements and the movie.
[0259] User Experience
[0260] 1. The user begins watching a movie and notices a billboard advertisement that is displayed naturally within the city scene.
[0261] 2. The movie is not interrupted, allowing users to remain immersed in the story while watching ads, and the ads are selected based on the user's emotional state, further enhancing the viewing experience.
[0262] 3. This allows users to enjoy movies without being stressed by advertisements.
[0263] The processing flow will be explained below.
[0264] Server Processing
[0265] Step 1:
[0266] The server receives the video content and performs video decoding to convert each frame into an analyzable format.
[0267] Step 2:
[0268] The server runs video analysis algorithms to identify important scenes and objects in the video, such as building walls, billboards, or the sides of cars, and find suitable spots for inserting ads.
[0269] Step 3:
[0270] The server uses an emotion engine to obtain real-time emotional data from users watching videos. This emotional data is determined by analyzing the user's facial expressions, tone of voice, eye movements, etc.
[0271] Step 4:
[0272] The server selects appropriate advertising materials from the advertising database based on the content of the video, the target audience, and the acquired user emotion data. For example, if the user emotion indicates enjoyment, entertainment-related advertisements will be selected.
[0273] Step 5:
[0274] The server uses generative AI to composite the selected ad material into specific points in the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the ad looks natural with the original footage. For example, a movie trailer ad can be composited onto a building billboard.
[0275] Step 6:
[0276] The server applies a compositing process to generate a new video file with ads embedded naturally and unobtrusively into the original content.
[0277] Step 7:
[0278] The server uploads the generated new video file to the distribution server, making it available for viewing by users.
[0279] Terminal handling
[0280] Step 1:
[0281] The user terminal accesses the distribution platform and selects the video that the user wants to watch.
[0282] Step 2:
[0283] The terminal transmits a request for the selected video and receives a new video file with an embedded advertisement from the distribution server.
[0284] Step 3:
[0285] The device decodes and buffers the received video file for playback.
[0286] Step 4:
[0287] The device plays the video and processes it so that the advertisements in the video and the original content are displayed seamlessly.
[0288] User Action
[0289] Step 1:
[0290] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[0291] Step 2:
[0292] A user watches a video playing on their device, and an ad appears, naturally embedded within the video, without interrupting the normal viewing experience. For example, a user might watch a billboard ad that is naturally embedded within a movie.
[0293] Step 3:
[0294] Real-time user emotional data is collected and used to select and adjust advertising materials, which may even dynamically change based on the user's emotional state while they are watching the video.
[0295] Step 4:
[0296] Users can enjoy videos continuously without feeling interrupted by advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[0297] Example 2
[0298] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0299] When inserting advertisements into video content, there is a need for a method to display advertisements naturally without disrupting the user's viewing experience. There is also a need for a method to appropriately select advertisements based on the viewer's emotional state and improve the viewing experience. However, conventional systems have difficulty in utilizing user emotional data to select advertisements and integrate them into videos in a natural way.
[0300] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for acquiring user emotion data and selecting advertising material based on the data, means for selecting appropriate advertising material from an advertisement database based on the identified points and the user emotion data, means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine the advertising material into the video, and means for generating a new video file including the combined advertisement and uploading it to a distribution device. This makes it possible to naturally combine optimal advertisements based on the user's emotions into the video, improving the user's viewing experience.
[0301] "Video Content" means a media file that combines video and audio and is provided for the enjoyment of viewers.
[0302] "Scene analysis" is the process of decoding each frame in a video and identifying important objects and points of interest.
[0303] "Good ad insertion points" are specific scenes or locations in a video where ads can be displayed naturally and effectively.
[0304] "User emotion data" is information indicating the emotional state of the viewer obtained from facial expressions, gaze, voice, and the like.
[0305] "Advertising Materials" means advertising footage, images, or other media content for display.
[0306] An "advertising database" is a searchable database that stores various advertising materials.
[0307] "Analysis and adjustment of color, lighting, and resolution" refers to the process of analyzing the quality and visual elements of a video and making adjustments to make the ads appear natural.
[0308] "Generative AI" is a type of artificial intelligence that is a technology that generates advertising materials and synthesizes them into videos.
[0309] "Distribution device" refers to the server and network infrastructure used to deliver the generated video files to users.
[0310] "Seamless playback" refers to a state in which the video and advertisements are played back consecutively, allowing the user to view them as a single entity without feeling any discomfort.
[0311] The "new video file" is a file in which the advertisement has been synthesized and the original video and advertisement are integrated.
[0312] MODE FOR CARRYING OUT THE INVENTION
[0313] The system of this invention receives and analyzes video content, and automatically inserts advertisements based on user emotion data to improve the viewing experience. The system is mainly composed of three elements: a server, a terminal, and a user.
[0314] server
[0315] The server receives the video content and securely stores it using Amazon S3. The stored video data is decoded into individual frames using FFmpeg. OpenCV is then used to analyze the video and identify important scenes and objects. For example, it can detect billboards suitable for inserting advertisements in urban scenes. The server then collects user emotion data using Microsoft® Azure®'s Emotion API. This technology analyzes the user's facial expressions, tone of voice, eye movements, and other factors to determine their emotional state in real time.
[0316] The server then uses Google Cloud's BigQuery to select appropriate ad materials from the ad database. The selected ad materials are then composited into specific points in the video using OpenAI's (registered trademark) generative AI. Blender is used to analyze the color tone, lighting, and resolution of the video and adjust them to make the ads appear natural. Finally, a new video file containing the composited ads is generated and uploaded to the distribution device via Amazon CloudFront. Users can then watch the video with the embedded ads.
[0317] Terminal
[0318] When a user selects a video they want to watch, the device sends a request to the distribution server. In response to the request, a new video file with embedded ads is received from Amazon CloudFront. The received video file is decoded and buffered using VLCKit, and then played back seamlessly. This allows the user to enjoy a natural viewing experience without any sense of incongruity between the video and the ads.
[0319] User
[0320] Users select a video they want to watch on the streaming platform and press the play button. This causes the device to send a request to the server and receive a video file containing the advertisement. As users watch the video on their device, they notice the advertisements, which are naturally embedded in billboards and backgrounds, without interrupting the normal viewing experience. In addition, users' emotional data is collected in real time and used to select and adjust advertising materials. Advertisements may also be dynamically changed based on the user's emotional state, further improving the viewing experience.
[0321] Specific examples
[0322] Server Processing
[0323] 1. Receive movie data and store it in Amazon S3.
[0324] 2. Decode the movie using FFmpeg and identify billboards in urban scenes with OpenCV.
[0325] 3. Use Microsoft Azure's Emotion API to analyze the user's facial expressions in real time and obtain positive emotions.
[0326] 4. Select entertainment advertising materials based on positive emotional states using Google Cloud BigQuery.
[0327] 5. Synthesize ads naturally using OpenAI's generative AI and Blender.
[0328] 6. Upload the composite video file to Amazon CloudFront and display it.
[0329] Terminal handling
[0330] 1. The user selects a movie and sends a request to the distribution server.
[0331] 2. Receive new video files from Amazon CloudFront.
[0332] 3. Use VLCKit to decode and buffer the video.
[0333] 4. Seamlessly play videos with natural ad insertions.
[0334] User Experience
[0335] 1. The user selects a movie and presses the play button.
[0336] 2. While watching the movie, you notice billboard advertisements appearing in city scenes.
[0337] 3. Watch movies without being interrupted by ads.
[0338] 4. While watching, ads may dynamically change based on emotions.
[0339] Prompt Sentence Examples
[0340] 1. "Can you give me an example of a diorama scene that displays entertainment advertisements for a movie?"
[0341] 2. "Please explain techniques for inserting ads naturally into videos, especially those that leverage sentiment data."
[0342] keyword
[0343] Generative AI model, prompt sentence
[0344] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0345] Step 1: Receive and save the video
[0346] The server receives video content from users. For example, a user uploads a movie. This received video data is stored in Amazon S3. This ensures that the data is securely stored and can be used for further processing. The input video data is stored in Amazon S3 as a stored output.
[0347] Step 2: Video decoding and analysis
[0348] The server decodes the video data stored in Amazon S3 into individual frames using FFmpeg. Each decoded frame undergoes video analysis using OpenCV. Specifically, it identifies objects and scenes suitable for inserting advertisements, such as billboards in urban scenes. The input of this step is the video data retrieved from Amazon S3, and the output is the analyzed frame data.
[0349] Step 3: Obtaining user emotion data
[0350] The server uses Microsoft Azure's Emotion API to obtain the user's emotional data in real time, which includes technology that analyzes the user's facial expressions, tone of voice, eye movements, etc. to determine their emotional state. The input for this step is the user's video feed and audio data, and the output is the analyzed emotional data.
[0351] Step 4: Selecting advertising materials
[0352] The server uses Google Cloud's BigQuery to select appropriate advertising materials based on the video content, target audience, and acquired user emotion data. For example, entertainment-related ads are selected for users who show positive emotions. The input of this step is the analyzed frame data and emotion data, and the output is the selected advertising materials.
[0353] Step 5: Adjust and combine content
[0354] The server uses OpenAI's generative AI to composite the selected advertising material into specific points in the video. Blender is used to analyze the color, lighting, and resolution of the video and adjust the video to make the advertisement appear natural. For example, a fashion brand advertisement can be composited naturally onto a billboard in an urban scene. The input to this step is the selected advertising material and the analyzed frame data, and the output is a new frame data with the advertisement composited into it.
[0355] Step 6: Generate and upload a new video file
[0356] The server generates a new video file containing the advertisements and uploads it to the distribution device via Amazon CloudFront. The uploaded video file is then accessible on the user's device. The input of this step is the new frame data with the advertisements mixed in, and the output is the final video file.
[0357] Step 7: Request and receive video
[0358] When the user selects a video to watch, the device sends a request to the distribution server. It receives a new video file delivered by Amazon CloudFront. The input of this step is the video viewing request, and the output is the received video file.
[0359] Step 8: Play the video
[0360] The device uses VLCKit to decode and buffer the received video file, allowing the user to seamlessly watch the video. The input of this step is the received video file, and the output is a decoded, buffered, and playable video.
[0361] Step 9: Obtaining and using user emotional feedback
[0362] While the user is watching the video, the server continues to capture the user's emotional data and use it to select and adjust advertising materials, which may be dynamically changed according to the user's real-time emotional state. The input of this step is real-time user emotional data, and the output is adapted advertising materials.
[0363] (Application example 2)
[0364] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0365] Conventional video ad insertion systems use simple ad synthesis methods, which often detract from the user's viewing experience. In particular, ads are displayed uniformly regardless of the video scene or the user's emotional state, which makes the ads visually unnatural and makes the user feel uncomfortable. Furthermore, there is a need for more personalized ad delivery by utilizing user emotional data to select ads.
[0366] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving video content and analyzing scenes in the video to identify points suitable for inserting advertisements; means for selecting appropriate advertising material from an advertising database based on the identified points; means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine advertising material into the video; means for analyzing user emotional data in real time and selecting an optimal advertisement based on the analysis results; means for combining the selected advertising material into specific points in the video and using a generative AI model to adjust the combination so that it is seamless; and means for generating a new video file including the combined advertisement and uploading it to a distribution server. This enables personalized advertising insertion based on the user's emotional state.
[0367] "Video content" refers to video data that provides information to users visually and audibly.
[0368] A "scene" is a set of consecutive video frames in a specific time range within a video.
[0369] "Advertisement" means visual or audio content intended to promote a product or service.
[0370] A "point" refers to a suitable location in a particular scene or frame of a video for inserting an advertisement.
[0371] "Means" refer to the methods or techniques used to achieve a particular goal.
[0372] A "database" is a collection of information that organizes large amounts of information and makes it possible to search it efficiently.
[0373] "Hue" is an attribute related to the color tone and color combination of an image.
[0374] "Lighting" refers to the intensity and directionality of the light used in the video.
[0375] "Resolution" is a measure of the ability to express detail in an image, and is generally expressed in terms of the number of pixels.
[0376] "Compositing" refers to the process of combining different video materials into one video.
[0377] "Emotion data" is information about the user's emotional state obtained from facial expressions, tone of voice, and the like.
[0378] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to generate new data or content.
[0379] "Uploading" is the act of transferring data from a local environment to a remote server.
[0380] "Terminal" refers to a computing device that is directly operated by a user, including smartphones and personal computers.
[0381] "Seamless" means that the operation is natural, without any gaps or interruptions.
[0382] "Real-time" refers to data acquisition and processing occurring immediately.
[0383] The system for implementing the invention is composed of three elements: a server, a user terminal, and a user. The specific operation process of this system is as follows:
[0384] Server Processing
[0385] The server receives the video content and analyzes it frame by frame. Using video analysis algorithms, it analyzes each scene in the video and identifies important elements (e.g., billboards or building walls) to find the optimal points for inserting ads.
[0386] The server then uses an emotion engine to analyze the user's emotional data in real time. It analyzes the user's facial image and tone of voice to understand their emotional state. Based on this information, the server selects the most suitable advertising material from its advertising database.
[0387] The system uses a generative AI model to seamlessly blend the selected ad material into the video, adjusting color, lighting, and resolution to seamlessly insert the ad material at specific points in the video. Finally, the server generates the composite video file and uploads it to the distribution server.
[0388] The specific hardware used is high-performance server equipment, and the software includes video analysis algorithms, emotion engines, advertising databases, and generative AI models.
[0389] User terminal processing
[0390] When a user selects a video they want to watch, the newly synthesized video file is received from the distribution server. The user's device decodes the video file and plays it back seamlessly, providing a seamless viewing experience for the user.
[0391] The specific hardware used is the user's device, such as a smartphone or PC, and the software includes the video playback application and decoding algorithm.
[0392] User Experience
[0393] Users watch videos through a video distribution platform, and ads are displayed naturally embedded within the video, so the viewing experience is not interrupted. In addition, user emotional data is acquired in real time, making it possible to dynamically display appropriate ads.
[0394] Examples of prompt statements
[0395] 1. "Analyze the user's emotional state using facial images and audio data."
[0396] 2. "Please identify scenes in your video file where billboards or other advertising can be inserted."
[0397] 3. "Choose the best ad based on the user's current emotional state."
[0398] 4. "Please integrate the selected ads into the video scenes in a natural way."
[0399] This format allows users to enjoy a highly satisfying video viewing experience without feeling stressed by advertisements.
[0400] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0401] Step 1:
[0402] The server receives video content and uses video analysis algorithms to convert each frame into an analyzable format. The input is a video file, and the output is frame-by-frame video data. Specifically, the server breaks down the video into frames and identifies important scenes and objects in each frame.
[0403] Step 2:
[0404] The server acquires the user's emotional data in real time. The input is the user's facial image and voice data, and the output is the user's emotional state (e.g., joy, excitement, sadness). Specifically, the server uses an emotion engine to analyze the user's facial expressions and tone of voice to determine the user's emotional state.
[0405] Step 3:
[0406] The server selects appropriate advertising materials from an advertising database based on video analysis data and user emotional data. The input is the analyzed scene data and the user's emotional state, and the output is the selected advertising materials. Specifically, the server uses an advertising identification algorithm to select the advertisement that best suits the user's emotional state.
[0407] Step 4:
[0408] The server utilizes a generative AI model to seamlessly composite selected ad material into specific points in the video. The input is the selected ad material and analyzed scene data, and the output is frame data with the ad composited. Specifically, the server adjusts the color tone, lighting, and resolution of the video to composite the ad material naturally.
[0409] Step 5:
[0410] The server generates a new video file containing the composited advertisements and uploads it to the distribution server. The input is the frame data with the composited advertisements, and the output is a new video file. Specifically, the server reconstructs the frame data into a single video file and transfers it to the distribution server.
[0411] Step 6:
[0412] The user terminal receives a new video file from the distribution server. The input is the video file sent from the distribution server, and the output is the saved video file. Specifically, the user terminal downloads the video file.
[0413] Step 7:
[0414] The user terminal decodes and buffers the received video file and plays it seamlessly. The input is the stored video file, and the output is the played video. Specifically, the user terminal uses a video decoder to buffer the video frame by frame and play it smoothly.
[0415] Step 8:
[0416] The user watches a video played on a device. The input is the video being played, and the output is the user's visual and auditory experience. The specific experience is that the user can enjoy the video containing naturally-mixed advertisements without any sense of incongruity.
[0417] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0418] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0419] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0420] [Second embodiment]
[0421] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0422] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0423] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0424] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0425] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0426] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0427] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0428] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0429] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0430] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0431] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0432] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0433] Explaining program processing in natural language
[0434] server
[0435] 1. Receiving and analyzing the video
[0436] The server receives the video content and analyzes each frame of the video, using video analysis algorithms to identify specific scenes or objects within the video (e.g., building walls, billboards, sides of cars, etc.) that are deemed suitable points for inserting advertisements.
[0437] 2. Selection of advertising materials
[0438] Based on the analysis results, the server selects the appropriate advertising material from its advertising database. This selection is based on an algorithm that selects the most suitable advertisement for the video content and target audience. For example, if a fashion brand advertisement is determined to be the most suitable for an urban scene, that advertisement will be selected.
[0439] 3. Content Coordination and Synthesis
[0440] The server uses generative AI to composite the selected advertising material into the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the advertising material is naturally integrated into the video. For example, a fashion brand advertisement can be composited onto a building sign, making it appear as if it were part of the original footage.
[0441] 4. Create and upload a new video file
[0442] A new video file containing the generated advertisement is generated and uploaded to a distribution server, allowing users to seamlessly watch the video with the embedded advertisement through a distribution platform.
[0443] Terminal
[0444] 1. Requesting and Receiving Videos
[0445] When a user selects a video they want to watch, the device sends a request to the distribution server, which then receives a new video file with an embedded advertisement.
[0446] 2. Play the video
[0447] The device then plays the received video file. During playback, the device buffers and decodes the video, providing a seamless viewing experience. There is no gap between the ads and the original video content.
[0448] User
[0449] 1. Select a video
[0450] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[0451] 2. Watching videos
[0452] A user watches a video playing on their device and notices ads that appear naturally within the video, but do not interrupt the normal viewing experience. For example, a user may see billboard ads or background ads that are naturally embedded in a movie.
[0453] 3. Continue watching
[0454] Users can enjoy videos continuously without interruptions caused by advertisements. They can immerse themselves in the content without feeling stressed by advertisements.
[0455] Specific examples
[0456] Server Processing
[0457] 1. The server receives the movie footage and performs video analysis, for example, identifying city scenes with billboards suitable for inserting advertisements.
[0458] 2. The server determines that an advertisement for a fashion brand would be ideal for this scene.
[0459] 3. The generative AI synthesizes a fashion brand's advertising video onto a specific billboard, integrating it into the scene.
[0460] 4. Upload the completed composite video file to the distribution server.
[0461] Terminal handling
[0462] 1. The user device accesses the distribution platform, selects the movie, and starts playback.
[0463] 2. City scenes from the film play, with fashion brand advertisements displayed on appropriate billboards, making them feel like part of a regular billboard.
[0464] 3. The user device provides a seamless viewing experience with no sense of incongruity between the advertisements and the movie.
[0465] User Experience
[0466] 1. The user begins watching a movie and notices a billboard advertisement that is displayed naturally within the city scene.
[0467] 2. The movie is not interrupted and users can watch ads while still being immersed in the story.
[0468] 3. This allows users to enjoy movies without being stressed by advertisements.
[0469] The processing flow will be explained below.
[0470] Server Processing
[0471] Step 1:
[0472] The server receives the video content and performs video decoding to convert each frame into an analyzable format.
[0473] Step 2:
[0474] The server runs video analysis algorithms to identify important scenes and objects in the video, such as building walls, billboards, or the sides of cars, and finds suitable spots for inserting ads.
[0475] Step 3:
[0476] Based on the analysis results, the server lists the identified points and plans to insert advertisements appropriate for those points.
[0477] Step 4:
[0478] The server selects appropriate advertising material from the advertising database based on the video content and target audience, and stores the selected advertising material in temporary memory.
[0479] Step 5:
[0480] The server uses generative AI to composite the selected ad material into specific points in the video, which involves analyzing the color, lighting, and resolution of the video and adjusting them to make the ad appear natural.
[0481] Step 6:
[0482] The server applies a compositing process to generate a new video file with ads embedded naturally and unobtrusively into the original content.
[0483] Step 7:
[0484] The server uploads the newly generated video file to the distribution server, making it available for viewing by users.
[0485] Terminal handling
[0486] Step 1:
[0487] The user terminal accesses the distribution platform and selects the video that the user wants to watch.
[0488] Step 2:
[0489] The terminal transmits a request for the selected video and receives a new video file with an embedded advertisement from the distribution server.
[0490] Step 3:
[0491] The device decodes and buffers the received video file for playback.
[0492] Step 4:
[0493] The device plays the video and processes it so that the advertisements in the video and the original content are displayed seamlessly.
[0494] User Action
[0495] Step 1:
[0496] Users select the video they want to watch on the video streaming platform and press the play button.
[0497] Step 2:
[0498] Users watch videos played on their device and are presented with ads that are naturally embedded within the video but do not interrupt the normal viewing experience.
[0499] Step 3:
[0500] Users can enjoy videos continuously without interruption due to advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[0501] Example 1
[0502] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0503] The challenge is to seamlessly integrate advertisements into video content without causing a sense of incongruity to users. Conventional methods often result in videos with advertisements inserted that look unnatural, detracting from the user's viewing experience.
[0504] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0505] In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for selecting appropriate advertising materials from an advertising database based on the identified points, means for analyzing and adjusting the color tone, lighting, and resolution of the video using a generative AI model to naturally combine the advertising materials into the video, means for generating a new video file including the generated advertisements and uploading it to a distribution server, means for a user to select a video they wish to watch on a video distribution platform, and means for a user terminal to send a request to the distribution server, receive the generated video file, and seamlessly play the video so that there is no sense of incongruity between the advertisements and the original video content. This allows users to enjoy a seamless and natural viewing experience while comfortably watching the advertisements.
[0506] "Video content" is a series of digital data including video and audio, and is a medium for users to view.
[0507] A "scene" refers to a specific frame or group of frames in video content, which represents a specific situation or background.
[0508] "Advertisement" means media content inserted into a video for the purpose of promoting a product or service.
[0509] A "point" refers to a specific position or timing within a scene that is suitable for inserting an advertisement.
[0510] "Means" refers to a method or mechanism for achieving a specific function or operation.
[0511] The "advertising database" is a database in which advertising materials are stored, and includes various advertising contents.
[0512] "Advertising materials" refers to specific media data used as advertising, including images, videos, text, etc.
[0513] "Analysis" is the process of examining data in detail and extracting specific information.
[0514] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to generate new data.
[0515] "Color tone" refers to the adjustment and balance of colors in videos and images.
[0516] "Lighting" refers to the intensity and direction of light in a scene, and is a factor in creating visual atmosphere.
[0517] "Resolution" is an indicator of how clearly the details of videos and images can be displayed.
[0518] A "distribution server" is a server for distributing videos and data to user terminals.
[0519] A "user terminal" is a device used by a user, including a PC, smartphone, tablet, etc.
[0520] A "request" is a request message sent from a user terminal to a server to obtain specific data.
[0521] "Prompt" refers to input text used to prompt a generative AI model to generate a particular result.
[0522] "Seamless" refers to a state without joints and indicates a natural, flowing continuity.
[0523] This invention provides a system for inserting advertisements into video content in a natural way, using a method for analyzing scenes and objects in the video, selecting the most suitable advertisements, and synthesizing them in a natural way. The following describes how to specifically implement this invention.
[0524] The server receives video content from users or content providers via HTTP requests or FTP. The received video content is broken down into frames using video analysis algorithms such as the OpenCV library or TensorFlow, and objects and scenes within each frame are analyzed. For example, distinctive objects such as building walls, signs, and the sides of cars are identified.
[0525] The server then selects the appropriate ad material from its advertising database based on the analysis results. This process uses recommendation algorithms such as Google AdSense API to select the most suitable ad based on the video content and target audience. For example, it may determine that a fashion brand ad is best suited for an urban scene.
[0526] Next, the server uses a generative AI model (such as DALL-E) to composite the selected advertising material into the video. Specifically, the server analyzes the color, lighting, and resolution of the video and adjusts it so that the advertising material is naturally integrated into the frame. At this time, the server uses prompts to the generative AI model to generate the desired results. For example, the server might use the prompt, "Generate a fashion brand advertisement on a billboard in an urban scene and display it naturally."
[0527] The completed video file is generated by the server, encoded, and uploaded to the distribution server via an HTTP POST request, allowing users to seamlessly watch the video with embedded ads through the video distribution platform.
[0528] When a user selects a video they want to watch on a video streaming platform and presses the play button, the device sends an HTTP GET request to the streaming server and receives the generated video file. The device decodes the video file and performs appropriate buffering during playback, providing the user with a seamless viewing experience. The user will notice advertisements embedded naturally within the original video, but the viewing experience is not interrupted.
[0529] As a specific example, the following prompt sentence is used:
[0530] "Generating and displaying fashion brand advertisements naturally on billboards in urban scenes"
[0531] "Naturally incorporate car ads into the background of movie scenes"
[0532] Users can enjoy the video without feeling uncomfortable by watching videos with these advertisements inserted naturally. In this way, the invention improves the user's viewing experience and maximizes the effectiveness of advertising.
[0533] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0534] Step 1:
[0535] Video reception and analysis
[0536] The server receives video files from users or content providers via HTTP requests or FTP. It uses the OpenCV library to break down the received video files into frames and analyze the objects and scenes contained in each frame. For example, it identifies color, lighting, and specific objects (buildings, signs, cars, etc.). The input video file is in digital format, and the analysis results for each frame are obtained as output.
[0537] Step 2:
[0538] Selection of advertising materials
[0539] The server selects appropriate advertising materials from an advertising database based on the analysis results. Recommendation algorithms such as the Google AdSense API are used to select the most suitable advertisements for the video content, scene, and target audience. For example, it may determine that a fashion brand advertisement is most suitable for an urban scene. The analysis results and scene information are used as input, and the selected advertising materials are generated as output.
[0540] Step 3:
[0541] Content adjustment and synthesis with generative AI models
[0542] The server uses a generative AI model to naturally incorporate the selected advertising material into the video. A prompt is created and input to the generative AI model (e.g., DALL-E). The prompt is used as an example: "Generate a fashion brand advertisement on a billboard in an urban scene and display it naturally." The generative AI model generates the required advertising images and uses OpenCV to composite the advertising material into the video frame. The prompt and advertising material are used as input, and a frame with the advertisement composited is obtained as output.
[0543] Step 4:
[0544] Creating and uploading a new video file
[0545] The server generates a new video file from all composited frames, encodes and compresses it, and saves it in the specified format (e.g., MP4 or MKV). It then uploads the resulting video file to the distribution server via an HTTP POST request. It uses each composited frame as input and generates a new video file as output, which is then uploaded to the distribution server.
[0546] Step 5:
[0547] Requesting and Receiving Videos
[0548] A user selects a video they want to watch on a video distribution platform and presses the play button. This causes the user device to send an HTTP GET request to the distribution server. A new video file with an embedded advertisement is sent from the distribution server to the user device as an HTTP response. The user's request information is used as input, and the new video file is provided to the user device as output.
[0549] Step 6:
[0550] Play video
[0551] The user device decodes the received video file and plays it through the appropriate player. The device buffers during playback to ensure a seamless viewing experience for the user. Advertisements are displayed seamlessly between the original video content. The new video file is used as input, and a seamlessly played video is provided as output.
[0552] The above steps realize a system that inserts advertisements naturally into video content, providing users with a natural viewing experience.
[0553] (Application example 1)
[0554] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0555] Conventional methods for inserting advertisements into video content often interrupt the viewer's experience, resulting in limited advertising effectiveness. Furthermore, it is difficult to integrate advertisements in a visually natural way, often causing viewers to feel uncomfortable. The present invention aims to provide a method for seamlessly and naturally inserting advertisements into video content without interrupting the viewer's experience.
[0556] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0557] In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for selecting appropriate advertising materials from an advertising database, means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine the advertising materials with the video, means for seamlessly receiving and playing the generated video file at the user terminal, and means for analyzing the effectiveness of the advertisements and generating a report of the results. This enables visually natural and seamless advertisement insertion, making it possible to provide video content without interrupting the viewer's experience while increasing the effectiveness of the advertisements.
[0558] "Video Content" means a digital video file containing visual and audio information.
[0559] A "scene" refers to a series of frames in video content, and signifies a sequence of images at a specific location or time axis.
[0560] An "advertising database" is a digital repository for storing advertising materials, and is a system that stores a wide variety of advertising data.
[0561] "Advertising Materials" refers to digital content such as images, videos, and text used as advertising.
[0562] "Generative AI" refers to algorithms or models that use artificial intelligence techniques to generate new data or content.
[0563] A "distribution server" is a centralized computer system that provides digital content to user terminals over a network.
[0564] "User terminal" refers to a device, such as a smartphone, tablet, or PC, that allows a user to access digital content via the Internet.
[0565] "Seamless" refers to a state in which the continuity of operations and actions is uninterrupted, and is carried out smoothly without any sense of discomfort.
[0566] "Analysis" refers to the methods and processes used to analyze data and information and understand its structure and meaning.
[0567] "Synthesis" refers to the process of combining different digital content to create new data or images.
[0568] "Effectiveness analysis" refers to statistical or quantitative research or analysis conducted to evaluate the effectiveness or impact of advertising.
[0569] "Report" means a document that summarizes the results of research and analysis and organizes them visually and written.
[0570] The system embodying the present invention performs a series of processes for inserting advertisements into video content in a natural way. This system is mainly composed of a server and a user terminal.
[0571] Server Processing
[0572] Video reception and analysis
[0573] The server first receives the video content sent from the user device, then analyzes each frame of the video using a video analysis algorithm (e.g., OpenCV) to identify scenes and objects suitable for inserting advertisements, thereby determining the advertisement insertion points.
[0574] Selection of advertising materials
[0575] The server then selects appropriate advertising material from its advertising database based on the analysis results, using algorithms to select ads that fit the video's content and target audience. For example, fashion brand ads are often selected for city scenes.
[0576] Content Alignment and Composition
[0577] The server uses a generative AI model (e.g., moviepy) to seamlessly blend the selected ad material into the video. It analyzes the color, lighting, and resolution of the video and adjusts it so that the ad material appears as a natural part of the video. This step provides a seamless visual experience for the viewer.
[0578] Creating and uploading a new video file
[0579] A new video file containing the advertisement is generated and uploaded to the distribution server, which then provides the new video file in response to a request from the user terminal.
[0580] User terminal processing
[0581] Requesting and Receiving Videos
[0582] When the user selects a video that he or she wants to watch, the user terminal sends a request to the distribution server and receives a new video file that includes advertisements.
[0583] Play video
[0584] The device then plays the received video file, buffering and decoding the video as it plays, providing a seamless viewing experience with no sense of incongruity between the ad and the original video content.
[0585] User Experience
[0586] Users select a video they want to watch on a video streaming platform and start playing it. While the video is playing, they may notice advertisements inserted naturally, but they do not interrupt the viewing experience. For example, they may see a billboard advertisement displayed naturally within a movie scene. This allows users to enjoy the video without feeling stressed by advertisements.
[0587] Specific examples
[0588] An example of a prompt sentence is as follows:
[0589] Video Analysis Prompt
[0590] Identify suitable scenes in the input video for ad insertion and composite the latest ads appropriately into those locations.
[0591] Advertisements used:
[0592] The latest advertising video.
[0593] Expected Results:
[0594] A new video file is generated in which the advertisements are seamlessly integrated into the original footage.
[0595] As described above, the system of the present invention provides a series of processes for inserting advertisements into video content in a natural way, thereby providing users with a seamless viewing experience.
[0596] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0597] Step 1:
[0598] Video reception and analysis
[0599] The server receives video content sent from the user's device. It analyzes the received video frame by frame using a video analysis algorithm (e.g., OpenCV) to identify scenes and objects suitable for inserting advertisements. The input to this step is the received video content, and the output is information identifying the ad insertion points. Specifically, it determines where advertisements can be inserted naturally based on the objects and background information recognized through frame-by-frame analysis.
[0600] Step 2:
[0601] Selection of advertising materials
[0602] The server selects appropriate advertising material from the advertising database based on the ad insertion point identified in step 1. This selection uses an algorithm to select advertisements that match the video content information and target audience information. The input to this step is the identification information of the ad insertion point, and the output is the selected advertising material. Specifically, the server searches the advertising database for the most suitable advertisement and extracts the corresponding advertising material.
[0603] Step 3:
[0604] Content Alignment and Composition
[0605] The server uses a generative AI model (e.g., moviepy) to composite the selected ad material into the video in a natural way. At this time, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the ad material blends in with the original footage. The input for this step is the video content and the selected ad material, and the output is a new video file with the ad composited into it. Specifically, the color tone and transparency of the ad material are adjusted based on the color tone and light direction of each frame of the video, resulting in a seamless composite.
[0606] Step 4:
[0607] Creating and uploading a new video file
[0608] The server generates a new video file containing the composited advertisement and uploads the video file to the distribution server. The input of this step is the new video file with the composited advertisement, and the output is the video file uploaded to the distribution server. Specifically, the generated video file is encoded into a specified format and transferred to the distribution server.
[0609] Step 5:
[0610] Requesting and Receiving Videos
[0611] When the user selects a video they want to watch, the user terminal sends a request to the distribution server and receives a new video file including advertisements. The input to this step is the user request, and the output is the video file received from the distribution server. Specifically, an HTTP request is sent based on the user's instructions, and the new video file is received as a response.
[0612] Step 6:
[0613] Play video
[0614] The user device plays the received video file. During video playback, buffering and decoding are performed, maintaining a seamless connection between the advertisement and the original video content. The input for this step is the received video file, and the output is a seamlessly played video. Specifically, the video player function is used to buffer and decode the video using a codec, ensuring smooth playback.
[0615] Step 7:
[0616] Providing a viewing experience
[0617] A user selects a video they want to watch on a video streaming platform and starts playback. While the video is playing, they may notice advertisements inserted naturally, but the viewing experience is not interrupted. Specifically, they can enjoy the video continuously while viewing billboard advertisements that are naturally displayed within movie scenes. The input for this step is a video file received from the streaming server, and the output is a video that plays continuously without interruption.
[0618] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0619] Explaining program processing in natural language
[0620] server
[0621] 1. Receiving and analyzing the video
[0622] The server receives the video content, performs video decoding to convert each frame into an analyzable format, and runs video analysis algorithms to identify important scenes and objects in the video (e.g., building walls, billboards, sides of cars, etc.) to find suitable points for inserting advertisements.
[0623] 2. Acquiring Emotion Data
[0624] The server uses an emotion engine to obtain the user's emotions in real time while watching a video, which includes technology to determine the user's emotional state by analyzing their facial expressions, tone of voice, and eye movements.
[0625] 3. Selection of advertising materials
[0626] The server selects appropriate advertising materials from the advertising database based on the content of the video, the target audience, and the acquired user emotion data. For example, if the user is enjoying watching the video, entertainment-related advertisements may be deemed appropriate.
[0627] 4. Content Coordination and Synthesis
[0628] The server uses generative AI to composite the selected advertising material into specific points in the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it to make the advertisement appear natural. For example, a fashion brand advertisement can be composited onto a building sign, making it appear as if it were part of the original footage.
[0629] 5. Create and upload a new video file
[0630] A new video file containing the generated advertisement is generated and uploaded to a distribution server, allowing users to seamlessly watch the video with the embedded advertisement through a distribution platform.
[0631] Terminal
[0632] 1. Requesting and Receiving Videos
[0633] When the user selects a video they want to watch, the device sends a request to the distribution server, which then receives a new video file with an embedded advertisement.
[0634] 2. Play the video
[0635] The device decodes and buffers the received video file for playback.
[0636] 3. Seamless Ad Display
[0637] The device plays the video and performs playback processing so that the advertisements in the video and the original content are displayed seamlessly, providing the user with a seamless viewing experience.
[0638] User
[0639] 1. Select a video
[0640] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[0641] 2. Watching videos
[0642] A user watches a video playing on their device and is presented with ads that are naturally embedded within the video, without interrupting the normal viewing experience. For example, a user might watch a movie with a billboard ad or background ad that is naturally embedded within the video.
[0643] 3. Emotional Feedback
[0644] Real-time user emotional data is collected and used to select and adjust advertising materials, which may even dynamically change based on the user's emotional state while they are watching the video.
[0645] 4. Continue watching
[0646] Users can enjoy videos continuously without feeling interrupted by advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[0647] Specific examples
[0648] Server Processing
[0649] 1. The server receives the movie footage and performs video analysis, for example, identifying city scenes with billboards suitable for inserting advertisements.
[0650] 2. The server uses an emotion engine to obtain the user's emotion in real time while watching. For example, if the user is enjoying the video, the emotion is judged to be positive.
[0651] 3. Based on the positive emotional state, the server determines that an entertainment advertisement (e.g., a trailer for a new movie) would be a good fit for this scene.
[0652] 4. The generative AI composites a movie trailer onto a specific billboard, integrating it into the scene.
[0653] 5. Upload the completed composite video file to the distribution server.
[0654] Terminal handling
[0655] 1. The user device accesses the distribution platform, selects the movie, and starts playback.
[0656] 2. City scenes from the film play, with entertainment ads on appropriate billboards, making them feel like part of the regular billboard scene.
[0657] 3. The user device provides a seamless viewing experience with no sense of incongruity between the advertisements and the movie.
[0658] User Experience
[0659] 1. The user begins watching a movie and notices a billboard advertisement that is displayed naturally within the city scene.
[0660] 2. The movie is not interrupted, allowing users to remain immersed in the story while watching ads, and the ads are selected based on the user's emotional state, further enhancing the viewing experience.
[0661] 3. This allows users to enjoy movies without being stressed by advertisements.
[0662] The processing flow will be explained below.
[0663] Server Processing
[0664] Step 1:
[0665] The server receives the video content and performs video decoding to convert each frame into an analyzable format.
[0666] Step 2:
[0667] The server runs video analysis algorithms to identify important scenes and objects in the video, such as building walls, billboards, or the sides of cars, and find suitable spots for inserting ads.
[0668] Step 3:
[0669] The server uses an emotion engine to obtain real-time emotional data from users watching videos. This emotional data is determined by analyzing the user's facial expressions, tone of voice, eye movements, etc.
[0670] Step 4:
[0671] The server selects appropriate advertising materials from the advertising database based on the content of the video, the target audience, and the acquired user emotion data. For example, if the user emotion indicates enjoyment, entertainment-related advertisements will be selected.
[0672] Step 5:
[0673] The server uses generative AI to composite the selected ad material into specific points in the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the ad looks natural with the original footage. For example, a movie trailer ad can be composited onto a building billboard.
[0674] Step 6:
[0675] The server applies a compositing process to generate a new video file with ads embedded naturally and unobtrusively into the original content.
[0676] Step 7:
[0677] The server uploads the generated new video file to the distribution server, making it available for viewing by users.
[0678] Terminal handling
[0679] Step 1:
[0680] The user terminal accesses the distribution platform and selects the video that the user wants to watch.
[0681] Step 2:
[0682] The terminal transmits a request for the selected video and receives a new video file with an embedded advertisement from the distribution server.
[0683] Step 3:
[0684] The device decodes and buffers the received video file for playback.
[0685] Step 4:
[0686] The device plays the video and processes it so that the advertisements in the video and the original content are displayed seamlessly.
[0687] User Action
[0688] Step 1:
[0689] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[0690] Step 2:
[0691] A user watches a video playing on their device, and an ad appears, naturally embedded within the video, without interrupting the normal viewing experience. For example, a user might watch a billboard ad that is naturally embedded within a movie.
[0692] Step 3:
[0693] Real-time user emotional data is collected and used to select and adjust advertising materials, which may even dynamically change based on the user's emotional state while they are watching the video.
[0694] Step 4:
[0695] Users can enjoy videos continuously without feeling interrupted by advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[0696] Example 2
[0697] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0698] When inserting advertisements into video content, there is a need for a method to display advertisements naturally without disrupting the user's viewing experience. There is also a need for a method to appropriately select advertisements based on the viewer's emotional state and improve the viewing experience. However, conventional systems have difficulty in utilizing user emotional data to select advertisements and integrate them into videos in a natural way.
[0699] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for acquiring user emotion data and selecting advertising material based on the data, means for selecting appropriate advertising material from an advertisement database based on the identified points and the user emotion data, means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine the advertising material into the video, and means for generating a new video file including the combined advertisement and uploading it to a distribution device. This makes it possible to naturally combine optimal advertisements based on the user's emotions into the video, improving the user's viewing experience.
[0700] "Video Content" means a media file that combines video and audio and is provided for the enjoyment of viewers.
[0701] "Scene analysis" is the process of decoding each frame in a video and identifying important objects and points of interest.
[0702] "Good ad insertion points" are specific scenes or locations in a video where ads can be displayed naturally and effectively.
[0703] "User emotion data" is information indicating the emotional state of the viewer obtained from facial expressions, gaze, voice, and the like.
[0704] "Advertising Materials" means advertising footage, images, or other media content for display.
[0705] An "advertising database" is a searchable database that stores various advertising materials.
[0706] "Analysis and adjustment of color, lighting, and resolution" refers to the process of analyzing the quality and visual elements of a video and making adjustments to make the ads appear natural.
[0707] "Generative AI" is a type of artificial intelligence that is a technology that generates advertising materials and synthesizes them into videos.
[0708] "Distribution device" refers to the server and network infrastructure used to deliver the generated video files to users.
[0709] "Seamless playback" refers to a state in which the video and advertisements are played back consecutively, allowing the user to view them as a single entity without feeling any discomfort.
[0710] The "new video file" is a file in which the advertisement has been synthesized and the original video and advertisement are integrated.
[0711] MODE FOR CARRYING OUT THE INVENTION
[0712] The system of this invention receives and analyzes video content, and automatically inserts advertisements based on user emotion data to improve the viewing experience. The system is mainly composed of three elements: a server, a terminal, and a user.
[0713] server
[0714] The server receives the video content and securely stores it using Amazon S3. The stored video data is decoded into individual frames using FFmpeg. OpenCV is then used to analyze the video and identify important scenes and objects. For example, it can detect billboards suitable for inserting advertisements in urban scenes. The server then collects user emotion data using Microsoft Azure's Emotion API. This technology analyzes the user's facial expressions, tone of voice, eye movements, and other factors to determine their emotional state in real time.
[0715] The server then uses Google Cloud's BigQuery to select appropriate ad materials from the ad database. The selected ad materials are then composited into specific points in the video using OpenAI's generative AI. Blender is used to analyze the color tone, lighting, and resolution of the video and adjust them to make the ads look natural. Finally, a new video file containing the composited ads is generated and uploaded to the distribution device via Amazon CloudFront. Users can then watch the video with the embedded ads.
[0716] Terminal
[0717] When a user selects a video they want to watch, the device sends a request to the distribution server. In response to the request, a new video file with embedded ads is received from Amazon CloudFront. The received video file is decoded and buffered using VLCKit, and then played back seamlessly. This allows the user to enjoy a natural viewing experience without any sense of incongruity between the video and the ads.
[0718] User
[0719] Users select a video they want to watch on the streaming platform and press the play button. This causes the device to send a request to the server and receive a video file containing the advertisement. As users watch the video on their device, they notice the advertisements, which are naturally embedded in billboards and backgrounds, without interrupting the normal viewing experience. In addition, users' emotional data is collected in real time and used to select and adjust advertising materials. Advertisements may also be dynamically changed based on the user's emotional state, further improving the viewing experience.
[0720] Specific examples
[0721] Server Processing
[0722] 1. Receive movie data and store it in Amazon S3.
[0723] 2. Decode the movie using FFmpeg and identify billboards in urban scenes with OpenCV.
[0724] 3. Use Microsoft Azure's Emotion API to analyze the user's facial expressions in real time and obtain positive emotions.
[0725] 4. Select entertainment advertising materials based on positive emotional states using Google Cloud BigQuery.
[0726] 5. Synthesize ads naturally using OpenAI's generative AI and Blender.
[0727] 6. Upload the composite video file to Amazon CloudFront and display it.
[0728] Terminal handling
[0729] 1. The user selects a movie and sends a request to the distribution server.
[0730] 2. Receive new video files from Amazon CloudFront.
[0731] 3. Use VLCKit to decode and buffer the video.
[0732] 4. Seamlessly play videos with natural ad insertions.
[0733] User Experience
[0734] 1. The user selects a movie and presses the play button.
[0735] 2. While watching the movie, you notice billboard advertisements appearing in city scenes.
[0736] 3. Watch movies without being interrupted by ads.
[0737] 4. While watching, ads may dynamically change based on emotions.
[0738] Prompt Sentence Examples
[0739] 1. "Can you give me an example of a diorama scene that displays entertainment advertisements for a movie?"
[0740] 2. "Please explain techniques for inserting ads naturally into videos, especially those that leverage sentiment data."
[0741] keyword
[0742] Generative AI model, prompt sentence
[0743] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0744] Step 1: Receive and save the video
[0745] The server receives video content from users. For example, a user uploads a movie. This received video data is stored in Amazon S3. This ensures that the data is securely stored and can be used for further processing. The input video data is stored in Amazon S3 as a stored output.
[0746] Step 2: Video decoding and analysis
[0747] The server decodes the video data stored in Amazon S3 into individual frames using FFmpeg. Each decoded frame undergoes video analysis using OpenCV. Specifically, it identifies objects and scenes suitable for inserting advertisements, such as billboards in urban scenes. The input of this step is the video data retrieved from Amazon S3, and the output is the analyzed frame data.
[0748] Step 3: Obtaining user emotion data
[0749] The server uses Microsoft Azure's Emotion API to obtain the user's emotional data in real time, which includes technology that analyzes the user's facial expressions, tone of voice, eye movements, etc. to determine their emotional state. The input for this step is the user's video feed and audio data, and the output is the analyzed emotional data.
[0750] Step 4: Selecting advertising materials
[0751] The server uses Google Cloud's BigQuery to select appropriate advertising materials based on the video content, target audience, and acquired user emotion data. For example, entertainment-related ads are selected for users who show positive emotions. The input of this step is the analyzed frame data and emotion data, and the output is the selected advertising materials.
[0752] Step 5: Adjust and combine content
[0753] The server uses OpenAI's generative AI to composite the selected advertising material into specific points in the video. Blender is used to analyze the color, lighting, and resolution of the video and adjust the video to make the advertisement appear natural. For example, a fashion brand advertisement can be composited naturally onto a billboard in an urban scene. The input to this step is the selected advertising material and the analyzed frame data, and the output is a new frame data with the advertisement composited into it.
[0754] Step 6: Generate and upload a new video file
[0755] The server generates a new video file containing the advertisements and uploads it to the distribution device via Amazon CloudFront. The uploaded video file is then accessible on the user's device. The input of this step is the new frame data with the advertisements mixed in, and the output is the final video file.
[0756] Step 7: Request and receive video
[0757] When the user selects a video to watch, the device sends a request to the distribution server. It receives a new video file delivered by Amazon CloudFront. The input of this step is the video viewing request, and the output is the received video file.
[0758] Step 8: Play the video
[0759] The device uses VLCKit to decode and buffer the received video file, allowing the user to seamlessly watch the video. The input of this step is the received video file, and the output is a decoded, buffered, and playable video.
[0760] Step 9: Obtaining and using user emotional feedback
[0761] While the user is watching the video, the server continues to capture the user's emotional data and use it to select and adjust advertising materials, which may be dynamically changed according to the user's real-time emotional state. The input of this step is real-time user emotional data, and the output is adapted advertising materials.
[0762] (Application example 2)
[0763] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0764] Conventional video ad insertion systems use simple ad synthesis methods, which often detract from the user's viewing experience. In particular, ads are displayed uniformly regardless of the video scene or the user's emotional state, which makes the ads visually unnatural and makes the user feel uncomfortable. Furthermore, there is a need for more personalized ad delivery by utilizing user emotional data to select ads.
[0765] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving video content and analyzing scenes in the video to identify points suitable for inserting advertisements; means for selecting appropriate advertising material from an advertising database based on the identified points; means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine advertising material into the video; means for analyzing user emotional data in real time and selecting an optimal advertisement based on the analysis results; means for combining the selected advertising material into specific points in the video and using a generative AI model to adjust the combination so that it is seamless; and means for generating a new video file including the combined advertisement and uploading it to a distribution server. This enables personalized advertising insertion based on the user's emotional state.
[0766] "Video content" refers to video data that provides information to users visually and audibly.
[0767] A "scene" is a set of consecutive video frames in a specific time range within a video.
[0768] "Advertisement" means visual or audio content intended to promote a product or service.
[0769] A "point" refers to a suitable location in a particular scene or frame of a video for inserting an advertisement.
[0770] "Means" refer to the methods or techniques used to achieve a particular goal.
[0771] A "database" is a collection of information that organizes large amounts of information and makes it possible to search it efficiently.
[0772] "Hue" is an attribute related to the color tone and color combination of an image.
[0773] "Lighting" refers to the intensity and directionality of the light used in the video.
[0774] "Resolution" is a measure of the ability to express detail in an image, and is generally expressed in terms of the number of pixels.
[0775] "Compositing" refers to the process of combining different video materials into one video.
[0776] "Emotion data" is information about the user's emotional state obtained from facial expressions, tone of voice, and the like.
[0777] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to generate new data or content.
[0778] "Uploading" is the act of transferring data from a local environment to a remote server.
[0779] "Terminal" refers to a computing device that is directly operated by a user, including smartphones and personal computers.
[0780] "Seamless" means that the operation is natural, without any gaps or interruptions.
[0781] "Real-time" refers to data acquisition and processing occurring immediately.
[0782] The system for implementing the invention is composed of three elements: a server, a user terminal, and a user. The specific operation process of this system is as follows:
[0783] Server Processing
[0784] The server receives the video content and analyzes it frame by frame. Using video analysis algorithms, it analyzes each scene in the video and identifies important elements (e.g., billboards or building walls) to find the optimal points for inserting ads.
[0785] The server then uses an emotion engine to analyze the user's emotional data in real time. It analyzes the user's facial image and tone of voice to understand their emotional state. Based on this information, the server selects the most suitable advertising material from its advertising database.
[0786] The system uses a generative AI model to seamlessly blend the selected ad material into the video, adjusting color, lighting, and resolution to seamlessly insert the ad material at specific points in the video. Finally, the server generates the composite video file and uploads it to the distribution server.
[0787] The specific hardware used is high-performance server equipment, and the software includes video analysis algorithms, emotion engines, advertising databases, and generative AI models.
[0788] User terminal processing
[0789] When a user selects a video they want to watch, the newly synthesized video file is received from the distribution server. The user's device decodes the video file and plays it back seamlessly, providing a seamless viewing experience for the user.
[0790] The specific hardware used is the user's device, such as a smartphone or PC, and the software includes the video playback application and decoding algorithm.
[0791] User Experience
[0792] Users watch videos through a video distribution platform, and ads are displayed naturally embedded within the video, so the viewing experience is not interrupted. In addition, user emotional data is acquired in real time, making it possible to dynamically display appropriate ads.
[0793] Examples of prompt statements
[0794] 1. "Analyze the user's emotional state using facial images and audio data."
[0795] 2. "Please identify scenes in your video file where billboards or other advertising can be inserted."
[0796] 3. "Choose the best ad based on the user's current emotional state."
[0797] 4. "Please integrate the selected ads into the video scenes in a natural way."
[0798] This format allows users to enjoy a highly satisfying video viewing experience without feeling stressed by advertisements.
[0799] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0800] Step 1:
[0801] The server receives video content and uses video analysis algorithms to convert each frame into an analyzable format. The input is a video file, and the output is frame-by-frame video data. Specifically, the server breaks down the video into frames and identifies important scenes and objects in each frame.
[0802] Step 2:
[0803] The server acquires the user's emotional data in real time. The input is the user's facial image and voice data, and the output is the user's emotional state (e.g., joy, excitement, sadness). Specifically, the server uses an emotion engine to analyze the user's facial expressions and tone of voice to determine the user's emotional state.
[0804] Step 3:
[0805] The server selects appropriate advertising materials from an advertising database based on video analysis data and user emotional data. The input is the analyzed scene data and the user's emotional state, and the output is the selected advertising materials. Specifically, the server uses an advertising identification algorithm to select the advertisement that best suits the user's emotional state.
[0806] Step 4:
[0807] The server utilizes a generative AI model to seamlessly composite selected ad material into specific points in the video. The input is the selected ad material and analyzed scene data, and the output is frame data with the ad composited. Specifically, the server adjusts the color tone, lighting, and resolution of the video to composite the ad material naturally.
[0808] Step 5:
[0809] The server generates a new video file containing the composited advertisements and uploads it to the distribution server. The input is the frame data with the composited advertisements, and the output is a new video file. Specifically, the server reconstructs the frame data into a single video file and transfers it to the distribution server.
[0810] Step 6:
[0811] The user terminal receives a new video file from the distribution server. The input is the video file sent from the distribution server, and the output is the saved video file. Specifically, the user terminal downloads the video file.
[0812] Step 7:
[0813] The user terminal decodes and buffers the received video file and plays it seamlessly. The input is the stored video file, and the output is the played video. Specifically, the user terminal uses a video decoder to buffer the video frame by frame and play it smoothly.
[0814] Step 8:
[0815] The user watches a video played on a device. The input is the video being played, and the output is the user's visual and auditory experience. The specific experience is that the user can enjoy the video containing naturally-mixed advertisements without any sense of incongruity.
[0816] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0817] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0818] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0819] [Third embodiment]
[0820] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0821] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0822] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0823] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0824] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0825] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0826] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0827] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0828] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0829] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0830] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0831] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0832] Explaining program processing in natural language
[0833] server
[0834] 1. Receiving and analyzing the video
[0835] The server receives the video content and analyzes each frame of the video, using video analysis algorithms to identify specific scenes or objects within the video (e.g., building walls, billboards, sides of cars, etc.) that are deemed suitable points for inserting advertisements.
[0836] 2. Selection of advertising materials
[0837] Based on the analysis results, the server selects the appropriate advertising material from its advertising database. This selection is based on an algorithm that selects the most suitable advertisement for the video content and target audience. For example, if a fashion brand advertisement is determined to be the most suitable for an urban scene, that advertisement will be selected.
[0838] 3. Content Coordination and Synthesis
[0839] The server uses generative AI to composite the selected advertising material into the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the advertising material is naturally integrated into the video. For example, a fashion brand advertisement can be composited onto a building sign, making it appear as if it were part of the original footage.
[0840] 4. Create and upload a new video file
[0841] A new video file containing the generated advertisement is generated and uploaded to a distribution server, allowing users to seamlessly watch the video with the embedded advertisement through a distribution platform.
[0842] Terminal
[0843] 1. Requesting and Receiving Videos
[0844] When a user selects a video they want to watch, the device sends a request to the distribution server, which then receives a new video file with an embedded advertisement.
[0845] 2. Play the video
[0846] The device then plays the received video file. During playback, the device buffers and decodes the video, providing a seamless viewing experience. There is no gap between the ads and the original video content.
[0847] User
[0848] 1. Select a video
[0849] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[0850] 2. Watching videos
[0851] A user watches a video playing on their device and notices ads that appear naturally within the video, but do not interrupt the normal viewing experience. For example, a user may see billboard ads or background ads that are naturally embedded in a movie.
[0852] 3. Continue watching
[0853] Users can enjoy videos continuously without interruptions caused by advertisements. They can immerse themselves in the content without feeling stressed by advertisements.
[0854] Specific examples
[0855] Server Processing
[0856] 1. The server receives the movie footage and performs video analysis, for example, identifying city scenes with billboards suitable for inserting advertisements.
[0857] 2. The server determines that an advertisement for a fashion brand would be ideal for this scene.
[0858] 3. The generative AI synthesizes a fashion brand's advertising video onto a specific billboard, integrating it into the scene.
[0859] 4. Upload the completed composite video file to the distribution server.
[0860] Terminal handling
[0861] 1. The user device accesses the distribution platform, selects the movie, and starts playback.
[0862] 2. City scenes from the film play, with fashion brand advertisements displayed on appropriate billboards, making them feel like part of a regular billboard.
[0863] 3. The user device provides a seamless viewing experience with no sense of incongruity between the advertisements and the movie.
[0864] User Experience
[0865] 1. The user begins watching a movie and notices a billboard advertisement that is displayed naturally within the city scene.
[0866] 2. The movie is not interrupted and users can watch ads while still being immersed in the story.
[0867] 3. This allows users to enjoy movies without being stressed by advertisements.
[0868] The processing flow will be explained below.
[0869] Server Processing
[0870] Step 1:
[0871] The server receives the video content and performs video decoding to convert each frame into an analyzable format.
[0872] Step 2:
[0873] The server runs video analysis algorithms to identify important scenes and objects in the video, such as building walls, billboards, or the sides of cars, and finds suitable spots for inserting ads.
[0874] Step 3:
[0875] Based on the analysis results, the server lists the identified points and plans to insert advertisements appropriate for those points.
[0876] Step 4:
[0877] The server selects appropriate advertising material from the advertising database based on the video content and target audience, and stores the selected advertising material in temporary memory.
[0878] Step 5:
[0879] The server uses generative AI to composite the selected ad material into specific points in the video, which involves analyzing the color, lighting, and resolution of the video and adjusting them to make the ad appear natural.
[0880] Step 6:
[0881] The server applies the compositing process and generates a new video file with the ads embedded naturally and unobtrusively into the original content.
[0882] Step 7:
[0883] The server uploads the newly generated video file to the distribution server, making it available for viewing by users.
[0884] Terminal handling
[0885] Step 1:
[0886] The user terminal accesses the distribution platform and selects the video that the user wants to watch.
[0887] Step 2:
[0888] The terminal transmits a request for the selected video and receives a new video file with an embedded advertisement from the distribution server.
[0889] Step 3:
[0890] The device decodes and buffers the received video file for playback.
[0891] Step 4:
[0892] The device plays the video and processes it so that the advertisements in the video and the original content are displayed seamlessly.
[0893] User Action
[0894] Step 1:
[0895] Users select the video they want to watch on the video streaming platform and press the play button.
[0896] Step 2:
[0897] Users watch videos played on their device and are presented with ads that are naturally embedded within the video but do not interrupt the normal viewing experience.
[0898] Step 3:
[0899] Users can enjoy videos continuously without interruption due to ads. Because ads are displayed naturally, users can immerse themselves in the content without stress.
[0900] Example 1
[0901] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0902] The challenge is to seamlessly integrate advertisements into video content without causing a sense of incongruity to users. Conventional methods often result in videos with advertisements inserted that look unnatural, detracting from the user's viewing experience.
[0903] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0904] In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for selecting appropriate advertising materials from an advertising database based on the identified points, means for analyzing and adjusting the color tone, lighting, and resolution of the video using a generative AI model to naturally combine the advertising materials into the video, means for generating a new video file including the generated advertisements and uploading it to a distribution server, means for a user to select a video they wish to watch on a video distribution platform, and means for a user terminal to send a request to the distribution server, receive the generated video file, and seamlessly play the video so that there is no sense of incongruity between the advertisements and the original video content. This allows users to enjoy a seamless and natural viewing experience while comfortably watching the advertisements.
[0905] "Video content" is a series of digital data including video and audio, and is a medium for users to view.
[0906] A "scene" refers to a specific frame or group of frames in video content, which represents a specific situation or background.
[0907] "Advertisement" means media content inserted into a video for the purpose of promoting a product or service.
[0908] A "point" refers to a specific position or timing within a scene that is suitable for inserting an advertisement.
[0909] "Means" refers to a method or mechanism for achieving a specific function or operation.
[0910] The "advertising database" is a database in which advertising materials are stored, and includes various advertising contents.
[0911] "Advertising materials" refers to specific media data used as advertising, including images, videos, text, etc.
[0912] "Analysis" is the process of examining data in detail and extracting specific information.
[0913] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to generate new data.
[0914] "Color tone" refers to the adjustment and balance of colors in videos and images.
[0915] "Lighting" refers to the intensity and direction of light in a scene, and is a factor in creating visual atmosphere.
[0916] "Resolution" is an indicator of how clearly the details of videos and images can be displayed.
[0917] A "distribution server" is a server for distributing videos and data to user terminals.
[0918] A "user terminal" is a device used by a user, including a PC, smartphone, tablet, etc.
[0919] A "request" is a request message sent from a user terminal to a server to obtain specific data.
[0920] "Prompt" refers to input text used to prompt a generative AI model to generate a particular result.
[0921] "Seamless" refers to a state without joints, indicating a natural, flowing continuity.
[0922] This invention provides a system for inserting advertisements into video content in a natural way, using a method for analyzing scenes and objects in the video, selecting the most suitable advertisements, and synthesizing them in a natural way. The following describes how to specifically implement this invention.
[0923] The server receives video content from users or content providers via HTTP requests or FTP. The received video content is broken down into frames using video analysis algorithms such as the OpenCV library or TensorFlow, and objects and scenes within each frame are analyzed. For example, distinctive objects such as building walls, signs, and the sides of cars are identified.
[0924] The server then selects the appropriate ad material from its advertising database based on the analysis results. This process uses recommendation algorithms such as Google AdSense API to select the most suitable ad based on the video content and target audience. For example, it may determine that a fashion brand ad is best suited for an urban scene.
[0925] Next, the server uses a generative AI model (such as DALL-E) to composite the selected advertising material into the video. Specifically, the server analyzes the color, lighting, and resolution of the video and adjusts it so that the advertising material is naturally integrated into the frame. At this time, the server uses prompts to the generative AI model to generate the desired results. For example, the server might use the prompt, "Generate a fashion brand advertisement on a billboard in an urban scene and display it naturally."
[0926] The completed video file is generated by the server, encoded, and uploaded to the distribution server via an HTTP POST request, allowing users to seamlessly watch the video with embedded ads through the video distribution platform.
[0927] When a user selects a video they want to watch on a video streaming platform and presses the play button, the device sends an HTTP GET request to the streaming server and receives the generated video file. The device decodes the video file and performs appropriate buffering during playback, providing the user with a seamless viewing experience. The user will notice advertisements embedded naturally within the original video, but the viewing experience is not interrupted.
[0928] As a specific example, the following prompt sentence is used:
[0929] "Generating and displaying fashion brand advertisements naturally on billboards in urban scenes"
[0930] "Naturally incorporate car ads into the background of movie scenes"
[0931] Users can enjoy the video without feeling uncomfortable by watching videos with these advertisements inserted naturally. In this way, the invention improves the user's viewing experience and maximizes the effectiveness of advertising.
[0932] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0933] Step 1:
[0934] Video reception and analysis
[0935] The server receives video files from users or content providers via HTTP requests or FTP. It uses the OpenCV library to break down the received video files into frames and analyze the objects and scenes contained in each frame. For example, it identifies color, lighting, and specific objects (buildings, signs, cars, etc.). The input video file is in digital format, and the analysis results for each frame are obtained as output.
[0936] Step 2:
[0937] Selection of advertising materials
[0938] The server selects appropriate advertising materials from an advertising database based on the analysis results. Recommendation algorithms such as the Google AdSense API are used to select the most suitable advertisements for the video content, scene, and target audience. For example, it may determine that a fashion brand advertisement is most suitable for an urban scene. The analysis results and scene information are used as input, and the selected advertising materials are generated as output.
[0939] Step 3:
[0940] Content adjustment and synthesis with generative AI models
[0941] The server uses a generative AI model to naturally incorporate the selected advertising material into the video. A prompt is created and input to the generative AI model (e.g., DALL-E). The prompt is used as an example: "Generate a fashion brand advertisement on a billboard in an urban scene and display it naturally." The generative AI model generates the required advertising images and uses OpenCV to composite the advertising material into the video frame. The prompt and advertising material are used as input, and a frame with the advertisement composited is obtained as output.
[0942] Step 4:
[0943] Creating and uploading a new video file
[0944] The server generates a new video file from all composited frames, encodes and compresses it, and saves it in the specified format (e.g., MP4 or MKV). It then uploads the resulting video file to the distribution server via an HTTP POST request. It uses each composited frame as input and generates a new video file as output, which is then uploaded to the distribution server.
[0945] Step 5:
[0946] Requesting and Receiving Videos
[0947] A user selects a video they want to watch on a video distribution platform and presses the play button. This causes the user device to send an HTTP GET request to the distribution server. A new video file with an embedded advertisement is sent from the distribution server to the user device as an HTTP response. The user's request information is used as input, and the new video file is provided to the user device as output.
[0948] Step 6:
[0949] Play video
[0950] The user device decodes the received video file and plays it through the appropriate player. The device buffers during playback to ensure a seamless viewing experience for the user. Advertisements are displayed seamlessly between the original video content. The new video file is used as input, and a seamlessly played video is provided as output.
[0951] The above steps realize a system that inserts advertisements naturally into video content, providing users with a natural viewing experience.
[0952] (Application example 1)
[0953] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0954] Conventional methods for inserting advertisements into video content often interrupt the viewer's experience, resulting in limited advertising effectiveness. Furthermore, it is difficult to integrate advertisements in a visually natural way, often causing viewers to feel uncomfortable. The present invention aims to provide a method for seamlessly and naturally inserting advertisements into video content without interrupting the viewer's experience.
[0955] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0956] In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for selecting appropriate advertising materials from an advertising database, means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine the advertising materials with the video, means for seamlessly receiving and playing the generated video file at the user terminal, and means for analyzing the effectiveness of the advertisements and generating a report of the results. This enables visually natural and seamless advertisement insertion, making it possible to provide video content without interrupting the viewer's experience while increasing the effectiveness of the advertisements.
[0957] "Video Content" means a digital video file containing visual and audio information.
[0958] A "scene" refers to a series of frames in video content, and signifies a sequence of images at a specific location or time axis.
[0959] An "advertising database" is a digital repository for storing advertising materials, and is a system that stores a wide variety of advertising data.
[0960] "Advertising Materials" refers to digital content such as images, videos, and text used as advertising.
[0961] "Generative AI" refers to algorithms or models that use artificial intelligence techniques to generate new data or content.
[0962] A "distribution server" is a centralized computer system that provides digital content to user terminals over a network.
[0963] "User terminal" refers to a device, such as a smartphone, tablet, or PC, that allows a user to access digital content via the Internet.
[0964] "Seamless" refers to a state in which the continuity of operations and actions is uninterrupted, and is carried out smoothly without any sense of discomfort.
[0965] "Analysis" refers to the methods and processes used to analyze data and information and understand its structure and meaning.
[0966] "Synthesis" refers to the process of combining different digital content to create new data or images.
[0967] "Effectiveness analysis" refers to statistical or quantitative research or analysis conducted to evaluate the effectiveness or impact of advertising.
[0968] "Report" means a document that summarizes the results of research and analysis and organizes them visually and written.
[0969] The system embodying the present invention performs a series of processes for inserting advertisements into video content in a natural way. This system is mainly composed of a server and a user terminal.
[0970] Server Processing
[0971] Video reception and analysis
[0972] The server first receives the video content sent from the user device, then analyzes each frame of the video using a video analysis algorithm (e.g., OpenCV) to identify scenes and objects suitable for inserting advertisements, thereby determining the advertisement insertion points.
[0973] Selection of advertising materials
[0974] The server then selects appropriate advertising material from its advertising database based on the analysis results, using algorithms to select ads that fit the video's content and target audience. For example, fashion brand ads are often selected for city scenes.
[0975] Content Alignment and Composition
[0976] The server uses a generative AI model (e.g., moviepy) to seamlessly blend the selected ad material into the video. It analyzes the color, lighting, and resolution of the video and adjusts it so that the ad material appears as a natural part of the video. This step provides a seamless visual experience for the viewer.
[0977] Creating and uploading a new video file
[0978] A new video file containing the advertisement is generated and uploaded to the distribution server, which then provides the new video file in response to a request from the user terminal.
[0979] User terminal processing
[0980] Requesting and Receiving Videos
[0981] When the user selects a video that he or she wants to watch, the user terminal sends a request to the distribution server and receives a new video file that includes advertisements.
[0982] Play video
[0983] The device then plays the received video file, buffering and decoding the video as it plays, providing a seamless viewing experience with no sense of incongruity between the ad and the original video content.
[0984] User Experience
[0985] Users select a video they want to watch on a video streaming platform and start playing it. While the video is playing, they may notice advertisements inserted naturally, but they do not interrupt the viewing experience. For example, they may see a billboard advertisement displayed naturally within a movie scene. This allows users to enjoy the video without feeling stressed by advertisements.
[0986] Specific examples
[0987] An example of a prompt sentence is as follows:
[0988] Video Analysis Prompt
[0989] Identify suitable scenes in the input video for ad insertion and composite the latest ads appropriately into those locations.
[0990] Advertisements used:
[0991] The latest advertising video.
[0992] Expected Results:
[0993] A new video file is generated in which the advertisements are seamlessly integrated into the original footage.
[0994] As described above, the system of the present invention provides a series of processes for inserting advertisements into video content in a natural way, thereby providing users with a seamless viewing experience.
[0995] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0996] Step 1:
[0997] Video reception and analysis
[0998] The server receives video content sent from the user's device. It analyzes the received video frame by frame using a video analysis algorithm (e.g., OpenCV) to identify scenes and objects suitable for inserting advertisements. The input to this step is the received video content, and the output is information identifying the ad insertion point. Specifically, it determines where an advertisement can be inserted naturally based on the objects and background information recognized through frame-by-frame analysis.
[0999] Step 2:
[1000] Selection of advertising materials
[1001] The server selects appropriate advertising material from the advertising database based on the ad insertion point identified in step 1. This selection uses an algorithm to select advertisements that match the video content information and target audience information. The input to this step is the identification information of the ad insertion point, and the output is the selected advertising material. Specifically, the server searches the advertising database for the most suitable advertisement and extracts the corresponding advertising material.
[1002] Step 3:
[1003] Content Alignment and Composition
[1004] The server uses a generative AI model (e.g., moviepy) to composite the selected ad material into the video in a natural way. At this time, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the ad material blends in with the original footage. The input for this step is the video content and the selected ad material, and the output is a new video file with the ad composited into it. Specifically, the color tone and transparency of the ad material are adjusted based on the color tone and light direction of each frame of the video, resulting in a seamless composite.
[1005] Step 4:
[1006] Creating and uploading a new video file
[1007] The server generates a new video file containing the composited advertisement and uploads the video file to the distribution server. The input of this step is the new video file with the composited advertisement, and the output is the video file uploaded to the distribution server. Specifically, the generated video file is encoded into a specified format and transferred to the distribution server.
[1008] Step 5:
[1009] Requesting and Receiving Videos
[1010] When the user selects a video they want to watch, the user terminal sends a request to the distribution server and receives a new video file including advertisements. The input to this step is the user request, and the output is the video file received from the distribution server. Specifically, an HTTP request is sent based on the user's instructions, and the new video file is received as a response.
[1011] Step 6:
[1012] Play video
[1013] The user device plays the received video file. During video playback, buffering and decoding are performed, maintaining a seamless connection between the advertisement and the original video content. The input for this step is the received video file, and the output is a seamlessly played video. Specifically, the video player function is used to buffer and decode the video using a codec, ensuring smooth playback.
[1014] Step 7:
[1015] Providing a viewing experience
[1016] A user selects a video they want to watch on a video streaming platform and starts playback. While the video is playing, they may notice advertisements inserted naturally, but the viewing experience is not interrupted. Specifically, they can enjoy the video continuously while viewing billboard advertisements that are naturally displayed within movie scenes. The input for this step is a video file received from the streaming server, and the output is a video that plays continuously without interruption.
[1017] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1018] Explaining program processing in natural language
[1019] server
[1020] 1. Receiving and analyzing the video
[1021] The server receives the video content, performs video decoding to convert each frame into an analyzable format, and runs video analysis algorithms to identify important scenes and objects in the video (e.g., building walls, billboards, sides of cars, etc.) to find suitable points for inserting advertisements.
[1022] 2. Acquiring Emotion Data
[1023] The server uses an emotion engine to obtain the user's emotions in real time while watching a video, which includes technology to determine the user's emotional state by analyzing their facial expressions, tone of voice, and eye movements.
[1024] 3. Selection of advertising materials
[1025] The server selects appropriate advertising materials from the advertising database based on the content of the video, the target audience, and the acquired user emotion data. For example, if the user is enjoying watching the video, entertainment-related advertisements may be deemed appropriate.
[1026] 4. Content Coordination and Synthesis
[1027] The server uses generative AI to composite the selected advertising material into specific points in the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it to make the advertisement appear natural. For example, a fashion brand advertisement can be composited onto a building sign, making it appear as if it were part of the original footage.
[1028] 5. Create and upload a new video file
[1029] A new video file containing the generated advertisement is generated and uploaded to a distribution server, allowing users to seamlessly watch the video with the embedded advertisement through a distribution platform.
[1030] Terminal
[1031] 1. Requesting and Receiving Videos
[1032] When a user selects a video they want to watch, the device sends a request to the distribution server, which then receives a new video file with an embedded advertisement.
[1033] 2. Play the video
[1034] The device decodes and buffers the received video file for playback.
[1035] 3. Seamless Ad Display
[1036] The device plays the video and performs playback processing so that the advertisements in the video and the original content are displayed seamlessly, providing the user with a seamless viewing experience.
[1037] User
[1038] 1. Select a video
[1039] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[1040] 2. Watching videos
[1041] A user watches a video playing on their device and is presented with ads that are naturally embedded within the video, without interrupting the normal viewing experience. For example, a user might watch a movie with a billboard ad or background ad that is naturally embedded within the video.
[1042] 3. Emotional Feedback
[1043] Real-time user emotional data is collected and used to select and adjust advertising materials, which may even dynamically change based on the user's emotional state while they are watching the video.
[1044] 4. Continue watching
[1045] Users can enjoy videos continuously without feeling interrupted by advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[1046] Specific examples
[1047] Server Processing
[1048] 1. The server receives the movie footage and performs video analysis, for example, identifying city scenes with billboards suitable for inserting advertisements.
[1049] 2. The server uses an emotion engine to obtain the user's emotion in real time while watching. For example, if the user is enjoying the video, the emotion is judged to be positive.
[1050] 3. Based on the positive emotional state, the server determines that an entertainment advertisement (e.g., a trailer for a new movie) would be a good fit for this scene.
[1051] 4. The generative AI composites a movie trailer onto a specific billboard, integrating it into the scene.
[1052] 5. Upload the completed composite video file to the distribution server.
[1053] Terminal handling
[1054] 1. The user device accesses the distribution platform, selects the movie, and starts playback.
[1055] 2. City scenes from the film play, with entertainment ads on appropriate billboards, making them feel like part of the regular billboard scene.
[1056] 3. The user device provides a seamless viewing experience with no sense of incongruity between the advertisements and the movie.
[1057] User Experience
[1058] 1. The user begins watching a movie and notices a billboard advertisement that is displayed naturally within the city scene.
[1059] 2. The movie is not interrupted, allowing users to remain immersed in the story while watching ads, and the ads are selected based on the user's emotional state, further enhancing the viewing experience.
[1060] 3. This allows users to enjoy movies without being stressed by advertisements.
[1061] The processing flow will be explained below.
[1062] Server Processing
[1063] Step 1:
[1064] The server receives the video content and performs video decoding to convert each frame into an analyzable format.
[1065] Step 2:
[1066] The server runs video analysis algorithms to identify important scenes and objects in the video, such as building walls, billboards, or the sides of cars, and find suitable spots for inserting ads.
[1067] Step 3:
[1068] The server uses an emotion engine to obtain real-time emotional data from users watching videos. This emotional data is determined by analyzing the user's facial expressions, tone of voice, eye movements, etc.
[1069] Step 4:
[1070] The server selects appropriate advertising materials from the advertising database based on the content of the video, the target audience, and the acquired user emotion data. For example, if the user emotion indicates enjoyment, entertainment-related advertisements will be selected.
[1071] Step 5:
[1072] The server uses generative AI to composite the selected ad material into specific points in the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the ad looks natural with the original footage. For example, a movie trailer ad can be composited onto a building billboard.
[1073] Step 6:
[1074] The server applies the compositing process and generates a new video file with the ads embedded naturally and unobtrusively into the original content.
[1075] Step 7:
[1076] The server uploads the generated new video file to the distribution server, making it available for viewing by users.
[1077] Terminal handling
[1078] Step 1:
[1079] The user terminal accesses the distribution platform and selects the video that the user wants to watch.
[1080] Step 2:
[1081] The terminal transmits a request for the selected video and receives a new video file with an embedded advertisement from the distribution server.
[1082] Step 3:
[1083] The device decodes and buffers the received video file for playback.
[1084] Step 4:
[1085] The device plays the video and processes it so that the advertisements in the video and the original content are displayed seamlessly.
[1086] User Action
[1087] Step 1:
[1088] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[1089] Step 2:
[1090] A user watches a video playing on their device, and an ad appears, naturally embedded within the video, without interrupting the normal viewing experience. For example, a user might watch a billboard ad that is naturally embedded within a movie.
[1091] Step 3:
[1092] Real-time user emotional data is collected and used to select and adjust advertising materials, which may even dynamically change based on the user's emotional state while they are watching the video.
[1093] Step 4:
[1094] Users can enjoy videos continuously without feeling interrupted by advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[1095] Example 2
[1096] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1097] When inserting advertisements into video content, there is a need for a method to display advertisements naturally without disrupting the user's viewing experience. There is also a need for a method to appropriately select advertisements based on the viewer's emotional state and improve the viewing experience. However, conventional systems have difficulty in utilizing user emotional data to select advertisements and integrate them into videos in a natural way.
[1098] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for acquiring user emotion data and selecting advertising material based on the data, means for selecting appropriate advertising material from an advertisement database based on the identified points and the user emotion data, means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine the advertising material into the video, and means for generating a new video file including the combined advertisement and uploading it to a distribution device. This makes it possible to naturally combine optimal advertisements based on the user's emotions into the video, improving the user's viewing experience.
[1099] "Video Content" means a media file that combines video and audio and is provided for the enjoyment of viewers.
[1100] "Scene analysis" is the process of decoding each frame in a video and identifying important objects and points of interest.
[1101] "Good ad insertion points" are specific scenes or locations in a video where ads can be displayed naturally and effectively.
[1102] "User emotion data" is information indicating the emotional state of the viewer obtained from facial expressions, gaze, voice, and the like.
[1103] "Advertising Materials" means advertising footage, images, or other media content for display.
[1104] An "advertising database" is a searchable database that stores various advertising materials.
[1105] "Analysis and adjustment of color, lighting, and resolution" refers to the process of analyzing the quality and visual elements of a video and making adjustments to make the ads appear natural.
[1106] "Generative AI" is a type of artificial intelligence that is a technology that generates advertising materials and synthesizes them into videos.
[1107] "Distribution device" refers to the server and network infrastructure used to deliver the generated video files to users.
[1108] "Seamless playback" refers to a state in which the video and advertisements are played back consecutively, allowing the user to view them as a single entity without feeling any discomfort.
[1109] The "new video file" is a file in which the advertisement has been synthesized and the original video and advertisement are integrated.
[1110] MODE FOR CARRYING OUT THE INVENTION
[1111] The system of this invention receives and analyzes video content, and automatically inserts advertisements based on user emotion data to improve the viewing experience. The system is mainly composed of three elements: a server, a terminal, and a user.
[1112] server
[1113] The server receives the video content and securely stores it using Amazon S3. The stored video data is decoded into individual frames using FFmpeg. OpenCV is then used to analyze the video and identify important scenes and objects. For example, it can detect billboards suitable for inserting advertisements in urban scenes. The server then collects user emotion data using Microsoft Azure's Emotion API. This technology analyzes the user's facial expressions, tone of voice, eye movements, and other factors to determine their emotional state in real time.
[1114] The server then uses Google Cloud's BigQuery to select appropriate ad materials from the ad database. The selected ad materials are then composited into specific points in the video using OpenAI's generative AI. Blender is used to analyze the color tone, lighting, and resolution of the video and adjust them to make the ads look natural. Finally, a new video file containing the composited ads is generated and uploaded to the distribution device via Amazon CloudFront. Users can then watch the video with the embedded ads.
[1115] Terminal
[1116] When a user selects a video they want to watch, the device sends a request to the distribution server. In response to the request, a new video file with embedded ads is received from Amazon CloudFront. The received video file is decoded and buffered using VLCKit, and then played back seamlessly. This allows the user to enjoy a natural viewing experience without any sense of incongruity between the video and the ads.
[1117] User
[1118] Users select a video they want to watch on the streaming platform and press the play button. This causes the device to send a request to the server and receive a video file containing the advertisement. As users watch the video on their device, they notice the advertisements, which are naturally embedded in billboards and backgrounds, without interrupting the normal viewing experience. In addition, users' emotional data is collected in real time and used to select and adjust advertising materials. Advertisements may also be dynamically changed based on the user's emotional state, further improving the viewing experience.
[1119] Specific examples
[1120] Server Processing
[1121] 1. Receive movie data and store it in Amazon S3.
[1122] 2. Decode the movie using FFmpeg and identify billboards in urban scenes with OpenCV.
[1123] 3. Use Microsoft Azure's Emotion API to analyze the user's facial expressions in real time and obtain positive emotions.
[1124] 4. Select entertainment advertising materials based on positive emotional states using Google Cloud BigQuery.
[1125] 5. Synthesize ads naturally using OpenAI's generative AI and Blender.
[1126] 6. Upload the composite video file to Amazon CloudFront and display it.
[1127] Terminal handling
[1128] 1. The user selects a movie and sends a request to the distribution server.
[1129] 2. Receive new video files from Amazon CloudFront.
[1130] 3. Use VLCKit to decode and buffer the video.
[1131] 4. Seamlessly play videos with natural ad insertions.
[1132] User Experience
[1133] 1. The user selects a movie and presses the play button.
[1134] 2. While watching the movie, you notice billboard advertisements appearing in city scenes.
[1135] 3. Watch movies without being interrupted by ads.
[1136] 4. While watching, ads may dynamically change based on emotions.
[1137] Prompt Sentence Examples
[1138] 1. "Can you give me an example of a diorama scene that displays entertainment advertisements for a movie?"
[1139] 2. "Please explain techniques for inserting ads naturally into videos, especially those that leverage sentiment data."
[1140] keyword
[1141] Generative AI model, prompt sentence
[1142] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1143] Step 1: Receive and save the video
[1144] The server receives video content from users. For example, a user uploads a movie. This received video data is stored in Amazon S3. This ensures that the data is securely stored and can be used for further processing. The input video data is stored in Amazon S3 as a stored output.
[1145] Step 2: Video decoding and analysis
[1146] The server decodes the video data stored in Amazon S3 into individual frames using FFmpeg. Each decoded frame undergoes video analysis using OpenCV. Specifically, it identifies objects and scenes suitable for inserting advertisements, such as billboards in urban scenes. The input of this step is the video data retrieved from Amazon S3, and the output is the analyzed frame data.
[1147] Step 3: Obtaining user emotion data
[1148] The server uses Microsoft Azure's Emotion API to obtain the user's emotional data in real time, which includes technology that analyzes the user's facial expressions, tone of voice, eye movements, etc. to determine their emotional state. The input for this step is the user's video feed and audio data, and the output is the analyzed emotional data.
[1149] Step 4: Selecting advertising materials
[1150] The server uses Google Cloud's BigQuery to select appropriate advertising materials based on the video content, target audience, and acquired user emotion data. For example, entertainment-related ads are selected for users who show positive emotions. The input of this step is the analyzed frame data and emotion data, and the output is the selected advertising materials.
[1151] Step 5: Adjust and combine content
[1152] The server uses OpenAI's generative AI to composite the selected advertising material into specific points in the video. Blender is used to analyze the color, lighting, and resolution of the video and adjust the video to make the advertisement appear natural. For example, a fashion brand advertisement can be composited naturally onto a billboard in an urban scene. The input to this step is the selected advertising material and the analyzed frame data, and the output is a new frame data with the advertisement composited into it.
[1153] Step 6: Generate and upload a new video file
[1154] The server generates a new video file containing the advertisements and uploads it to the distribution device via Amazon CloudFront. The uploaded video file is then accessible on the user's device. The input of this step is the new frame data with the advertisements mixed in, and the output is the final video file.
[1155] Step 7: Requesting and receiving video
[1156] When the user selects a video to watch, the device sends a request to the distribution server. It receives a new video file delivered by Amazon CloudFront. The input of this step is the video viewing request, and the output is the received video file.
[1157] Step 8: Play the video
[1158] The device uses VLCKit to decode and buffer the received video file, allowing the user to seamlessly watch the video. The input of this step is the received video file, and the output is a decoded, buffered, and playable video.
[1159] Step 9: Obtaining and using user emotional feedback
[1160] While the user is watching the video, the server continues to capture the user's emotional data and use it to select and adjust advertising materials, which may be dynamically changed according to the user's real-time emotional state. The input of this step is real-time user emotional data, and the output is adapted advertising materials.
[1161] (Application example 2)
[1162] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1163] Conventional video ad insertion systems use simple ad synthesis methods, which often detract from the user's viewing experience. In particular, ads are displayed uniformly regardless of the video scene or the user's emotional state, which makes the ads visually unnatural and makes the user feel uncomfortable. Furthermore, there is a need for more personalized ad delivery by utilizing user emotional data to select ads.
[1164] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving video content and analyzing scenes in the video to identify points suitable for inserting advertisements; means for selecting appropriate advertising material from an advertising database based on the identified points; means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine advertising material into the video; means for analyzing user emotional data in real time and selecting an optimal advertisement based on the analysis results; means for combining the selected advertising material into specific points in the video and using a generative AI model to adjust the combination so that it is seamless; and means for generating a new video file including the combined advertisement and uploading it to a distribution server. This enables personalized advertising insertion based on the user's emotional state.
[1165] "Video content" refers to video data that provides information to users visually and audibly.
[1166] A "scene" is a set of consecutive video frames in a specific time range within a video.
[1167] "Advertisement" means visual or audio content intended to promote a product or service.
[1168] A "point" refers to a suitable location in a particular scene or frame of a video for inserting an advertisement.
[1169] "Means" refer to the methods or techniques used to achieve a particular goal.
[1170] A "database" is a collection of information that organizes large amounts of information and makes it possible to search it efficiently.
[1171] "Hue" is an attribute related to the color tone and color combination of an image.
[1172] "Lighting" refers to the intensity and directionality of the light used in the video.
[1173] "Resolution" is a measure of the ability to express detail in an image, and is generally expressed in terms of the number of pixels.
[1174] "Compositing" refers to the process of combining different video materials into one video.
[1175] "Emotion data" is information about the user's emotional state obtained from facial expressions, tone of voice, and the like.
[1176] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to generate new data or content.
[1177] "Uploading" is the act of transferring data from a local environment to a remote server.
[1178] "Terminal" refers to a computing device that is directly operated by a user, including smartphones and personal computers.
[1179] "Seamless" means that the operation is natural, without any gaps or interruptions.
[1180] "Real-time" refers to data acquisition and processing occurring immediately.
[1181] The system for implementing the invention is composed of three elements: a server, a user terminal, and a user. The specific operation process of this system is as follows:
[1182] Server Processing
[1183] The server receives the video content and analyzes it frame by frame. Using video analysis algorithms, it analyzes each scene in the video and identifies important elements (e.g., billboards or building walls) to find the optimal points for inserting ads.
[1184] The server then uses an emotion engine to analyze the user's emotional data in real time. It analyzes the user's facial image and tone of voice to understand their emotional state. Based on this information, the server selects the most suitable advertising material from its advertising database.
[1185] The system uses a generative AI model to seamlessly blend the selected ad material into the video, adjusting color, lighting, and resolution to seamlessly insert the ad material at specific points in the video. Finally, the server generates the composite video file and uploads it to the distribution server.
[1186] The specific hardware used is high-performance server equipment, and the software includes video analysis algorithms, emotion engines, advertising databases, and generative AI models.
[1187] User terminal processing
[1188] When a user selects a video they want to watch, the newly synthesized video file is received from the distribution server. The user's device decodes the video file and plays it back seamlessly, providing a seamless viewing experience for the user.
[1189] The specific hardware used is the user's device, such as a smartphone or PC, and the software includes the video playback application and decoding algorithm.
[1190] User Experience
[1191] Users watch videos through a video distribution platform, and ads are displayed naturally embedded within the video, so the viewing experience is not interrupted. In addition, user emotional data is acquired in real time, making it possible to dynamically display appropriate ads.
[1192] Examples of prompt statements
[1193] 1. "Analyze the user's emotional state using facial images and audio data."
[1194] 2. "Please identify scenes in your video file where billboards or other advertising can be inserted."
[1195] 3. "Choose the best ad based on the user's current emotional state."
[1196] 4. "Please integrate the selected ads into the video scenes in a natural way."
[1197] This format allows users to enjoy a highly satisfying video viewing experience without feeling stressed by advertisements.
[1198] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1199] Step 1:
[1200] The server receives video content and uses video analysis algorithms to convert each frame into an analyzable format. The input is a video file, and the output is frame-by-frame video data. Specifically, the server breaks down the video into frames and identifies important scenes and objects in each frame.
[1201] Step 2:
[1202] The server acquires the user's emotional data in real time. The input is the user's facial image and voice data, and the output is the user's emotional state (e.g., joy, excitement, sadness). Specifically, the server uses an emotion engine to analyze the user's facial expressions and tone of voice to determine the user's emotional state.
[1203] Step 3:
[1204] The server selects appropriate advertising materials from an advertising database based on video analysis data and user emotional data. The input is the analyzed scene data and the user's emotional state, and the output is the selected advertising materials. Specifically, the server uses an advertising identification algorithm to select the advertisement that best suits the user's emotional state.
[1205] Step 4:
[1206] The server utilizes a generative AI model to seamlessly composite selected ad material into specific points in the video. The input is the selected ad material and analyzed scene data, and the output is frame data with the ad composited. Specifically, the server adjusts the color tone, lighting, and resolution of the video to composite the ad material naturally.
[1207] Step 5:
[1208] The server generates a new video file containing the composited advertisements and uploads it to the distribution server. The input is the frame data with the composited advertisements, and the output is a new video file. Specifically, the server reconstructs the frame data into a single video file and transfers it to the distribution server.
[1209] Step 6:
[1210] The user terminal receives a new video file from the distribution server. The input is the video file sent from the distribution server, and the output is the saved video file. Specifically, the user terminal downloads the video file.
[1211] Step 7:
[1212] The user terminal decodes and buffers the received video file and plays it seamlessly. The input is the stored video file, and the output is the played video. Specifically, the user terminal uses a video decoder to buffer the video frame by frame and play it smoothly.
[1213] Step 8:
[1214] The user watches a video played on a device. The input is the video being played, and the output is the user's visual and auditory experience. The specific experience is that the user can enjoy the video containing naturally-mixed advertisements without any sense of incongruity.
[1215] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1216] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1217] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1218] [Fourth embodiment]
[1219] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1220] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1221] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1222] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1223] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1224] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1225] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1226] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1227] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1228] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1229] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1230] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1231] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1232] Explaining program processing in natural language
[1233] server
[1234] 1. Receiving and analyzing the video
[1235] The server receives the video content and analyzes each frame of the video, using video analysis algorithms to identify specific scenes or objects within the video (e.g., building walls, billboards, sides of cars, etc.) that are deemed suitable points for inserting advertisements.
[1236] 2. Selection of advertising materials
[1237] Based on the analysis results, the server selects the appropriate advertising material from its advertising database. This selection is based on an algorithm that selects the most suitable advertisement for the video content and target audience. For example, if a fashion brand advertisement is determined to be the most suitable for an urban scene, that advertisement will be selected.
[1238] 3. Content Coordination and Synthesis
[1239] The server uses generative AI to composite the selected advertising material into the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the advertising material is naturally integrated into the video. For example, a fashion brand advertisement can be composited onto a building sign, making it appear as if it were part of the original footage.
[1240] 4. Create and upload a new video file
[1241] A new video file containing the generated advertisement is generated and uploaded to a distribution server, allowing users to seamlessly watch the video with the embedded advertisement through a distribution platform.
[1242] Terminal
[1243] 1. Requesting and Receiving Videos
[1244] When a user selects a video they want to watch, the device sends a request to the distribution server, which then receives a new video file with an embedded advertisement.
[1245] 2. Play the video
[1246] The device then plays the received video file. During playback, the device buffers and decodes the video, providing a seamless viewing experience. There is no gap between the ads and the original video content.
[1247] User
[1248] 1. Select a video
[1249] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[1250] 2. Watching videos
[1251] A user watches a video playing on their device and notices ads that appear naturally within the video, but do not interrupt the normal viewing experience. For example, a user may see billboard ads or background ads that are naturally embedded in a movie.
[1252] 3. Continue watching
[1253] Users can enjoy videos continuously without interruptions caused by advertisements. They can immerse themselves in the content without feeling stressed by advertisements.
[1254] Specific examples
[1255] Server Processing
[1256] 1. The server receives the movie footage and performs video analysis, for example, identifying city scenes with billboards suitable for inserting advertisements.
[1257] 2. The server determines that an advertisement for a fashion brand would be ideal for this scene.
[1258] 3. The generative AI synthesizes a fashion brand's advertising video onto a specific billboard, integrating it into the scene.
[1259] 4. Upload the completed composite video file to the distribution server.
[1260] Terminal handling
[1261] 1. The user device accesses the distribution platform, selects the movie, and starts playback.
[1262] 2. City scenes from the film play, with fashion brand advertisements displayed on appropriate billboards, making them feel like part of a regular billboard.
[1263] 3. The user device provides a seamless viewing experience with no sense of incongruity between the advertisements and the movie.
[1264] User Experience
[1265] 1. The user begins watching a movie and notices a billboard advertisement that is displayed naturally within the city scene.
[1266] 2. The movie is not interrupted and users can watch ads while still being immersed in the story.
[1267] 3. This allows users to enjoy movies without being stressed by advertisements.
[1268] The processing flow will be explained below.
[1269] Server Processing
[1270] Step 1:
[1271] The server receives the video content and performs video decoding to convert each frame into an analyzable format.
[1272] Step 2:
[1273] The server runs video analysis algorithms to identify important scenes and objects in the video, such as building walls, billboards, or the sides of cars, and finds suitable spots for inserting ads.
[1274] Step 3:
[1275] Based on the analysis results, the server lists the identified points and plans to insert advertisements appropriate for those points.
[1276] Step 4:
[1277] The server selects appropriate advertising material from the advertising database based on the video content and target audience, and stores the selected advertising material in temporary memory.
[1278] Step 5:
[1279] The server uses generative AI to composite the selected ad material into specific points in the video, which involves analyzing the color, lighting, and resolution of the video and adjusting them to make the ad appear natural.
[1280] Step 6:
[1281] The server applies the compositing process and generates a new video file with the ads embedded naturally and unobtrusively into the original content.
[1282] Step 7:
[1283] The server uploads the newly generated video file to the distribution server, making it available for viewing by users.
[1284] Terminal handling
[1285] Step 1:
[1286] The user terminal accesses the distribution platform and selects the video that the user wants to watch.
[1287] Step 2:
[1288] The terminal transmits a request for the selected video and receives a new video file with an embedded advertisement from the distribution server.
[1289] Step 3:
[1290] The device decodes and buffers the received video file for playback.
[1291] Step 4:
[1292] The device plays the video and processes it so that the advertisements in the video and the original content are displayed seamlessly.
[1293] User Action
[1294] Step 1:
[1295] Users select the video they want to watch on the video streaming platform and press the play button.
[1296] Step 2:
[1297] Users watch videos played on their device and are presented with ads that are naturally embedded within the video but do not interrupt the normal viewing experience.
[1298] Step 3:
[1299] Users can enjoy videos continuously without interruption due to ads. Because ads are displayed naturally, users can immerse themselves in the content without stress.
[1300] Example 1
[1301] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1302] The challenge is to seamlessly integrate advertisements into video content without causing a sense of incongruity to users. Conventional methods often result in videos with advertisements inserted that look unnatural, detracting from the user's viewing experience.
[1303] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1304] In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for selecting appropriate advertising materials from an advertising database based on the identified points, means for analyzing and adjusting the color tone, lighting, and resolution of the video using a generative AI model to naturally combine the advertising materials into the video, means for generating a new video file including the generated advertisements and uploading it to a distribution server, means for a user to select a video they wish to watch on a video distribution platform, and means for a user terminal to send a request to the distribution server, receive the generated video file, and seamlessly play the video so that there is no sense of incongruity between the advertisements and the original video content. This allows users to enjoy a seamless and natural viewing experience while comfortably watching the advertisements.
[1305] "Video content" is a series of digital data including video and audio, and is a medium for users to view.
[1306] A "scene" refers to a specific frame or group of frames in video content, which represents a specific situation or background.
[1307] "Advertisement" means media content inserted into a video for the purpose of promoting a product or service.
[1308] A "point" refers to a specific position or timing within a scene that is suitable for inserting an advertisement.
[1309] "Means" refers to a method or mechanism for achieving a specific function or operation.
[1310] The "advertising database" is a database in which advertising materials are stored, and includes various advertising contents.
[1311] "Advertising materials" refers to specific media data used as advertising, including images, videos, text, etc.
[1312] "Analysis" is the process of examining data in detail and extracting specific information.
[1313] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to generate new data.
[1314] "Color tone" refers to the adjustment and balance of colors in videos and images.
[1315] "Lighting" refers to the intensity and direction of light in a scene, and is a factor in creating visual atmosphere.
[1316] "Resolution" is an indicator of how clearly the details of videos and images can be displayed.
[1317] A "distribution server" is a server for distributing videos and data to user terminals.
[1318] A "user terminal" is a device used by a user, including a PC, smartphone, tablet, etc.
[1319] A "request" is a request message sent from a user terminal to a server to obtain specific data.
[1320] "Prompt" refers to input text used to prompt a generative AI model to generate a particular result.
[1321] "Seamless" refers to a state without joints, indicating a natural, flowing continuity.
[1322] This invention provides a system for inserting advertisements into video content in a natural way, using a method for analyzing scenes and objects in the video, selecting the most suitable advertisements, and synthesizing them in a natural way. The following describes how to specifically implement this invention.
[1323] The server receives video content from users or content providers via HTTP requests or FTP. The received video content is broken down into frames using video analysis algorithms such as the OpenCV library or TensorFlow, and objects and scenes within each frame are analyzed. For example, distinctive objects such as building walls, signs, and the sides of cars are identified.
[1324] The server then selects the appropriate ad material from its advertising database based on the analysis results. This process uses recommendation algorithms such as Google AdSense API to select the most suitable ad based on the video content and target audience. For example, it may determine that a fashion brand ad is best suited for an urban scene.
[1325] Next, the server uses a generative AI model (such as DALL-E) to composite the selected advertising material into the video. Specifically, the server analyzes the color, lighting, and resolution of the video and adjusts it so that the advertising material is naturally integrated into the frame. At this time, the server uses prompts to the generative AI model to generate the desired results. For example, the server might use the prompt, "Generate a fashion brand advertisement on a billboard in an urban scene and display it naturally."
[1326] The completed video file is generated by the server, encoded, and uploaded to the distribution server via an HTTP POST request, allowing users to seamlessly watch the video with embedded ads through the video distribution platform.
[1327] When a user selects a video they want to watch on a video streaming platform and presses the play button, the device sends an HTTP GET request to the streaming server and receives the generated video file. The device decodes the video file and performs appropriate buffering during playback, providing the user with a seamless viewing experience. The user will notice advertisements embedded naturally within the original video, but the viewing experience is not interrupted.
[1328] As a specific example, the following prompt sentence is used:
[1329] "Generating and displaying fashion brand advertisements naturally on billboards in urban scenes"
[1330] "Naturally incorporate car ads into the background of movie scenes"
[1331] Users can enjoy the video without feeling uncomfortable by watching videos with these advertisements inserted naturally. In this way, the invention improves the user's viewing experience and maximizes the effectiveness of advertising.
[1332] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1333] Step 1:
[1334] Video reception and analysis
[1335] The server receives video files from users or content providers via HTTP requests or FTP. It uses the OpenCV library to break down the received video files into frames and analyze the objects and scenes contained in each frame. For example, it identifies color, lighting, and specific objects (buildings, signs, cars, etc.). The input video file is in digital format, and the analysis results for each frame are obtained as output.
[1336] Step 2:
[1337] Selection of advertising materials
[1338] The server selects appropriate advertising materials from an advertising database based on the analysis results. Recommendation algorithms such as the Google AdSense API are used to select the most suitable advertisements for the video content, scene, and target audience. For example, it may determine that a fashion brand advertisement is most suitable for an urban scene. The analysis results and scene information are used as input, and the selected advertising materials are generated as output.
[1339] Step 3:
[1340] Content adjustment and synthesis with generative AI models
[1341] The server uses a generative AI model to naturally incorporate the selected advertising material into the video. A prompt is created and input to the generative AI model (e.g., DALL-E). The prompt is used as an example: "Generate a fashion brand advertisement on a billboard in an urban scene and display it naturally." The generative AI model generates the required advertising images and uses OpenCV to composite the advertising material into the video frame. The prompt and advertising material are used as input, and a frame with the advertisement composited is obtained as output.
[1342] Step 4:
[1343] Creating and uploading a new video file
[1344] The server generates a new video file from all composited frames, encodes and compresses it, and saves it in the specified format (e.g., MP4 or MKV). It then uploads the resulting video file to the distribution server via an HTTP POST request. It uses each composited frame as input and generates a new video file as output, which is then uploaded to the distribution server.
[1345] Step 5:
[1346] Requesting and Receiving Videos
[1347] A user selects a video they want to watch on a video distribution platform and presses the play button. This causes the user device to send an HTTP GET request to the distribution server. A new video file with an embedded advertisement is sent from the distribution server to the user device as an HTTP response. The user's request information is used as input, and the new video file is provided to the user device as output.
[1348] Step 6:
[1349] Play video
[1350] The user device decodes the received video file and plays it through the appropriate player. The device buffers during playback to ensure a seamless viewing experience for the user. Advertisements are displayed seamlessly between the original video content. The new video file is used as input, and a seamlessly played video is provided as output.
[1351] The above steps realize a system that inserts advertisements naturally into video content, providing users with a natural viewing experience.
[1352] (Application example 1)
[1353] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1354] Conventional methods for inserting advertisements into video content often interrupt the viewer's experience, resulting in limited advertising effectiveness. Furthermore, it is difficult to integrate advertisements in a visually natural way, often causing viewers to feel uncomfortable. The present invention aims to provide a method for seamlessly and naturally inserting advertisements into video content without interrupting the viewer's experience.
[1355] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1356] In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for selecting appropriate advertising materials from an advertising database, means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine the advertising materials with the video, means for seamlessly receiving and playing the generated video file at the user terminal, and means for analyzing the effectiveness of the advertisements and generating a report of the results. This enables visually natural and seamless advertisement insertion, making it possible to provide video content without interrupting the viewer's experience while increasing the effectiveness of the advertisements.
[1357] "Video Content" means a digital video file containing visual and audio information.
[1358] A "scene" refers to a series of frames in video content, and signifies a sequence of images at a specific location or time axis.
[1359] An "advertising database" is a digital repository for storing advertising materials, and is a system that stores a wide variety of advertising data.
[1360] "Advertising Materials" refers to digital content such as images, videos, and text used as advertising.
[1361] "Generative AI" refers to algorithms or models that use artificial intelligence techniques to generate new data or content.
[1362] A "distribution server" is a centralized computer system that provides digital content to user terminals over a network.
[1363] "User terminal" refers to a device, such as a smartphone, tablet, or PC, that allows a user to access digital content via the Internet.
[1364] "Seamless" refers to a state in which the continuity of operations and actions is uninterrupted, and is carried out smoothly without any sense of discomfort.
[1365] "Analysis" refers to the methods and processes used to analyze data and information and understand its structure and meaning.
[1366] "Synthesis" refers to the process of combining different digital content to create new data or images.
[1367] "Effectiveness analysis" refers to statistical or quantitative research or analysis conducted to evaluate the effectiveness or impact of advertising.
[1368] "Report" means a document that summarizes the results of research and analysis and organizes them visually and written.
[1369] The system embodying the present invention performs a series of processes for inserting advertisements into video content in a natural way. This system is mainly composed of a server and a user terminal.
[1370] Server Processing
[1371] Video reception and analysis
[1372] The server first receives the video content sent from the user device, then analyzes each frame of the video using a video analysis algorithm (e.g., OpenCV) to identify scenes and objects suitable for inserting advertisements, thereby determining the advertisement insertion points.
[1373] Selection of advertising materials
[1374] The server then selects appropriate advertising material from its advertising database based on the analysis results, using algorithms to select ads that fit the video's content and target audience. For example, fashion brand ads are often selected for city scenes.
[1375] Content Alignment and Composition
[1376] The server uses a generative AI model (e.g., moviepy) to seamlessly blend the selected ad material into the video. It analyzes the color, lighting, and resolution of the video and adjusts it so that the ad material appears as a natural part of the video. This step provides a seamless visual experience for the viewer.
[1377] Creating and uploading a new video file
[1378] A new video file containing the advertisement is generated and uploaded to the distribution server, which then provides the new video file in response to a request from the user terminal.
[1379] User terminal processing
[1380] Requesting and Receiving Videos
[1381] When the user selects a video that he or she wants to watch, the user terminal sends a request to the distribution server and receives a new video file that includes advertisements.
[1382] Play video
[1383] The device then plays the received video file, buffering and decoding the video as it plays, providing a seamless viewing experience with no sense of incongruity between the ad and the original video content.
[1384] User Experience
[1385] Users select a video they want to watch on a video streaming platform and start playing it. While the video is playing, they may notice advertisements inserted naturally, but they do not interrupt the viewing experience. For example, they may see a billboard advertisement displayed naturally within a movie scene. This allows users to enjoy the video without feeling stressed by advertisements.
[1386] Specific examples
[1387] An example of a prompt sentence is as follows:
[1388] Video Analysis Prompt
[1389] Identify suitable scenes in the input video for ad insertion and composite the latest ads appropriately into those locations.
[1390] Advertisements used:
[1391] The latest advertising video.
[1392] Expected Results:
[1393] A new video file is generated in which the advertisements are seamlessly integrated into the original footage.
[1394] As described above, the system of the present invention provides a series of processes for inserting advertisements into video content in a natural way, thereby providing users with a seamless viewing experience.
[1395] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1396] Step 1:
[1397] Video reception and analysis
[1398] The server receives video content sent from the user's device. It analyzes the received video frame by frame using a video analysis algorithm (e.g., OpenCV) to identify scenes and objects suitable for inserting advertisements. The input to this step is the received video content, and the output is information identifying the ad insertion point. Specifically, it determines where an advertisement can be inserted naturally based on the objects and background information recognized through frame-by-frame analysis.
[1399] Step 2:
[1400] Selection of advertising materials
[1401] The server selects appropriate advertising material from the advertising database based on the ad insertion point identified in step 1. This selection uses an algorithm to select advertisements that match the video content information and target audience information. The input to this step is the identification information of the ad insertion point, and the output is the selected advertising material. Specifically, the server searches the advertising database for the most suitable advertisement and extracts the corresponding advertising material.
[1402] Step 3:
[1403] Content Alignment and Composition
[1404] The server uses a generative AI model (e.g., moviepy) to composite the selected ad material into the video in a natural way. At this time, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the ad material blends in with the original footage. The input for this step is the video content and the selected ad material, and the output is a new video file with the ad composited into it. Specifically, the color tone and transparency of the ad material are adjusted based on the color tone and light direction of each frame of the video, resulting in a seamless composite.
[1405] Step 4:
[1406] Creating and uploading a new video file
[1407] The server generates a new video file containing the composited advertisement and uploads the video file to the distribution server. The input of this step is the new video file with the composited advertisement, and the output is the video file uploaded to the distribution server. Specifically, the generated video file is encoded into a specified format and transferred to the distribution server.
[1408] Step 5:
[1409] Requesting and Receiving Videos
[1410] When the user selects a video they want to watch, the user terminal sends a request to the distribution server and receives a new video file including advertisements. The input to this step is the user request, and the output is the video file received from the distribution server. Specifically, an HTTP request is sent based on the user's instructions, and the new video file is received as a response.
[1411] Step 6:
[1412] Play video
[1413] The user device plays the received video file. During video playback, buffering and decoding are performed, maintaining a seamless connection between the advertisement and the original video content. The input for this step is the received video file, and the output is a seamlessly played video. Specifically, the video player function is used to buffer and decode the video using a codec, ensuring smooth playback.
[1414] Step 7:
[1415] Providing a viewing experience
[1416] A user selects a video they want to watch on a video streaming platform and starts playback. While the video is playing, they may notice advertisements inserted naturally, but the viewing experience is not interrupted. Specifically, they can enjoy the video continuously while viewing billboard advertisements that are naturally displayed within movie scenes. The input for this step is a video file received from the streaming server, and the output is a video that plays continuously without interruption.
[1417] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1418] Explaining program processing in natural language
[1419] server
[1420] 1. Receiving and analyzing the video
[1421] The server receives the video content, performs video decoding to convert each frame into an analyzable format, and runs video analysis algorithms to identify important scenes and objects in the video (e.g., building walls, billboards, sides of cars, etc.) to find suitable points for inserting advertisements.
[1422] 2. Acquiring Emotion Data
[1423] The server uses an emotion engine to obtain the user's emotions in real time while watching a video, which includes technology to determine the user's emotional state by analyzing their facial expressions, tone of voice, and eye movements.
[1424] 3. Selection of advertising materials
[1425] The server selects appropriate advertising materials from the advertising database based on the content of the video, the target audience, and the acquired user emotion data. For example, if the user is enjoying watching the video, entertainment-related advertisements may be deemed appropriate.
[1426] 4. Content Coordination and Synthesis
[1427] The server uses generative AI to composite the selected advertising material into specific points in the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it to make the advertisement appear natural. For example, a fashion brand advertisement can be composited onto a building sign, making it appear as if it were part of the original footage.
[1428] 5. Create and upload a new video file
[1429] A new video file containing the generated advertisement is generated and uploaded to a distribution server, allowing users to seamlessly watch the video with the embedded advertisement through a distribution platform.
[1430] Terminal
[1431] 1. Requesting and Receiving Videos
[1432] When a user selects a video they want to watch, the device sends a request to the distribution server, which then receives a new video file with an embedded advertisement.
[1433] 2. Play the video
[1434] The device decodes and buffers the received video file for playback.
[1435] 3. Seamless Ad Display
[1436] The device plays the video and performs playback processing so that the advertisements in the video and the original content are displayed seamlessly, providing the user with a seamless viewing experience.
[1437] User
[1438] 1. Select a video
[1439] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[1440] 2. Watching videos
[1441] A user watches a video playing on their device and is presented with ads that are naturally embedded within the video, without interrupting the normal viewing experience. For example, a user might watch a movie with a billboard ad or background ad that is naturally embedded within the video.
[1442] 3. Emotional Feedback
[1443] Real-time user emotional data is collected and used to select and adjust advertising materials, which may even dynamically change based on the user's emotional state while they are watching the video.
[1444] 4. Continue watching
[1445] Users can enjoy videos continuously without feeling interrupted by advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[1446] Specific examples
[1447] Server Processing
[1448] 1. The server receives the movie footage and performs video analysis, for example, identifying city scenes with billboards suitable for inserting advertisements.
[1449] 2. The server uses an emotion engine to obtain the user's emotion in real time while watching. For example, if the user is enjoying the video, the emotion is judged to be positive.
[1450] 3. Based on the positive emotional state, the server determines that an entertainment advertisement (e.g., a trailer for a new movie) would be a good fit for this scene.
[1451] 4. The generative AI composites a movie trailer onto a specific billboard, integrating it into the scene.
[1452] 5. Upload the completed composite video file to the distribution server.
[1453] Terminal handling
[1454] 1. The user device accesses the distribution platform, selects the movie, and starts playback.
[1455] 2. City scenes from the film play, with entertainment ads on appropriate billboards, making them feel like part of the regular billboard scene.
[1456] 3. The user device provides a seamless viewing experience with no sense of incongruity between the advertisements and the movie.
[1457] User Experience
[1458] 1. The user begins watching a movie and notices a billboard advertisement that is displayed naturally within the city scene.
[1459] 2. The movie is not interrupted, allowing users to remain immersed in the story while watching ads, and the ads are selected based on the user's emotional state, further enhancing the viewing experience.
[1460] 3. This allows users to enjoy movies without being stressed by advertisements.
[1461] The processing flow will be explained below.
[1462] Server Processing
[1463] Step 1:
[1464] The server receives the video content and performs video decoding to convert each frame into an analyzable format.
[1465] Step 2:
[1466] The server runs video analysis algorithms to identify important scenes and objects in the video, such as building walls, billboards, or the sides of cars, and find suitable spots for inserting ads.
[1467] Step 3:
[1468] The server uses an emotion engine to obtain real-time emotional data from users watching videos. This emotional data is determined by analyzing the user's facial expressions, tone of voice, eye movements, etc.
[1469] Step 4:
[1470] The server selects appropriate advertising materials from the advertising database based on the content of the video, the target audience, and the acquired user emotion data. For example, if the user emotion indicates enjoyment, entertainment-related advertisements will be selected.
[1471] Step 5:
[1472] The server uses generative AI to composite the selected ad material into specific points in the video. Specifically, it analyzes the color tone, lighting, and resolution of the video and adjusts it so that the ad looks natural with the original footage. For example, a movie trailer ad can be composited onto a building billboard.
[1473] Step 6:
[1474] The server applies the compositing process and generates a new video file with the ads embedded naturally and unobtrusively into the original content.
[1475] Step 7:
[1476] The server uploads the generated new video file to the distribution server, making it available for viewing by users.
[1477] Terminal handling
[1478] Step 1:
[1479] The user terminal accesses the distribution platform and selects the video that the user wants to watch.
[1480] Step 2:
[1481] The terminal transmits a request for the selected video and receives a new video file with an embedded advertisement from the distribution server.
[1482] Step 3:
[1483] The device decodes and buffers the received video file for playback.
[1484] Step 4:
[1485] The device plays the video and processes it so that the advertisements in the video and the original content are displayed seamlessly.
[1486] User Action
[1487] Step 1:
[1488] The user selects the video they want to watch on the video distribution platform and presses the play button, which causes the device to send a request to the server and receive the generated video file.
[1489] Step 2:
[1490] A user watches a video playing on their device, and an ad appears, naturally embedded within the video, without interrupting the normal viewing experience. For example, a user might watch a billboard ad that is naturally embedded within a movie.
[1491] Step 3:
[1492] Real-time user emotional data is collected and used to select and adjust advertising materials, which may even dynamically change based on the user's emotional state while they are watching the video.
[1493] Step 4:
[1494] Users can enjoy videos continuously without feeling interrupted by advertisements. Because advertisements are displayed naturally, users can immerse themselves in the content without stress.
[1495] Example 2
[1496] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1497] When inserting advertisements into video content, there is a need for a method to display advertisements naturally without disrupting the user's viewing experience. There is also a need for a method to appropriately select advertisements based on the viewer's emotional state and improve the viewing experience. However, conventional systems have difficulty in utilizing user emotional data to select advertisements and integrate them into videos in a natural way.
[1498] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving video content and analyzing scenes in the video to identify suitable points for inserting advertisements, means for acquiring user emotion data and selecting advertising material based on the data, means for selecting appropriate advertising material from an advertisement database based on the identified points and the user emotion data, means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine the advertising material into the video, and means for generating a new video file including the combined advertisement and uploading it to a distribution device. This makes it possible to naturally combine optimal advertisements based on the user's emotions into the video, improving the user's viewing experience.
[1499] "Video Content" means a media file that combines video and audio and is provided for the enjoyment of viewers.
[1500] "Scene analysis" is the process of decoding each frame in a video and identifying important objects and points of interest.
[1501] "Good ad insertion points" are specific scenes or locations in a video where ads can be displayed naturally and effectively.
[1502] "User emotion data" is information indicating the emotional state of the viewer obtained from facial expressions, gaze, voice, and the like.
[1503] "Advertising Materials" means advertising footage, images, or other media content for display.
[1504] An "advertising database" is a searchable database that stores various advertising materials.
[1505] "Analysis and adjustment of color, lighting, and resolution" refers to the process of analyzing the quality and visual elements of a video and making adjustments to make the ads appear natural.
[1506] "Generative AI" is a type of artificial intelligence that is a technology that generates advertising materials and synthesizes them into videos.
[1507] "Distribution device" refers to the server and network infrastructure used to deliver the generated video files to users.
[1508] "Seamless playback" refers to a state in which the video and advertisements are played back consecutively, allowing the user to view them as a single entity without feeling any discomfort.
[1509] The "new video file" is a file in which the advertisement has been synthesized and the original video and advertisement are integrated.
[1510] MODE FOR CARRYING OUT THE INVENTION
[1511] The system of this invention receives and analyzes video content, and automatically inserts advertisements based on user emotion data to improve the viewing experience. The system is mainly composed of three elements: a server, a terminal, and a user.
[1512] server
[1513] The server receives the video content and securely stores it using Amazon S3. The stored video data is decoded into individual frames using FFmpeg. OpenCV is then used to analyze the video and identify important scenes and objects. For example, it can detect billboards suitable for inserting advertisements in urban scenes. The server then collects user emotion data using Microsoft Azure's Emotion API. This technology analyzes the user's facial expressions, tone of voice, eye movements, and other factors to determine their emotional state in real time.
[1514] The server then uses Google Cloud's BigQuery to select appropriate ad materials from the ad database. The selected ad materials are then composited into specific points in the video using OpenAI's generative AI. Blender is used to analyze the color tone, lighting, and resolution of the video and adjust them to make the ads look natural. Finally, a new video file containing the composited ads is generated and uploaded to the distribution device via Amazon CloudFront. Users can then watch the video with the embedded ads.
[1515] Terminal
[1516] When a user selects a video they want to watch, the device sends a request to the distribution server. In response to the request, a new video file with embedded ads is received from Amazon CloudFront. The received video file is decoded and buffered using VLCKit, and then played back seamlessly. This allows the user to enjoy a natural viewing experience without any sense of incongruity between the video and the ads.
[1517] User
[1518] Users select a video they want to watch on the streaming platform and press the play button. This causes the device to send a request to the server and receive a video file containing the advertisement. As users watch the video on their device, they notice the advertisements, which are naturally embedded in billboards and backgrounds, without interrupting the normal viewing experience. In addition, users' emotional data is collected in real time and used to select and adjust advertising materials. Advertisements may also be dynamically changed based on the user's emotional state, further improving the viewing experience.
[1519] Specific examples
[1520] Server Processing
[1521] 1. Receive movie data and store it in Amazon S3.
[1522] 2. Decode the movie using FFmpeg and identify billboards in urban scenes with OpenCV.
[1523] 3. Use Microsoft Azure's Emotion API to analyze the user's facial expressions in real time and obtain positive emotions.
[1524] 4. Select entertainment advertising materials based on positive emotional states using Google Cloud BigQuery.
[1525] 5. Synthesize ads naturally using OpenAI's generative AI and Blender.
[1526] 6. Upload the composite video file to Amazon CloudFront and display it.
[1527] Terminal handling
[1528] 1. The user selects a movie and sends a request to the distribution server.
[1529] 2. Receive new video files from Amazon CloudFront.
[1530] 3. Use VLCKit to decode and buffer the video.
[1531] 4. Seamlessly play videos with natural ad insertions.
[1532] User Experience
[1533] 1. The user selects a movie and presses the play button.
[1534] 2. While watching the movie, you notice billboard advertisements appearing in city scenes.
[1535] 3. Watch movies without being interrupted by ads.
[1536] 4. While watching, ads may dynamically change based on emotions.
[1537] Prompt Sentence Examples
[1538] 1. "Can you give me an example of a diorama scene that displays entertainment advertisements for a movie?"
[1539] 2. "Please explain techniques for inserting ads naturally into videos, especially those that leverage sentiment data."
[1540] keyword
[1541] Generative AI model, prompt sentence
[1542] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1543] Step 1: Receive and save the video
[1544] The server receives video content from users. For example, a user uploads a movie. This received video data is stored in Amazon S3. This ensures that the data is securely stored and can be used for further processing. The input video data is stored in Amazon S3 as a stored output.
[1545] Step 2: Video decoding and analysis
[1546] The server decodes the video data stored in Amazon S3 into individual frames using FFmpeg. Each decoded frame undergoes video analysis using OpenCV. Specifically, it identifies objects and scenes suitable for inserting advertisements, such as billboards in urban scenes. The input of this step is the video data retrieved from Amazon S3, and the output is the analyzed frame data.
[1547] Step 3: Obtaining user emotion data
[1548] The server uses Microsoft Azure's Emotion API to obtain the user's emotional data in real time, which includes technology that analyzes the user's facial expressions, tone of voice, eye movements, etc. to determine their emotional state. The input for this step is the user's video feed and audio data, and the output is the analyzed emotional data.
[1549] Step 4: Selecting advertising materials
[1550] The server uses Google Cloud's BigQuery to select appropriate advertising materials based on the video content, target audience, and acquired user emotion data. For example, entertainment-related ads are selected for users who show positive emotions. The input of this step is the analyzed frame data and emotion data, and the output is the selected advertising materials.
[1551] Step 5: Adjust and combine content
[1552] The server uses OpenAI's generative AI to composite the selected advertising material into specific points in the video. Blender is used to analyze the color, lighting, and resolution of the video and adjust the video to make the advertisement appear natural. For example, a fashion brand advertisement can be composited naturally onto a billboard in an urban scene. The input to this step is the selected advertising material and the analyzed frame data, and the output is a new frame data with the advertisement composited into it.
[1553] Step 6: Generate and upload a new video file
[1554] The server generates a new video file containing the advertisements and uploads it to the distribution device via Amazon CloudFront. The uploaded video file is then accessible on the user's device. The input of this step is the new frame data with the advertisements mixed in, and the output is the final video file.
[1555] Step 7: Requesting and receiving video
[1556] When the user selects a video to watch, the device sends a request to the distribution server. It receives a new video file delivered by Amazon CloudFront. The input of this step is the video viewing request, and the output is the received video file.
[1557] Step 8: Play the video
[1558] The device uses VLCKit to decode and buffer the received video file, allowing the user to seamlessly watch the video. The input of this step is the received video file, and the output is a decoded, buffered, and playable video.
[1559] Step 9: Obtaining and using user emotional feedback
[1560] While the user is watching the video, the server continues to capture the user's emotional data and use it to select and adjust advertising materials, which may be dynamically changed according to the user's real-time emotional state. The input of this step is real-time user emotional data, and the output is adapted advertising materials.
[1561] (Application example 2)
[1562] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1563] Conventional video ad insertion systems use simple ad synthesis methods, which often detract from the user's viewing experience. In particular, ads are displayed uniformly regardless of the video scene or the user's emotional state, which makes the ads visually unnatural and makes the user feel uncomfortable. Furthermore, there is a need for more personalized ad delivery by utilizing user emotional data to select ads.
[1564] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving video content and analyzing scenes in the video to identify points suitable for inserting advertisements; means for selecting appropriate advertising material from an advertising database based on the identified points; means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally combine advertising material into the video; means for analyzing user emotional data in real time and selecting an optimal advertisement based on the analysis results; means for combining the selected advertising material into specific points in the video and using a generative AI model to adjust the combination so that it is seamless; and means for generating a new video file including the combined advertisement and uploading it to a distribution server. This enables personalized advertising insertion based on the user's emotional state.
[1565] "Video content" refers to video data that provides information to users visually and audibly.
[1566] A "scene" is a set of consecutive video frames in a specific time range within a video.
[1567] "Advertisement" means visual or audio content intended to promote a product or service.
[1568] A "point" refers to a suitable location in a particular scene or frame of a video for inserting an advertisement.
[1569] "Means" refer to the methods or techniques used to achieve a particular goal.
[1570] A "database" is a collection of information that organizes large amounts of information and makes it possible to search it efficiently.
[1571] "Hue" is an attribute related to the color tone and color combination of an image.
[1572] "Lighting" refers to the intensity and directionality of the light used in the video.
[1573] "Resolution" is a measure of the ability to express detail in an image, and is generally expressed in terms of the number of pixels.
[1574] "Compositing" refers to the process of combining different video materials into one video.
[1575] "Emotion data" is information about the user's emotional state obtained from facial expressions, tone of voice, and the like.
[1576] A "generative AI model" refers to an algorithm or system that uses artificial intelligence to generate new data or content.
[1577] "Uploading" is the act of transferring data from a local environment to a remote server.
[1578] "Terminal" refers to a computing device that is directly operated by a user, including smartphones and personal computers.
[1579] "Seamless" means that the operation is natural, without any gaps or interruptions.
[1580] "Real-time" refers to data acquisition and processing occurring immediately.
[1581] The system for implementing the invention is composed of three elements: a server, a user terminal, and a user. The specific operation process of this system is as follows:
[1582] Server Processing
[1583] The server receives the video content and analyzes it frame by frame. Using video analysis algorithms, it analyzes each scene in the video and identifies important elements (e.g., billboards or building walls) to find the optimal points for inserting ads.
[1584] The server then uses an emotion engine to analyze the user's emotional data in real time. It analyzes the user's facial image and tone of voice to understand their emotional state. Based on this information, the server selects the most suitable advertising material from its advertising database.
[1585] The system uses a generative AI model to seamlessly blend the selected ad material into the video, adjusting color, lighting, and resolution to seamlessly insert the ad material at specific points in the video. Finally, the server generates the composite video file and uploads it to the distribution server.
[1586] The specific hardware used is high-performance server equipment, and the software includes video analysis algorithms, emotion engines, advertising databases, and generative AI models.
[1587] User terminal processing
[1588] When a user selects a video they want to watch, the newly synthesized video file is received from the distribution server. The user's device decodes the video file and plays it back seamlessly, providing a seamless viewing experience for the user.
[1589] The specific hardware used is the user's device, such as a smartphone or PC, and the software includes the video playback application and decoding algorithm.
[1590] User Experience
[1591] When users watch videos through a video distribution platform, ads are displayed naturally embedded within the video, so the normal viewing experience is not interrupted. In addition, by acquiring user emotional data in real time, it is possible to dynamically display appropriate ads.
[1592] Examples of prompt statements
[1593] 1. "Analyze the user's emotional state using facial images and audio data."
[1594] 2. "Please identify scenes in your video file where billboards or other advertising can be inserted."
[1595] 3. "Choose the best ad based on the user's current emotional state."
[1596] 4. "Please integrate the selected ads into the video scenes in a natural way."
[1597] This format allows users to enjoy a highly satisfying video viewing experience without feeling stressed by advertisements.
[1598] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1599] Step 1:
[1600] The server receives video content and uses video analysis algorithms to convert each frame into an analyzable format. The input is a video file, and the output is frame-by-frame video data. Specifically, the server breaks down the video into frames and identifies important scenes and objects in each frame.
[1601] Step 2:
[1602] The server acquires the user's emotional data in real time. The input is the user's facial image and voice data, and the output is the user's emotional state (e.g., joy, excitement, sadness). Specifically, the server uses an emotion engine to analyze the user's facial expressions and tone of voice to determine the user's emotional state.
[1603] Step 3:
[1604] The server selects appropriate advertising materials from an advertising database based on video analysis data and user emotional data. The input is the analyzed scene data and the user's emotional state, and the output is the selected advertising materials. Specifically, the server uses an advertising identification algorithm to select the advertisement that best suits the user's emotional state.
[1605] Step 4:
[1606] The server utilizes a generative AI model to seamlessly composite selected ad material into specific points in the video. The input is the selected ad material and analyzed scene data, and the output is frame data with the ad composited. Specifically, the server adjusts the color tone, lighting, and resolution of the video to composite the ad material naturally.
[1607] Step 5:
[1608] The server generates a new video file containing the composited advertisements and uploads it to the distribution server. The input is the frame data with the composited advertisements, and the output is a new video file. Specifically, the server reconstructs the frame data into a single video file and transfers it to the distribution server.
[1609] Step 6:
[1610] The user terminal receives a new video file from the distribution server. The input is the video file sent from the distribution server, and the output is the saved video file. Specifically, the user terminal downloads the video file.
[1611] Step 7:
[1612] The user terminal decodes and buffers the received video file and plays it seamlessly. The input is the stored video file, and the output is the played video. Specifically, the user terminal uses a video decoder to buffer the video frame by frame and play it smoothly.
[1613] Step 8:
[1614] The user watches a video played on a device. The input is the video being played, and the output is the user's visual and auditory experience. The specific experience is that the user can enjoy the video containing naturally-mixed advertisements without any sense of incongruity.
[1615] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1616] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1617] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1618] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1619] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1620] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1621] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1622] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1623] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1624] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1625] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1626] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1627] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1628] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1629] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1630] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1631] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1632] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1633] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1634] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1635] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1636] The following is further disclosed regarding the above embodiment.
[1637] (Claim 1)
[1638] means for receiving video content and analyzing scenes within the video to identify suitable points for inserting advertisements;
[1639] means for selecting appropriate advertising material from an advertising database based on the identified points;
[1640] means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally synthesize advertising materials into the video;
[1641] means for generating a new video file including the combined advertisement and uploading the file to a distribution server;
[1642] A system including:
[1643] (Claim 2)
[1644] The system according to claim 1, wherein the advertising material is adjusted and synthesized using a generation AI when being synthesized into a video.
[1645] (Claim 3)
[1646] 2. The system according to claim 1, further comprising means for receiving the generated video file in a user terminal and playing it back seamlessly.
[1647] "Example 1"
[1648] (Claim 1)
[1649] means for receiving video content and analyzing scenes within the video to identify suitable points for inserting advertisements;
[1650] means for selecting appropriate advertising material from an advertising database based on the identified points;
[1651] means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally synthesize advertising materials into the video;
[1652] means for generating a new video file including the generated advertisement and uploading the file to a distribution server;
[1653] A means for a user to select a video they wish to watch on the video distribution platform;
[1654] A means for a user terminal to send a request to the distribution server, receive the generated video file, and seamlessly play it;
[1655] A system including:
[1656] (Claim 2)
[1657] The system of claim 1, wherein the advertising material is adjusted and synthesized using a generative AI model when being synthesized into a video.
[1658] (Claim 3)
[1659] 2. The system according to claim 1, further comprising means for receiving the generated video file in a user terminal and playing the video file so that the advertisement and the original video content do not look out of place.
[1660] "Application Example 1"
[1661] (Claim 1)
[1662] means for receiving video content and analyzing scenes within the video to identify suitable points for inserting advertisements;
[1663] means for selecting appropriate advertising material from an advertising database based on the identified points;
[1664] means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally synthesize advertising materials into the video;
[1665] means for generating a new video file including the combined advertisement and uploading the file to a distribution server;
[1666] means for seamlessly receiving and playing the generated video file in a user terminal;
[1667] A system including:
[1668] (Claim 2)
[1669] The system according to claim 1, wherein the advertising material is adjusted and synthesized using a generation AI when being synthesized into a video.
[1670] (Claim 3)
[1671] 2. The system according to claim 1, further comprising means for analyzing advertising effectiveness and generating a report based on the results.
[1672] "Example 2: Combining Emotion Engines"
[1673] (Claim 1)
[1674] means for receiving video content and analyzing scenes within the video to identify suitable points for inserting advertisements;
[1675] means for acquiring user emotion data and selecting advertising materials based on the data;
[1676] means for selecting appropriate advertising materials from an advertising database based on the identified points and user emotion data;
[1677] means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally synthesize advertising materials into the video;
[1678] means for generating a new video file including the synthesized advertisement and uploading the file to a distribution device;
[1679] A system including:
[1680] (Claim 2)
[1681] The system according to claim 1, wherein the advertising material is adjusted and synthesized using a generation AI when being synthesized into a video.
[1682] (Claim 3)
[1683] 2. The system according to claim 1, further comprising means for receiving the generated video file in a user terminal and playing it back seamlessly.
[1684] "Application example 2 when combining emotion engines"
[1685] (Claim 1)
[1686] means for receiving video content and analyzing scenes within the video to identify suitable points for inserting advertisements;
[1687] means for selecting appropriate advertising material from an advertising database based on the identified points;
[1688] means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally synthesize advertising materials into the video;
[1689] a means for analyzing user emotion data in real time and selecting an optimal advertisement based on the analysis results;
[1690] a means for using a generative AI model to synthesize selected advertising material into specific points in the video and adjust the synthesis so that it is seamless;
[1691] means for generating a new video file including the combined advertisement and uploading the file to a distribution server;
[1692] A system including:
[1693] (Claim 2)
[1694] The system according to claim 1, wherein the advertising material is adjusted and synthesized using a generation AI when being synthesized into a video.
[1695] (Claim 3)
[1696] 2. The system according to claim 1, further comprising means for receiving the generated video file in a user terminal and playing it back seamlessly. [Explanation of symbols]
[1697] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving video content and analyzing scenes within the video to identify suitable points for inserting advertisements; means for selecting appropriate advertising material from an advertising database based on the identified points; means for analyzing and adjusting the color tone, lighting, and resolution of the video to naturally synthesize advertising materials into the video; means for generating a new video file including the combined advertisement and uploading the file to a distribution server; A system including:
2. The system according to claim 1, wherein the advertising material is adjusted and synthesized using a generation AI when being synthesized into a video.
3. 2. The system according to claim 1, further comprising means for receiving the generated video file and playing it back seamlessly in a user terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A