system
A system that collects and analyzes talent data to generate AI talents, integrating them into advertising content and registering as NFTs on a blockchain addresses the challenge of maintaining talent value, ensuring sustained brand image and advertising effectiveness.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
The entertainment industry faces challenges in maintaining the market value and brand image of famous talents due to aging or death, and long-term advertising strategies incur high costs without effectively sustaining the talent's peak value.
A system that collects visual, audio, and thought patterns of a talent, generates an AI talent using a generative AI model, incorporates it into advertising content, and saves the data as an NFT on a blockchain to ensure the talent's peak value is maintained and utilized efficiently.
This system recreates the talent's peak value, enabling long-term use in advertising strategies and branding, ensuring the immortality of the talent's value by converting the generated AI talent data into a tradable digital asset.
Smart Images

Figure 2026060643000001_ABST
Abstract
Description
Technical Field
[0005]
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the entertainment industry, there is a risk that the popularity and market value of famous talents will decrease due to aging or death. In addition, when a company hires talents for long-term advertising strategies and branding, it incurs high costs. Furthermore, it is difficult to sustain the value of a talent at its peak, and as a result, the brand image of the company is also likely to fluctuate. There is a need for a method to solve such problems, maintain the value of a talent at its peak over a long period, and efficiently and effectively implement the advertising strategy of a company.
Means for Solving the Problems
[0005] This invention is a system that includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content such as commercials; and means for saving the generated AI talent data as an NFT and registering it on a blockchain. This system allows for the reproduction of the talent's peak value and its long-term use in advertising strategies and branding. Furthermore, by converting the generated AI talent data into an NFT, it can be traded on the market as a new digital asset, thereby ensuring the immortality of the talent's value.
[0006] The term "talent" refers to individuals who work in the entertainment industry, and includes actors, singers, models, comedians, and others.
[0007] "Visual information" refers to information related to videos and images of talent, including data such as facial features, expressions, and movements.
[0008] "Audio information" refers to information related to a talent's voice, including data such as voice tone, intonation, and speaking style.
[0009] "Thinking patterns" refer to information related to the thoughts and attitudes that a talent displays through their statements and actions, and include data extracted from interviews and statements.
[0010] A "generative AI model" refers to a data generation system that uses artificial intelligence to create AI talents, taking visual information, audio information, and thought patterns of talents as input data.
[0011] An "AI talent" is a synthetic talent created using a generative AI model, which reproduces the visual information, auditory information, and thought patterns of a real talent during their prime.
[0012] "Advertising content" refers to content, including videos, images, and text, produced by companies to promote their products or services.
[0013] An "ad template" refers to the basic design of advertising content, including the storyline and design.
[0014] "NFT" stands for Non-Fungible Token, and refers to a unique digital asset created using blockchain technology.
[0015] "Blockchain" refers to a system that uses distributed ledger technology to securely store data and record transaction information. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.
[0018] First, the language used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak period, with the aim of establishing advertising strategies and brand image. This system is implemented as follows.
[0038] 1. Data collection:
[0039] The terminal collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak period.
[0040] 2. Data Analysis:
[0041] The device extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm.
[0042] 3. Preparing the Generative AI Model:
[0043] The server prepares a generative AI model based on the analyzed data. The generative AI model is trained using deep learning to recreate the celebrity's appearance during their prime, and processes visual information, auditory information, and thought patterns as input data.
[0044] 4. AI Talent Generation:
[0045] The server uses a prepared generative AI model to generate realistic AI talents. These AI talents represent the real-life talents in their prime and possess high fidelity in terms of visual, auditory, and cognitive aspects.
[0046] 5. Ad generation:
[0047] The terminal retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into the advertising template. This generates a finished advertisement with visual consistency and audio adjustments. For example, the generated AI talent can be used to create new commercial content and promote products.
[0048] 6. NFT creation and blockchain registration:
[0049] The server stores the generated AI talent data as an NFT. This NFT is assigned a unique digital ID and registered on the blockchain. As a result, the generated AI talent data is recognized as a new digital asset and becomes tradable on the market.
[0050] Specific example
[0051] For example, consider a case where past data of a famous celebrity A is collected and analyzed, and then an AI celebrity is generated based on that data.
[0052] 1. The device collects past video files, audio files, and interview records of talent A.
[0053] 2. The device analyzes collected video files to extract visual data, and analyzes audio files to extract audio data. Furthermore, it analyzes interview recordings to extract thought data.
[0054] 3. The server trains an AI model based on this data to generate an AI talent that represents Talent A in their prime.
[0055] 4. The server incorporates the generated AI talent into an advertising template provided by the company to produce the final advertisement. This advertisement features an AI talent that represents Talent A in their prime.
[0056] 5. The server stores the generated AI talent data as an NFT and registers it on the blockchain. This NFT can be traded on the market as a digital asset.
[0057] In this way, the present invention uses AI to recreate a talent's peak period, enabling the maintenance of advertising strategies and brand image, and the creation of new business models.
[0058] The following describes the processing flow.
[0059] Program processing flow: Detailed step-by-step explanation
[0060] Step 1: Data Collection
[0061] The terminal collects past video files, audio files, and interview recordings of the talent from a database. This includes searching for and retrieving all related media files using a specified talent ID.
[0062] Step 2: Analysis of the month
[0063] The device processes the collected video files using a video analysis algorithm to extract visual data. This analysis captures the talent's facial features, movements, and expressions in high definition.
[0064] The device extracts audio data from the collected audio files by running them through an audio analysis algorithm. This analysis examines the tone of voice, intonation, and speaking style characteristics.
[0065] The device processes the collected interview records using a natural language processing algorithm to extract thought data. This analysis examines the content of the talent's statements and logical patterns.
[0066] Step 3: Preparing the Generative AI Model
[0067] The server prepares a generative AI model using the analyzed visual, audio, and thought data. Specifically, it initializes a deep learning model (e.g., a GAN or Transformer model) and supplies the collected data as training data.
[0068] Step 4: Generating AI Talent
[0069] The server generates AI talent using a pre-trained generative AI model. This generation process creates realistic AI talent based on input visual data, audio data, and thought data.
[0070] Step 5: Generate Ads
[0071] The device retrieves an advertising template provided by the company. This template includes storyline and design specifications.
[0072] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency.
[0073] Step 6: Review and correct the generated ads
[0074] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[0075] Step 7: NFT conversion
[0076] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[0077] Step 8: Register on the blockchain
[0078] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[0079] This allows for the recreation of a talent's peak period using AI, ensuring the sustainability of advertising strategies and enabling the utilization of talent value as a new digital asset.
[0080] (Example 1)
[0081] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0082] Currently, leveraging the influence of popular celebrities is common for advertising strategies and brand image building. However, once a celebrity's peak is over, or due to their constantly busy schedule, it often becomes difficult to secure their appearances in advertisements. To solve this problem, a system is needed that can recreate a celebrity's peak performance based on visual, auditory, and thought patterns, and effectively utilize this in advertising.
[0083] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0084] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for storing the generated AI talent data as an NFT and registering it on a blockchain; and means for obtaining an advertising template and generating an advertisement with visual and audio consistency. This makes it possible to recreate the talent's appearance during their prime while establishing an appropriate advertising strategy and brand image.
[0085] A "talent" is a famous or well-known person who is active in the media or advertising.
[0086] "Visual information" refers to digital data such as videos and images that pertain to the appearance and actions of a talent.
[0087] "Audio information" refers to digital audio data related to the voice and speaking style of a talent.
[0088] "Thinking patterns" refer to data about a celebrity's tendencies in thinking and emotions, extracted from their statements, interview content, and other sources.
[0089] "Means of collection" refers to the methods and technologies used to obtain necessary data from databases, the internet, and other sources.
[0090] "Means of analysis" refers to methods of analyzing collected data using technologies such as video analysis algorithms, audio analysis algorithms, and natural language processing algorithms.
[0091] A "generative AI model" is an artificial intelligence model that can reproduce the characteristics of a talent with high accuracy, based on their visual information, auditory information, and thought patterns during their peak.
[0092] An "AI talent" is a virtual talent created using a generative AI model.
[0093] An "ad template" refers to a basic format or design layout used to create advertising content.
[0094] "NFT" stands for Non-Fungible Token, and it refers to a unique token that uses blockchain technology to prove ownership of a digital item.
[0095] Blockchain is a technology that securely manages and trades digital data on a decentralized network.
[0096] Modes for carrying out the invention
[0097] This invention is a system that analyzes the visual information, auditory information, and thought patterns of a talent during their peak period and reproduces them using a generative AI, and is used to improve advertising strategies and establish brand image. Specific embodiments are described in detail below.
[0098] Program Description
[0099] 1. Data collection:
[0100] The terminal connects to a database and collects past video files, audio files, and interview recordings of the talent. This is done using automated scripts, which download the necessary data from sources such as YouTube® and television station archives.
[0101] 2. Data Analysis:
[0102] The device uses OpenCV to analyze video files and extract facial features and movements as visual data. For example, it can analyze the microexpressions and body movements of a celebrity in detail.
[0103] The device uses LibROSA to analyze audio files and extract audio features (pitch, tone, volume, etc.). For example, it can analyze the intonation and tone of a specific voice of a talent.
[0104] The device uses Spacy and BERT to analyze interview recordings and extract the talent's thought patterns and frequently occurring words as natural language data. For example, it analyzes characteristics such as "frequently using positive words" and "having a tendency to respond to stress."
[0105] 3. Preparing the Generative AI Model:
[0106] The server integrates the analyzed visual, audio, and thought data to train a generative AI model. This model utilizes deep learning frameworks such as TENSORFLOW® and PyTorch. For example, a large dataset containing tens of thousands of image and audio data is used to recreate a celebrity's appearance during their prime.
[0107] 4. AI Talent Generation:
[0108] The server uses a trained generative AI model to generate an AI talent that looks like the real talent in their prime. This AI talent has a very realistic appearance and voice based on actual video and audio data. For example, it can generate a scene in which the talent introduces a specific product.
[0109] 5. Ad generation:
[0110] The device retrieves advertising templates provided by companies, and the server integrates AI talent into these templates. This process utilizes video editing software such as After Effects or Premiere Pro. For example, the AI talent could create a commercial explaining the features of a new smartphone.
[0111] 6. NFT creation and blockchain registration:
[0112] The server stores the generated AI talent data as an NFT using platforms such as OpenSea and Rarible, and registers it on a blockchain such as Ethereum. This makes the AI talent tradable on the market as a unique digital asset. For example, the AI talent could be transformed into an NFT as a signed digital poster.
[0113] Specific example
[0114] For example, consider a case where past data of famous celebrities is collected and analyzed, and then AI celebrities are generated based on that data. Specifically, the following procedure is used:
[0115] 1. The device collects past video files, audio files, and interview recordings of famous celebrities.
[0116] 2. The device analyzes this data and extracts visual data, audio data, and thought data.
[0117] 3. The server uses this data to train an AI model and generate an AI talent that represents the talent in their prime.
[0118] 4. The server incorporates the generated AI talent into the ad template and generates the finished ad. This ad features the AI talent in their prime.
[0119] 5. The server stores the generated AI talent data as an NFT and registers it on the blockchain. This NFT can be traded on the market as a digital asset.
[0120] Example of a prompt
[0121] "Train an AI model to recreate the peak appearance of famous celebrities based on their past data. Provide advertising templates using the generated AI celebrity and create completed advertisements with visual and audio consistency. Also, save the data of this AI celebrity as an NFT and register it on the blockchain."
[0122] The above describes specific embodiments for carrying out the present invention.
[0123] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0124] Step 1:
[0125] Data collection
[0126] The terminal connects to a database to search for and download past video files, audio files, and interview recordings of a specified talent. Input is information such as the talent's name and the URL of the data source, and output is a set of downloaded video files, audio files, and interview recordings. For example, it can run a script that automatically collects past appearance videos using the YouTube API.
[0127] Step 2:
[0128] Analysis of video data
[0129] The device analyzes video files collected using OpenCV frame by frame. It extracts facial features and movements and saves them as visual data. The input is a video file, and the output is visual data including facial features and movements. Specifically, it uses a face detection algorithm to extract facial contours, microexpressions, and body movements.
[0130] Step 3:
[0131] Analysis of audio data
[0132] The device uses LibROSA to analyze collected audio files and extract audio features (pitch, tone, volume, etc.). The input is an audio file, and the output is audio feature data. Specifically, it converts the audio file into waveform data and analyzes the pitch range and intonation.
[0133] Step 4:
[0134] Analysis of interview data
[0135] The device uses Spacy and BERT to process interview records collected using natural language processing, extracting the talent's thought patterns and frequently occurring words. The input is the interview records, and the output is a list of thought patterns and frequently occurring words. Specifically, the text data is tokenized, and sentiment analysis and theme extraction are performed.
[0136] Step 5:
[0137] Preparing a Generative AI Model
[0138] The server integrates the visual, auditory, and thought data obtained in steps 2 through 4 to train a generative AI model. Deep learning frameworks such as TensorFlow and PyTorch are used. The input is visual, auditory, and thought data, and the output is the generative AI model. Specifically, a deep neural network is designed and trained on these datasets.
[0139] Step 6:
[0140] AI Talent Generate
[0141] The server uses a trained generative AI model to generate an AI talent that represents the talent in their prime. The input is the generative AI model, and the output is the generated AI talent. Specifically, it generates realistic video and audio from the input data to enhance the accuracy of the talent's representation.
[0142] Step 7:
[0143] Ad generation
[0144] The terminal retrieves advertising templates provided by companies, and the server integrates the generated AI talent into the advertising templates. The input is the advertising template and the generated AI talent, and the output is the completed advertising video. Specifically, video editing software such as After Effects or Premiere Pro is used to create a commercial in which the AI talent introduces a new product.
[0145] Step 8:
[0146] NFT creation and blockchain registration
[0147] The server stores the generated AI talent data as an NFT and registers it on a blockchain such as Ethereum. The input is the generated AI talent data, and the output is the NFT and its registration information on the blockchain. Specifically, it uses platforms such as OpenSea and Rarible to create the digital assets of the AI talent and records them on the blockchain.
[0148] Through the steps described above, this system can recreate a talent's peak period with high fidelity, contributing to advertising strategies and the establishment of a brand image.
[0149] (Application Example 1)
[0150] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0151] In today's advertising industry, there is a growing demand to effectively leverage the image of a celebrity during their peak to promote products and services. However, integrating the latest advertising content while recreating the video, audio, and thought patterns of a celebrity from their prime is difficult. Furthermore, there is a lack of mechanisms to manage this generated data and preserve its value as a digital asset. Therefore, there is a need for an advertising generation system that utilizes celebrity data.
[0152] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0153] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content automatically generated based on user input; and means for storing the generated AI talent data as an NFT and registering it on the blockchain. This makes it possible to automatically generate highly accurate advertising content using data from the talent's peak period and guarantee its asset value.
[0154] "Visual information" refers to video data related to the talent's appearance, movements, posture, facial expressions, etc.
[0155] "Audio information" refers to audio data related to a talent's voice quality, pronunciation, intonation, speaking style, etc.
[0156] "Thinking patterns" refer to tendencies in a celebrity's thinking, opinions, and ideas, based on their past interviews and statements.
[0157] A "generative AI model" refers to an artificial intelligence model that generates AI talent based on collected and analyzed visual information, audio information, and thought patterns.
[0158] "AI talent" refers to a virtual character created based on collected and analyzed data, possessing the appearance, voice, and thought patterns of a talent during their prime.
[0159] "Advertising content" refers to digital media, including videos and audio, used to promote products and services.
[0160] An "ad template" refers to a format or design template used when generating advertising content.
[0161] "NFT" stands for Non-Fungible Token, and refers to a token used to prove ownership or uniqueness of a digital asset.
[0162] "Blockchain" refers to a database technology that manages digital data in a decentralized manner and is resistant to tampering.
[0163] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak period, with the aim of establishing advertising strategies and brand image. This system is implemented as follows.
[0164] 1. Data collection:
[0165] Users collect past video files, audio files, and interview records of talents from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak. To do this, they extract the necessary data from the database using programming languages such as Python.
[0166] 2. Data Analysis:
[0167] The server extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm. Specifically, it uses PIL (Python Imaging Library) for video analysis, the TextToSpeech library for speech analysis, and spaCy and transformers for natural language processing.
[0168] 3. Preparing the Generative AI Model:
[0169] The server prepares a generative AI model based on the analyzed data. This generative AI model is trained using deep learning to recreate the talent's peak performance, processing visual information, audio information, and thought patterns as input data. Frameworks such as TensorFlow and PyTorch are used to train the deep learning model.
[0170] 4. AI Talent Generation:
[0171] The server generates realistic AI talents using a prepared generative AI model. These AI talents represent real-life talents in their prime and possess high fidelity in terms of visual, auditory, and cognitive aspects. The generated AI talents are then used in advertising content for products and services.
[0172] 5. Ad generation:
[0173] The server retrieves the ad template provided by the user and incorporates the generated AI talent into the ad template. It then adjusts the visual consistency and audio to generate the finished ad. The generated ad is converted into a video file as visual data and an audio file as audio data.
[0174] 6. NFT creation and blockchain registration:
[0175] The server stores the generated AI talent data as NFTs. Each NFT is assigned a unique digital ID and registered using blockchain technology. This process makes the generated AI talent data tradable on the market as a new digital asset.
[0176] Specific example
[0177] For example, suppose a company wants to create a commercial for a new product launch that uses the image of a famous celebrity in their prime. The user collects past data on the celebrity and inputs it into the system. Next, the server generates an AI celebrity based on the analyzed data. Then, the AI celebrity is incorporated into a provided advertising template to generate the completed advertisement. The generated advertisement is used to promote the sale of the product or service.
[0178] Example of a prompt
[0179] "Based on data from Talent A's heyday, recreate the product's visual characteristics, voice, and thought patterns to create a 30-second advertisement promoting the new BeeWatch product. The advertisement features Talent A discussing the new product's features and explaining its ease of use. The background should feature a city nightscape and a luxurious set."
[0180] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0181] Step 1:
[0182] The user inputs prompt text and advertising content details to recreate the talent's heyday. This provides the system with specific target talent information and advertising direction.
[0183] Input: Talent identification information, advertising content details, prompt text
[0184] Output: List of data to be collected
[0185] Specific operation: The user enters information into an input form via a smartphone application and sends it to the server.
[0186] Step 2:
[0187] The device collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information to recreate the talent's heyday.
[0188] Input: List of data to be collected
[0189] Output: Collected video files, audio files, interview recordings
[0190] Specific operation: The device accesses the database and downloads the necessary video, audio, and interview recordings.
[0191] Step 3:
[0192] The server extracts visual data from collected video files by applying a video analysis algorithm, and extracts audio data from audio files by analyzing them with an audio analysis algorithm. Similarly, it extracts thought data from interview recordings by applying a natural language processing algorithm.
[0193] Input: Collected video files, audio files, interview recordings
[0194] Output: Visual data, audio data, thought data
[0195] Specific operation: The server analyzes video data using PIL (Python Imaging Library), analyzes audio data using the TextToSpeech library, and analyzes interview recordings using spaCy and transformers.
[0196] Step 4:
[0197] The server trains a generative AI model based on analyzed visual, audio, and thought data. This generative AI model is used to recreate the talent's peak performance. TensorFlow and PyTorch are used to train the deep learning model.
[0198] Input: Visual data, audio data, thought data
[0199] Output: Trained generative AI model
[0200] Specific operation: The server uses a deep learning framework to train an AI model and optimize the model parameters.
[0201] Step 5:
[0202] The server uses a trained generative AI model to generate AI talents that represent the talents in their prime. The AI talents are generated as virtual characters with high visual, auditory, and cognitive accuracy.
[0203] Input: Trained generative AI model
[0204] Output: Generated AI talents
[0205] Specific operation: The server runs the generated AI model to produce 3D models and voice samples that recreate the appearance of the talent.
[0206] Step 6:
[0207] The server retrieves the ad template provided by the user and incorporates the generated AI talent into the ad template. This results in the creation of a finished ad with visual consistency and audio adjustments.
[0208] Input: Generated AI talent, ad template
[0209] Output: Completed advertising content
[0210] Specific operation: The server inserts AI talent data into the ad template and uses a video editing library (e.g., moviepy) to create the final ad content.
[0211] Step 7:
[0212] The server stores the generated AI talent data as an NFT, assigns a unique digital ID to it, and registers it on the blockchain. This allows the generated AI talent data to be recognized as a digital asset and traded on the market.
[0213] Input: Generated AI talent data
[0214] Output: AI talent data as NFTs, digital ID registered on the blockchain
[0215] Specific operation: The server uses an NFT generation tool to assign a digital ID and register the data on the blockchain network.
[0216] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0217] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak, aiming to solidify advertising strategies and brand image. Furthermore, by combining it with an emotion engine that recognizes user emotions, the effectiveness of advertising content is enhanced.
[0218] 1. Data collection:
[0219] The terminal collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak period.
[0220] 2. Data Analysis:
[0221] The device extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm.
[0222] 3. Preparing the Generative AI Model:
[0223] The server prepares a generative AI model using the analyzed visual, audio, and thought data. The generative AI model initializes a deep learning model (e.g., GAN or Transformer model) and supplies the collected data as training data.
[0224] 4. AI Talent Generation:
[0225] The server generates AI talent using a pre-trained generative AI model. This generation process creates realistic AI talent based on input visual data, audio data, and thought data.
[0226] 5. Ad generation:
[0227] The device retrieves an advertising template provided by the company. This template includes storyline and design specifications.
[0228] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency.
[0229] 6. Review and correct the generated ads:
[0230] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[0231] 7. NFT creation and blockchain registration:
[0232] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[0233] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[0234] Embedding an emotion engine
[0235] Furthermore, by incorporating an emotion engine that recognizes user emotions, advertising content will be personalized.
[0236] 1. Collecting user sentiment data:
[0237] The device collects the user's facial expressions and tone of voice in real time from visual and auditory information. This captures the user's emotional state while they are watching the advertisement.
[0238] 2. Emotion analysis:
[0239] The server processes the collected visual and audio information of the user through an emotion analysis algorithm to detect the user's emotions. This analysis identifies emotions such as happiness, sadness, surprise, and excitement.
[0240] 3. Adjusting advertising content:
[0241] Based on the results of emotion analysis, the server adjusts the facial expressions and dialogue of the generated AI talent in real time. For example, if the user is excited, the server enhances the energetic expressions of the AI talent.
[0242] Specific example
[0243] For example, if you collect data on talent A during their peak, analyze it to generate an AI talent, and then use an emotion engine to personalize advertising content:
[0244] 1. The device collects past video files, audio files, and interview records of talent A.
[0245] 2. The device analyzes the collected data and extracts visual data, audio data, and thought data.
[0246] 3. The server trains a generation AI model to create an AI talent that resembles Talent A in his prime.
[0247] 4. The terminal retrieves the advertising template provided by the company, and the server incorporates the AI talent into the advertising template.
[0248] 5. The user (company representative) reviews the generated advertising content and makes corrections as needed.
[0249] 6. The server converts the generated AI talent data into an NFT and registers it on the blockchain.
[0250] 7. The terminal collects the user's visual and auditory information, and the server detects the user's emotions using an emotion analysis algorithm.
[0251] 8. Based on the results of the emotion analysis, the server adjusts the facial expressions and speech content of the AI talent to provide personalized advertising content.
[0252] This invention enables the reproduction of a talent's peak period using AI generation, ensuring the sustainability of advertising strategies and providing highly accurate advertising content that takes user emotions into consideration.
[0253] The following describes the processing flow.
[0254] Program processing flow: Detailed step-by-step explanation
[0255] Step 1: Data Collection
[0256] The terminal collects past video files, audio files, and interview recordings of the talent from a database. This includes searching for and retrieving all related media files using a specified talent ID.
[0257] Step 2: Visual Data Analysis
[0258] The device processes the collected video files using a video analysis algorithm to extract visual data. This analysis captures the talent's facial features, movements, and expressions in high definition. Specifically, it uses facial recognition and motion analysis technologies.
[0259] Step 3: Analyzing the audio data
[0260] The device extracts audio data from collected audio files by running them through an audio analysis algorithm. This analysis examines voice tone, intonation, and speaking style characteristics. For example, it may use speech recognition technology or emotion analysis technology.
[0261] Step 4: Analysis of thought data
[0262] The device processes the collected interview records using natural language processing algorithms to extract thought data. This analysis examines the content and logical patterns of the talent's statements. For example, it employs text mining and semantic analysis techniques.
[0263] Step 5: Preparing the Generative AI Model
[0264] The server prepares generative AI models using the analyzed visual, audio, and thought data. Specifically, it initializes deep learning models (such as GANs and Transformer models) and supplies the collected data as training data.
[0265] Step 6: Generating AI Talents
[0266] The server generates AI talents using a pre-trained generative AI model. The generation process creates realistic AI talents based on input visual, audio, and thought data. For example, it utilizes technologies such as synthetic facial image generation and speech synthesis.
[0267] Step 7: Obtain an ad template
[0268] The device retrieves an ad template provided by the company. This template includes the ad's storyline and design specifications.
[0269] Step 8: Generate Ads
[0270] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency. For example, it might use composite video editing technology.
[0271] Step 9: Review and correct the generated ads
[0272] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[0273] Step 10: NFT conversion
[0274] The server stores the data of the generated AI talent as an NFT. In this process, a unique digital ID of the AI talent is generated, and an NFT is created based on this.
[0275] Step 11: Registration on the blockchain
[0276] The server registers the generated NFT on the blockchain. Through this registration process, the ownership and transaction history of the NFT are recorded in the distributed ledger, ensuring its asset value.
[0277] Step 12: Collection of user's emotional data
[0278] The terminal collects the user's expressions and vocal tones in real time from visual information and audio information. Specifically, it captures the user's reactions using a camera and a microphone.
[0279] Step 13: Emotional analysis
[0280] The server applies the collected user's visual information and audio information to an emotional analysis algorithm to detect the user's emotions. In this analysis, emotions such as the user's happiness, sadness, surprise, excitement, etc. are identified.
[0281] Step 14: Adjustment of advertising content
[0282] The server adjusts the expressions and speech content of the generated AI talent in real time based on the results of the emotional analysis. For example, when the user is excited, the energetic expressions of the AI talent are enhanced.
[0283] Specific example
[0284] For example, a case where the past data of famous talent A is collected and analyzed to generate an AI talent, and at the same time, the user's emotions are analyzed to provide appropriate advertising content:
[0285] 1. The terminal collects the past video files, audio files, and interview records of Talent A.
[0286] 2. The terminal analyzes the collected video files to extract visual data, analyzes the audio files to extract audio data, and further analyzes the interview records to extract thinking data.
[0287] 3. The server trains a generated AI model based on these data and generates an AI talent with the appearance of Talent A at its prime.
[0288] 4. The terminal obtains the advertising template provided by the enterprise, and the server incorporates the generated AI talent into the template.
[0289] 5. The user (enterprise staff) checks the generated advertising content and modifies it as necessary.
[0290] 6. The server NFTs the data of the generated AI talent and registers it on the blockchain.
[0291] 7. The terminal collects the visual information and audio information of the user, and the server uses an emotion analysis algorithm to detect the user's emotion.
[0292] 8. The server adjusts the expression and speech content of the AI talent based on the result of the emotion analysis and provides advertising content optimized for the user's emotion.
[0293] According to the present invention, it is possible to reproduce the prime of a talent with a generated AI, ensure the sustainability of an advertising strategy, and provide highly accurate advertising content considering the emotions of users.
[0294] (Example 2)
[0295] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".
[0296] Traditional advertising content generation systems struggled to effectively reproduce the visual, auditory, and thought patterns of celebrities during their peak, resulting in limited advertising effectiveness. Furthermore, there was a lack of means to analyze user emotions in real time and provide personalized advertising content accordingly. Additionally, centralized management of ownership and transaction history of generated data was a challenge.
[0297] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0298] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generative AI model based on the analyzed data; means for incorporating the AI talent into advertising content; means for storing the generated AI talent data as a non-fungible token and registering it in a distributed ledger; means for collecting and analyzing user emotion data; and means for adjusting the advertising content in real time based on the emotion analysis results. This makes it possible to effectively recreate the talent's peak performance and provide personalized advertising content that responds to the user's emotions. Furthermore, registering the ownership and transaction history of the generated data in a distributed ledger enhances the reliability and transparency of the data.
[0299] "Visual information" refers to visual data extracted from video files, such as the face, body movements, and facial expressions of the talent.
[0300] "Audio information" refers to acoustic data extracted from audio files, such as the characteristics of a talent's voice, tone, and speaking style.
[0301] "Thinking patterns" refer to the way of thinking, the content of what is said, and the choice of words that can be extracted from interview transcripts and speeches of celebrities.
[0302] The "generative AI model" refers to an artificial intelligence model for generating AI talents based on data collected and analyzed using deep learning technology.
[0303] The "AI talent" refers to a virtual person created by the generative AI model that reproduces the visual information, audio information, and thinking patterns of the talent.
[0304] The "advertising content" refers to media such as videos, images, audio, and articles for the purpose of corporate promotion and marketing.
[0305] The "non-fungible token" refers to a unique token that uses blockchain technology to prove the ownership of digital data.
[0306] The "distributed ledger" refers to a distributed management system that uses blockchain technology to centrally manage data ownership and transaction histories.
[0307] The "emotion data" refers to data related to the emotional state collected from the user's expressions and voice tones.
[0308] The "emotion analysis" refers to the process of analyzing the collected emotion data to identify the user's emotional state.
[0309] The present invention is a system that uses generative AI to reproduce the visual information, audio information, and thinking patterns of a talent at their prime, aiming to establish an advertising strategy and brand image. Furthermore, by combining an emotion engine that recognizes the user's emotions, the effectiveness of the advertising content is enhanced.
[0310] First, the device collects past video files, audio files, and interview records of the talent from a database. Specifically, videos are downloaded from the internet and saved to the local disk. Audio files are retrieved from cloud storage and saved in the same way. Interview records are collected as digital data from scanned documents using OCR technology (e.g., Tesseract).
[0311] Next, the device analyzes the collected video files using a video analysis algorithm (e.g., OpenCV) to extract visual data. Similarly, audio files are analyzed using a speech analysis algorithm (e.g., Librosa) to extract audio data. Interview recordings are analyzed using a natural language processing algorithm (e.g., NLTK) to extract thought data.
[0312] Subsequently, the server prepares a generative AI model (e.g., a GAN or Transformer model) using the analyzed visual, audio, and thought data. Specifically, it initializes the model using a deep learning framework (e.g., TensorFlow or PyTorch) and supplies the collected data as training data. The server then adjusts the model's hyperparameters and selects the optimal model.
[0313] The server generates AI talents using a trained generative AI model. This process involves inputting visual, auditory, and thought data to create realistic AI talents. The generated AI talents undergo filtering and enhancement to reproduce natural facial expressions and voice tones.
[0314] The device then retrieves an advertising template provided by the company. This template includes storyline and design specifications. For example, it retrieves a Google® Slides template and downloads it from a cloud service (e.g., AWS® S3). The server incorporates the generated AI talent into the advertising template and edits the advertising video according to the specified storyline. This process uses video editing software such as Adobe Premiere Pro.
[0315] The user (company representative) reviews the generated advertising content and makes corrections as needed. For example, they play the advertising video in Adobe Premiere Pro and mark the areas that need correction. The user then checks if it matches the brand image and provides feedback.
[0316] The server stores the generated AI talent data as non-fungible tokens (NFTs) and registers them on a distributed ledger. This process involves executing smart contracts to generate a unique digital ID for the AI talent and then creating the NFT based on that ID. Blockchain platforms such as Ethereum and Polygon are used as the distributed ledger.
[0317] Furthermore, by incorporating an emotion engine that recognizes user emotions, advertising content can be personalized. The device collects the user's facial expressions and tone of voice in real time using visual information (webcam) and audio information (microphone). This data is sent to a server, where an emotion analysis algorithm (e.g., DeepFace) analyzes the user's emotions. Based on the analysis results, the server adjusts the AI talent's facial expressions and speech in real time, tailoring the advertising content to best match the user's emotional state. For example, if the user shows a surprised reaction, a corresponding catchphrase is displayed to emphasize the visual impact.
[0318] Example of a prompt
[0319] An example of a prompt message is: "Generate an advertisement based on an interview with talent A during their heyday. Analyze user sentiment data in real time to personalize the ad content."
[0320] The above describes a specific embodiment for carrying out the present invention. This makes it possible to effectively recreate a talent's heyday and provide personalized advertising content that responds to the user's emotions. Furthermore, by registering the ownership and transaction history of the generated data in a distributed ledger, the reliability and transparency of the data can be enhanced.
[0321] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0322] Step 1: Data Collection
[0323] The device collects past video files, audio files, and interview records of the talent. Specifically, it downloads video files from the internet (input) and saves them to the local disk (output). It retrieves audio files from cloud storage (input) and saves them similarly (output). Interview records are collected as digital data from scanned documents using OCR technology (e.g., Tesseract) (input) and saved as text data (output).
[0324] Step 2: Analysis of the month
[0325] The terminal analyzes collected video files using a video analysis algorithm (e.g., OpenCV) (input) and extracts visual data (output). Specifically, it performs face detection frame by frame and extracts features of facial expressions and movements. Similarly, it analyzes audio files using a speech analysis algorithm (e.g., Librosa) (input) and extracts audio data (output). Furthermore, it analyzes interview recordings using a natural language processing algorithm (e.g., NLTK) (input) and extracts thought data (output). This allows the characteristics of the talent during their peak to be obtained as data.
[0326] Step 3: Preparing the Generative AI Model
[0327] The server prepares a generative AI model (e.g., a GAN or Transformer model) using collected and analyzed visual, audio, and thought data (input). Specifically, it initializes the model using a deep learning framework (e.g., TensorFlow or PyTorch) (input) and supplies the analyzed data as training data (output). The server then adjusts the hyperparameters of the model and selects the optimal model (output).
[0328] Step 4: Generating AI Talent
[0329] The server generates AI talent using a trained generative AI model (input). Based on visual data, audio data, and thought data, it creates realistic AI talent (output). During this process, the generated AI is fine-tuned to match the actual video and audio (specifically, this includes filtering and enhancement to naturally reproduce facial expressions and voice tone).
[0330] Step 5: Generate Ads
[0331] The device retrieves an advertising template provided by the company (input). This template includes storyline and design specifications. For example, a Google Slides template is downloaded from a cloud service (e.g., AWS S3) (output). The server incorporates the generated AI talent into the advertising template (input) and edits it according to the specified storyline (output). This process uses video editing software (e.g., Adobe Premiere Pro) to add movement and effects.
[0332] Step 6: Review and correct the generated ads
[0333] The user (company representative) reviews the generated advertising content (input) and makes corrections as needed (output). For example, they might play the advertising video in Adobe Premiere Pro and mark the areas that need correction. The user then checks for consistency with the brand image and provides feedback (specifically, by marking up the corrections on the screen).
[0334] Step 7: NFT creation and blockchain registration
[0335] The server stores the generated AI talent data as non-fungible tokens (NFTs) (input) and registers them on a distributed ledger (output). This process involves executing smart contracts to generate a unique digital ID for the AI talent (input) and then creating an NFT based on that ID (output). Blockchain platforms such as Ethereum and Polygon are used as the distributed ledger (specifically, this involves issuing tokens and writing them to the blockchain).
[0336] Step 8: Integrating the Emotion Engine
[0337] The device uses a webcam and microphone to collect the user's facial expressions and voice tone in real time (input). The server processes the collected data using an emotion analysis algorithm (e.g., DeepFace) to analyze the user's emotions (output). Based on the analysis results, the server adjusts the AI talent's facial expressions and speech in real time (input), and adjusts the advertisement content to best match the user's emotional state (output). For example, if the user shows a surprised reaction, a corresponding catchphrase is displayed to emphasize the visual impact (specific actions include real-time changes in facial expressions and adjustments to voice tone).
[0338] Example of a prompt
[0339] "Generate an advertisement based on an interview with talent A during their heyday. Analyze user sentiment data in real time and personalize the ad content."
[0340] (Application Example 2)
[0341] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0342] While using celebrities and characters is common in current advertising strategies, their effectiveness has certain limitations, and providing advertising content tailored to individual user emotions in real time is a particularly challenging task. Furthermore, there is a need to generate more effective advertisements by leveraging data from celebrities' peak periods. Against this backdrop, there is a need for methods to increase user engagement by providing more personalized advertising content.
[0343] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0344] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content; means for recognizing emotions from the user's visual information and audio information and adjusting the AI talent's facial expressions and statements in real time; and means for saving the generated AI talent data as an NFT and registering it on the blockchain. This makes it possible to effectively provide personalized advertising content that responds to the user's emotions.
[0345] A "talent" is a person or character used in advertising and entertainment, primarily based on visual and auditory information.
[0346] "Visual information" refers to information that can be recognized visually, such as videos and image data of celebrities.
[0347] "Audio information" refers to information that can be recognized as sound, such as the voice and statements of a celebrity.
[0348] "Thinking patterns" refer to the tendencies in thinking and statements that a celebrity has shown in past interviews and statements.
[0349] A "generative AI model" is an artificial intelligence model that uses deep learning to analyze data and generate new data. Specifically, this includes GANs (Generative Adversarial Networks) and Transformer models.
[0350] An "AI talent" is a virtual talent generated using a generative AI model, based on collected visual information, audio information, and thought patterns of real talents.
[0351] "Advertising content" refers to media content that conveys advertising messages through visual information, audio information, text information, etc.
[0352] "Emotion recognition" is a technology that analyzes a user's visual and auditory information to determine their emotional state.
[0353] "Real-time adjustment" refers to instantly reflecting the user's current emotional state and changing the content accordingly.
[0354] "NFT" stands for "Non-Fungible Token," which is a non-fungible token that uses blockchain technology to prove ownership of digital assets.
[0355] Blockchain is a technology that uses distributed ledger technology to prevent data tampering and ensure transparency.
[0356] This invention provides personalized advertising content that reflects the user's emotions based on the following procedure. Specific hardware and software configurations are shown for each step.
[0357] Data collection
[0358] To collect visual, auditory, and thought patterns of talent, the server retrieves past video files, audio files, and interview records from a database. The databases and storage used for this purpose include cloud-based storage services such as Google Cloud Storage and Amazon S3.
[0359] Data Analysis
[0360] To analyze the collected visual, auditory, and thought patterns, the terminal uses video analysis algorithms (such as OpenCV), audio analysis algorithms (such as Librosa), and natural language processing algorithms (such as the transformers library). This allows for the extraction of the talent's visual, auditory, and thought data.
[0361] Training of generative AI models
[0362] The server prepares a generative AI model using the analyzed visual, auditory, and thought data. Generative AI models used include GANs (Generative Adversarial Networks) and Transformer models (e.g., GPT-3®). These models are based on deep learning and generate AI talents by supplying collected data as training data.
[0363] AI Talent Generate
[0364] The server uses a trained generative AI model to generate AI talents that replicate the visual information, auditory information, and thought patterns of the talent during their prime.
[0365] Creating advertising content
[0366] The device retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into this template. The advertising template includes storyline and design specifications, and the AI talent generates advertising content based on these.
[0367] User emotion recognition
[0368] When a user views advertising content, the device collects the user's visual and auditory information in real time through the smartphone's camera and microphone. Using this data, the server identifies the user's emotions using sentiment analysis algorithms (such as dlib or OpenCV).
[0369] Real-time adjustment of advertising content
[0370] The server adjusts the AI talent's facial expressions and speech in real time based on the results of sentiment analysis. This process uses GPT-2 and GPT-3 models to generate appropriate prompt sentences, which are then reflected in the advertising content.
[0371] Storage and management of AI talent data
[0372] The generated AI talent data is stored as an NFT (Non-Fungible Token) and registered using blockchain technology (such as Ethereum or Hyperledger). This generates a unique digital ID for the AI talent, and ownership and transaction history are recorded on a distributed ledger.
[0373] Specific example
[0374] If a user is watching a video ad on their smartphone, and the camera detects their facial expression and recognizes that they are "surprised," the ad message will automatically adjust to a tone such as "Amazing product!". An example of a prompt message might be, "When the user looks happy, our AI talent will generate the most appropriate ad message."
[0375] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0376] Step 1:
[0377] Collect visual information, auditory information, and thought patterns of the talent.
[0378] The server retrieves past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak. Specifically, it downloads data from cloud-based storage services (e.g., Google Cloud Storage, Amazon S3).
[0379] Step 2:
[0380] The collected visual information, auditory information, and thought patterns are analyzed.
[0381] The device analyzes data using video analysis algorithms (OpenCV), audio analysis algorithms (Librosa), and natural language processing algorithms (transformers library). Visual data, audio data, and thought data are extracted. Specifically, it detects face regions frame by frame from video files, extracts audio features from audio files, and performs semantic analysis on interview recordings.
[0382] Step 3:
[0383] Train a generative AI model.
[0384] The server prepares a generative AI model using the analyzed visual, audio, and thought data. This generative AI model (GAN or Transformer model) is trained based on deep learning. Specifically, it uses a Python deep learning framework (e.g., TensorFlow, PyTorch) to feed the dataset into the training model and optimize the model parameters.
[0385] Step 4:
[0386] Generate AI talent.
[0387] The server uses a trained generative AI model to generate AI talents that recreate the visual, auditory, and thought patterns of the talent during their prime. Specifically, when a user provides input data, the model generates realistic video and audio of the talent based on that data. This generated data is then used in the next step.
[0388] Step 5:
[0389] Create advertising content.
[0390] The device retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into this template. Specifically, the template includes storyline and design specifications, and the AI talent generates advertising content based on these.
[0391] Step 6:
[0392] Collect and analyze user sentiment data.
[0393] The device collects the user's visual and audio information in real time through the smartphone's camera and microphone. The server processes the collected data using emotion analysis algorithms (such as dlib or OpenCV) to detect emotions. Specifically, it identifies the user's emotions by analyzing facial expressions from the camera image and voice tone from the audio.
[0394] Step 7:
[0395] Adjust advertising content in real time.
[0396] Based on the results of sentiment analysis, the server adjusts the facial expressions and dialogue of the AI talent in the generated advertising content in real time. Specifically, it takes specific prompt text as input to the generating AI model and reflects the resulting text and video in the advertisement, thereby customizing it according to the user's emotions.
[0397] Step 8:
[0398] Store and manage AI talent data.
[0399] The server stores the generated AI talent data as an NFT and registers it using blockchain technology (e.g., Ethereum, Hyperledger). This generates a unique digital ID for the AI talent, and ownership and transaction history are recorded on a distributed ledger. Specifically, it sends transactions to the blockchain network, ensuring that the data is permanently recorded on the blockchain.
[0400] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0401] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0402] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0403] [Second Embodiment]
[0404] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0405] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0406] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0407] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0408] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0409] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0410] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0411] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0412] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0413] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0414] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0415] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0416] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak period, with the aim of establishing advertising strategies and brand image. This system is implemented as follows.
[0417] 1. Data collection:
[0418] The terminal collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak period.
[0419] 2. Data Analysis:
[0420] The device extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm.
[0421] 3. Preparing the Generative AI Model:
[0422] The server prepares a generative AI model based on the analyzed data. The generative AI model is trained using deep learning to recreate the celebrity's appearance during their prime, and processes visual information, auditory information, and thought patterns as input data.
[0423] 4. AI Talent Generation:
[0424] The server uses a prepared generative AI model to generate realistic AI talents. These AI talents represent the real-life talents in their prime and possess high fidelity in terms of visual, auditory, and cognitive aspects.
[0425] 5. Ad generation:
[0426] The terminal retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into the advertising template. This generates a finished advertisement with visual consistency and audio adjustments. For example, the generated AI talent can be used to create new commercial content and promote products.
[0427] 6. NFT creation and blockchain registration:
[0428] The server stores the generated AI talent data as an NFT. This NFT is assigned a unique digital ID and registered on the blockchain. As a result, the generated AI talent data is recognized as a new digital asset and becomes tradable on the market.
[0429] Specific example
[0430] For example, consider a case where past data of a famous celebrity A is collected and analyzed, and then an AI celebrity is generated based on that data.
[0431] 1. The device collects past video files, audio files, and interview records of talent A.
[0432] 2. The device analyzes collected video files to extract visual data, and analyzes audio files to extract audio data. Furthermore, it analyzes interview recordings to extract thought data.
[0433] 3. The server trains an AI model based on this data to generate an AI talent that represents Talent A in their prime.
[0434] 4. The server incorporates the generated AI talent into an advertising template provided by the company to produce the final advertisement. This advertisement features an AI talent that represents Talent A in their prime.
[0435] 5. The server stores the generated AI talent data as an NFT and registers it on the blockchain. This NFT can be traded on the market as a digital asset.
[0436] In this way, the present invention uses AI to recreate a talent's peak period, enabling the maintenance of advertising strategies and brand image, and the creation of new business models.
[0437] The following describes the processing flow.
[0438] Program processing flow: Detailed step-by-step explanation
[0439] Step 1: Data Collection
[0440] The terminal collects past video files, audio files, and interview recordings of the talent from a database. This includes searching for and retrieving all related media files using a specified talent ID.
[0441] Step 2: Analysis of the month
[0442] The device processes the collected video files using a video analysis algorithm to extract visual data. This analysis captures the talent's facial features, movements, and expressions in high definition.
[0443] The device extracts audio data from the collected audio files by running them through an audio analysis algorithm. This analysis examines the tone of voice, intonation, and speaking style characteristics.
[0444] The device processes the collected interview records using a natural language processing algorithm to extract thought data. This analysis examines the content of the talent's statements and logical patterns.
[0445] Step 3: Preparing the Generative AI Model
[0446] The server prepares a generative AI model using the analyzed visual, audio, and thought data. Specifically, it initializes a deep learning model (e.g., a GAN or Transformer model) and supplies the collected data as training data.
[0447] Step 4: Generating AI Talent
[0448] The server generates AI talent using a pre-trained generative AI model. This generation process creates realistic AI talent based on input visual data, audio data, and thought data.
[0449] Step 5: Generate Ads
[0450] The device retrieves an advertising template provided by the company. This template includes storyline and design specifications.
[0451] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency.
[0452] Step 6: Review and correct the generated ads
[0453] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[0454] Step 7: NFT conversion
[0455] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[0456] Step 8: Register on the blockchain
[0457] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[0458] This allows for the recreation of a talent's peak period using AI, ensuring the sustainability of advertising strategies and enabling the utilization of talent value as a new digital asset.
[0459] (Example 1)
[0460] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0461] Currently, leveraging the influence of popular celebrities is common for advertising strategies and brand image building. However, once a celebrity's peak is over, or due to their constantly busy schedule, it often becomes difficult to secure their appearances in advertisements. To solve this problem, a system is needed that can recreate a celebrity's peak performance based on visual, auditory, and thought patterns, and effectively utilize this in advertising.
[0462] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0463] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for storing the generated AI talent data as an NFT and registering it on a blockchain; and means for obtaining an advertising template and generating an advertisement with visual and audio consistency. This makes it possible to recreate the talent's appearance during their prime while establishing an appropriate advertising strategy and brand image.
[0464] A "talent" is a famous or well-known person who is active in the media or advertising.
[0465] "Visual information" refers to digital data such as videos and images that pertain to the appearance and actions of a talent.
[0466] "Audio information" refers to digital audio data related to the voice and speaking style of a talent.
[0467] "Thinking patterns" refer to data about a celebrity's tendencies in thinking and emotions, extracted from their statements, interview content, and other sources.
[0468] "Means of collection" refers to the methods and technologies used to obtain necessary data from databases, the internet, and other sources.
[0469] "Means of analysis" refers to methods of analyzing collected data using technologies such as video analysis algorithms, audio analysis algorithms, and natural language processing algorithms.
[0470] A "generative AI model" is an artificial intelligence model that can reproduce the characteristics of a talent with high accuracy, based on their visual information, auditory information, and thought patterns during their peak.
[0471] An "AI talent" is a virtual talent created using a generative AI model.
[0472] An "ad template" refers to a basic format or design layout used to create advertising content.
[0473] "NFT" stands for Non-Fungible Token, and it refers to a unique token that uses blockchain technology to prove ownership of a digital item.
[0474] Blockchain is a technology that securely manages and trades digital data on a decentralized network.
[0475] Modes for carrying out the invention
[0476] This invention is a system that analyzes the visual information, auditory information, and thought patterns of a talent during their peak period and reproduces them using a generative AI, and is used to improve advertising strategies and establish brand image. Specific embodiments are described in detail below.
[0477] Program Description
[0478] 1. Data collection:
[0479] The terminal connects to a database and collects past video files, audio files, and interview recordings of the talent. This uses automated scripts to download necessary data from sources such as YouTube and television station archives.
[0480] 2. Data Analysis:
[0481] The device uses OpenCV to analyze video files and extract facial features and movements as visual data. For example, it can analyze the microexpressions and body movements of a celebrity in detail.
[0482] The device uses LibROSA to analyze audio files and extract audio features (pitch, tone, volume, etc.). For example, it can analyze the intonation and tone of a specific voice of a talent.
[0483] The device uses Spacy and BERT to analyze interview recordings and extract the talent's thought patterns and frequently occurring words as natural language data. For example, it analyzes characteristics such as "frequently using positive words" and "having a tendency to respond to stress."
[0484] 3. Preparing the Generative AI Model:
[0485] The server integrates the analyzed visual, audio, and thought data to train a generative AI model. This model utilizes deep learning frameworks such as TensorFlow and PyTorch. For example, a large dataset containing tens of thousands of image and audio data is used to recreate a celebrity's appearance during their prime.
[0486] 4. AI Talent Generation:
[0487] The server uses a trained generative AI model to generate an AI talent that looks like the real talent in their prime. This AI talent has a very realistic appearance and voice based on actual video and audio data. For example, it can generate a scene in which the talent introduces a specific product.
[0488] 5. Ad generation:
[0489] The device retrieves advertising templates provided by companies, and the server integrates AI talent into these templates. This process utilizes video editing software such as After Effects or Premiere Pro. For example, the AI talent could create a commercial explaining the features of a new smartphone.
[0490] 6. NFT creation and blockchain registration:
[0491] The server stores the generated AI talent data as an NFT using platforms such as OpenSea and Rarible, and registers it on a blockchain such as Ethereum. This makes the AI talent tradable on the market as a unique digital asset. For example, the AI talent could be transformed into an NFT as a signed digital poster.
[0492] Specific example
[0493] For example, consider a case where past data of famous celebrities is collected and analyzed, and then AI celebrities are generated based on that data. Specifically, the following procedure is used:
[0494] 1. The device collects past video files, audio files, and interview recordings of famous celebrities.
[0495] 2. The device analyzes this data and extracts visual data, audio data, and thought data.
[0496] 3. The server uses this data to train an AI model and generate an AI talent that represents the talent in their prime.
[0497] 4. The server incorporates the generated AI talent into the ad template and generates the finished ad. This ad features the AI talent in their prime.
[0498] 5. The server stores the generated AI talent data as an NFT and registers it on the blockchain. This NFT can be traded on the market as a digital asset.
[0499] Example of a prompt
[0500] "Train an AI model to recreate the peak appearance of famous celebrities based on their past data. Provide advertising templates using the generated AI celebrity and create completed advertisements with visual and audio consistency. Also, save the data of this AI celebrity as an NFT and register it on the blockchain."
[0501] The above describes specific embodiments for carrying out the present invention.
[0502] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0503] Step 1:
[0504] Data collection
[0505] The terminal connects to a database to search for and download past video files, audio files, and interview recordings of a specified talent. Input is information such as the talent's name and the URL of the data source, and output is a set of downloaded video files, audio files, and interview recordings. For example, it can run a script that automatically collects past appearance videos using the YouTube API.
[0506] Step 2:
[0507] Analysis of video data
[0508] The device analyzes video files collected using OpenCV frame by frame. It extracts facial features and movements and saves them as visual data. The input is a video file, and the output is visual data including facial features and movements. Specifically, it uses a face detection algorithm to extract facial contours, microexpressions, and body movements.
[0509] Step 3:
[0510] Analysis of audio data
[0511] The device uses LibROSA to analyze collected audio files and extract audio features (pitch, tone, volume, etc.). The input is an audio file, and the output is audio feature data. Specifically, it converts the audio file into waveform data and analyzes the pitch range and intonation.
[0512] Step 4:
[0513] Analysis of interview data
[0514] The device uses Spacy and BERT to process interview records collected using natural language processing, extracting the talent's thought patterns and frequently occurring words. The input is the interview records, and the output is a list of thought patterns and frequently occurring words. Specifically, the text data is tokenized, and sentiment analysis and theme extraction are performed.
[0515] Step 5:
[0516] Preparing a Generative AI Model
[0517] The server integrates the visual, auditory, and thought data obtained in steps 2 through 4 to train a generative AI model. Deep learning frameworks such as TensorFlow and PyTorch are used. The input is visual, auditory, and thought data, and the output is the generative AI model. Specifically, a deep neural network is designed and trained on these datasets.
[0518] Step 6:
[0519] AI Talent Generate
[0520] The server uses a trained generative AI model to generate an AI talent that represents the talent in their prime. The input is the generative AI model, and the output is the generated AI talent. Specifically, it generates realistic video and audio from the input data to enhance the accuracy of the talent's representation.
[0521] Step 7:
[0522] Ad generation
[0523] The terminal retrieves advertising templates provided by companies, and the server integrates the generated AI talent into the advertising templates. The input is the advertising template and the generated AI talent, and the output is the completed advertising video. Specifically, video editing software such as After Effects or Premiere Pro is used to create a commercial in which the AI talent introduces a new product.
[0524] Step 8:
[0525] NFT creation and blockchain registration
[0526] The server stores the generated AI talent data as an NFT and registers it on a blockchain such as Ethereum. The input is the generated AI talent data, and the output is the NFT and its registration information on the blockchain. Specifically, it uses platforms such as OpenSea and Rarible to create the digital assets of the AI talent and records them on the blockchain.
[0527] Through the steps described above, this system can recreate a talent's peak period with high fidelity, contributing to advertising strategies and the establishment of a brand image.
[0528] (Application Example 1)
[0529] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0530] In today's advertising industry, there is a growing demand to effectively leverage the image of a celebrity during their peak to promote products and services. However, integrating the latest advertising content while recreating the video, audio, and thought patterns of a celebrity from their prime is difficult. Furthermore, there is a lack of mechanisms to manage this generated data and preserve its value as a digital asset. Therefore, there is a need for an advertising generation system that utilizes celebrity data.
[0531] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0532] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content automatically generated based on user input; and means for storing the generated AI talent data as an NFT and registering it on the blockchain. This makes it possible to automatically generate highly accurate advertising content using data from the talent's peak period and guarantee its asset value.
[0533] "Visual information" refers to video data related to the talent's appearance, movements, posture, facial expressions, etc.
[0534] "Audio information" refers to audio data related to a talent's voice quality, pronunciation, intonation, speaking style, etc.
[0535] "Thinking patterns" refer to tendencies in a celebrity's thinking, opinions, and ideas, based on their past interviews and statements.
[0536] A "generative AI model" refers to an artificial intelligence model that generates AI talent based on collected and analyzed visual information, audio information, and thought patterns.
[0537] "AI talent" refers to a virtual character created based on collected and analyzed data, possessing the appearance, voice, and thought patterns of a talent during their prime.
[0538] "Advertising content" refers to digital media, including videos and audio, used to promote products and services.
[0539] An "ad template" refers to a format or design template used when generating advertising content.
[0540] "NFT" stands for Non-Fungible Token, and refers to a token used to prove ownership or uniqueness of a digital asset.
[0541] "Blockchain" refers to a database technology that manages digital data in a decentralized manner and is resistant to tampering.
[0542] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak period, with the aim of establishing advertising strategies and brand image. This system is implemented as follows.
[0543] 1. Data collection:
[0544] Users collect past video files, audio files, and interview records of talents from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak. To do this, they extract the necessary data from the database using programming languages such as Python.
[0545] 2. Data Analysis:
[0546] The server extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm. Specifically, it uses PIL (Python Imaging Library) for video analysis, the TextToSpeech library for speech analysis, and spaCy and transformers for natural language processing.
[0547] 3. Preparing the Generative AI Model:
[0548] The server prepares a generative AI model based on the analyzed data. This generative AI model is trained using deep learning to recreate the talent's peak performance, processing visual information, audio information, and thought patterns as input data. Frameworks such as TensorFlow and PyTorch are used to train the deep learning model.
[0549] 4. AI Talent Generation:
[0550] The server generates realistic AI talents using a prepared generative AI model. These AI talents represent real-life talents in their prime and possess high fidelity in terms of visual, auditory, and cognitive aspects. The generated AI talents are then used in advertising content for products and services.
[0551] 5. Ad generation:
[0552] The server retrieves the ad template provided by the user and incorporates the generated AI talent into the ad template. It then adjusts the visual consistency and audio to generate the finished ad. The generated ad is converted into a video file as visual data and an audio file as audio data.
[0553] 6. NFT creation and blockchain registration:
[0554] The server stores the generated AI talent data as NFTs. Each NFT is assigned a unique digital ID and registered using blockchain technology. This process makes the generated AI talent data tradable on the market as a new digital asset.
[0555] Specific example
[0556] For example, suppose a company wants to create a commercial for a new product launch that uses the image of a famous celebrity in their prime. The user collects past data on the celebrity and inputs it into the system. Next, the server generates an AI celebrity based on the analyzed data. Then, the AI celebrity is incorporated into a provided advertising template to generate the completed advertisement. The generated advertisement is used to promote the sale of the product or service.
[0557] Example of a prompt
[0558] "Based on data from Talent A's heyday, recreate the product's visual characteristics, voice, and thought patterns to create a 30-second advertisement promoting the new BeeWatch product. The advertisement features Talent A discussing the new product's features and explaining its ease of use. The background should feature a city nightscape and a luxurious set."
[0559] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0560] Step 1:
[0561] The user inputs prompt text and advertising content details to recreate the talent's heyday. This provides the system with specific target talent information and advertising direction.
[0562] Input: Talent identification information, advertising content details, prompt text
[0563] Output: List of data to be collected
[0564] Specific operation: The user enters information into an input form via a smartphone application and sends it to the server.
[0565] Step 2:
[0566] The device collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information to recreate the talent's heyday.
[0567] Input: List of data to be collected
[0568] Output: Collected video files, audio files, interview recordings
[0569] Specific operation: The device accesses the database and downloads the necessary video, audio, and interview recordings.
[0570] Step 3:
[0571] The server extracts visual data from collected video files by applying a video analysis algorithm, and extracts audio data from audio files by analyzing them with an audio analysis algorithm. Similarly, it extracts thought data from interview recordings by applying a natural language processing algorithm.
[0572] Input: Collected video files, audio files, interview recordings
[0573] Output: Visual data, audio data, thought data
[0574] Specific operation: The server analyzes video data using PIL (Python Imaging Library), analyzes audio data using the TextToSpeech library, and analyzes interview recordings using spaCy and transformers.
[0575] Step 4:
[0576] The server trains a generative AI model based on analyzed visual, audio, and thought data. This generative AI model is used to recreate the talent's peak performance. TensorFlow and PyTorch are used to train the deep learning model.
[0577] Input: Visual data, audio data, thought data
[0578] Output: Trained generative AI model
[0579] Specific operation: The server uses a deep learning framework to train an AI model and optimize the model parameters.
[0580] Step 5:
[0581] The server uses a trained generative AI model to generate AI talents that represent the talents in their prime. The AI talents are generated as virtual characters with high visual, auditory, and cognitive accuracy.
[0582] Input: Trained generative AI model
[0583] Output: Generated AI talents
[0584] Specific operation: The server runs the generated AI model to produce 3D models and voice samples that recreate the appearance of the talent.
[0585] Step 6:
[0586] The server retrieves the ad template provided by the user and incorporates the generated AI talent into the ad template. This results in the creation of a finished ad with visual consistency and audio adjustments.
[0587] Input: Generated AI talent, ad template
[0588] Output: Completed advertising content
[0589] Specific operation: The server inserts AI talent data into the ad template and uses a video editing library (e.g., moviepy) to create the final ad content.
[0590] Step 7:
[0591] The server stores the generated AI talent data as an NFT, assigns a unique digital ID to it, and registers it on the blockchain. This allows the generated AI talent data to be recognized as a digital asset and traded on the market.
[0592] Input: Generated AI talent data
[0593] Output: AI talent data as NFTs, digital ID registered on the blockchain
[0594] Specific operation: The server uses an NFT generation tool to assign a digital ID and register the data on the blockchain network.
[0595] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0596] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak, aiming to solidify advertising strategies and brand image. Furthermore, by combining it with an emotion engine that recognizes user emotions, the effectiveness of advertising content is enhanced.
[0597] 1. Data collection:
[0598] The terminal collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak period.
[0599] 2. Data Analysis:
[0600] The device extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm.
[0601] 3. Preparing the Generative AI Model:
[0602] The server prepares a generative AI model using the analyzed visual, audio, and thought data. The generative AI model initializes a deep learning model (e.g., GAN or Transformer model) and supplies the collected data as training data.
[0603] 4. AI Talent Generation:
[0604] The server generates AI talent using a pre-trained generative AI model. This generation process creates realistic AI talent based on input visual data, audio data, and thought data.
[0605] 5. Ad generation:
[0606] The device retrieves an advertising template provided by the company. This template includes storyline and design specifications.
[0607] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency.
[0608] 6. Review and correct the generated ads:
[0609] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[0610] 7. NFT creation and blockchain registration:
[0611] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[0612] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[0613] Embedding an emotion engine
[0614] Furthermore, by incorporating an emotion engine that recognizes user emotions, advertising content will be personalized.
[0615] 1. Collecting user sentiment data:
[0616] The device collects the user's facial expressions and tone of voice in real time from visual and auditory information. This captures the user's emotional state while they are watching the advertisement.
[0617] 2. Emotion analysis:
[0618] The server processes the collected visual and audio information of the user through an emotion analysis algorithm to detect the user's emotions. This analysis identifies emotions such as happiness, sadness, surprise, and excitement.
[0619] 3. Adjusting advertising content:
[0620] Based on the results of emotion analysis, the server adjusts the facial expressions and dialogue of the generated AI talent in real time. For example, if the user is excited, the server enhances the energetic expressions of the AI talent.
[0621] Specific example
[0622] For example, if you collect data on talent A during their peak, analyze it to generate an AI talent, and then use an emotion engine to personalize advertising content:
[0623] 1. The device collects past video files, audio files, and interview records of talent A.
[0624] 2. The device analyzes the collected data and extracts visual data, audio data, and thought data.
[0625] 3. The server trains a generation AI model to create an AI talent that resembles Talent A in his prime.
[0626] 4. The terminal retrieves the advertising template provided by the company, and the server incorporates the AI talent into the advertising template.
[0627] 5. The user (company representative) reviews the generated advertising content and makes corrections as needed.
[0628] 6. The server converts the generated AI talent data into an NFT and registers it on the blockchain.
[0629] 7. The terminal collects the user's visual and auditory information, and the server detects the user's emotions using an emotion analysis algorithm.
[0630] 8. Based on the results of the emotion analysis, the server adjusts the facial expressions and speech content of the AI talent to provide personalized advertising content.
[0631] This invention enables the reproduction of a talent's peak period using AI generation, ensuring the sustainability of advertising strategies and providing highly accurate advertising content that takes user emotions into consideration.
[0632] The following describes the processing flow.
[0633] Program processing flow: Detailed step-by-step explanation
[0634] Step 1: Data Collection
[0635] The terminal collects past video files, audio files, and interview recordings of the talent from a database. This includes searching for and retrieving all related media files using a specified talent ID.
[0636] Step 2: Visual Data Analysis
[0637] The device processes the collected video files using a video analysis algorithm to extract visual data. This analysis captures the talent's facial features, movements, and expressions in high definition. Specifically, it uses facial recognition and motion analysis technologies.
[0638] Step 3: Analyzing the audio data
[0639] The device extracts audio data from collected audio files by running them through an audio analysis algorithm. This analysis examines voice tone, intonation, and speaking style characteristics. For example, it may use speech recognition technology or emotion analysis technology.
[0640] Step 4: Analysis of thought data
[0641] The device processes the collected interview records using natural language processing algorithms to extract thought data. This analysis examines the content and logical patterns of the talent's statements. For example, it employs text mining and semantic analysis techniques.
[0642] Step 5: Preparing the Generative AI Model
[0643] The server prepares generative AI models using the analyzed visual, audio, and thought data. Specifically, it initializes deep learning models (such as GANs and Transformer models) and supplies the collected data as training data.
[0644] Step 6: Generating AI Talents
[0645] The server generates AI talents using a pre-trained generative AI model. The generation process creates realistic AI talents based on input visual, audio, and thought data. For example, it utilizes technologies such as synthetic facial image generation and speech synthesis.
[0646] Step 7: Obtain an ad template
[0647] The device retrieves an ad template provided by the company. This template includes the ad's storyline and design specifications.
[0648] Step 8: Generate Ads
[0649] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency. For example, it might use composite video editing technology.
[0650] Step 9: Review and correct the generated ads
[0651] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[0652] Step 10: NFT conversion
[0653] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[0654] Step 11: Register on the blockchain
[0655] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[0656] Step 12: Collecting user sentiment data
[0657] The device collects the user's facial expressions and tone of voice in real time from visual and auditory information. Specifically, it uses a camera and microphone to capture the user's reactions.
[0658] Step 13: Emotion Analysis
[0659] The server uses an emotion analysis algorithm to process the collected visual and audio information of the user to detect their emotions. This analysis identifies emotions such as happiness, sadness, surprise, and excitement.
[0660] Step 14: Adjusting ad content
[0661] Based on the results of emotion analysis, the server adjusts the facial expressions and dialogue of the generated AI talent in real time. For example, if the user is excited, the server enhances the energetic expressions of the AI talent.
[0662] Specific example
[0663] For example, consider a case where past data of a famous celebrity A is collected and analyzed to generate an AI talent, while simultaneously analyzing user emotions to provide appropriate advertising content:
[0664] 1. The device collects past video files, audio files, and interview records of talent A.
[0665] 2. The device analyzes collected video files to extract visual data, and analyzes audio files to extract audio data. Furthermore, it analyzes interview recordings to extract thought data.
[0666] 3. The server trains a generation AI model based on this data and generates an AI talent that resembles Talent A in his prime.
[0667] 4. The terminal retrieves the advertising template provided by the company, and the server incorporates the generated AI talent into the template.
[0668] 5. The user (company representative) reviews the generated advertising content and makes corrections as necessary.
[0669] 6. The server converts the generated AI talent data into an NFT and registers it on the blockchain.
[0670] 7. The terminal collects the user's visual and auditory information, and the server uses an emotion analysis algorithm to detect the user's emotions.
[0671] 8. The server adjusts the AI talent's facial expressions and dialogue based on the results of the emotion analysis, providing advertising content optimized for the user's emotions.
[0672] This invention enables the reproduction of a talent's peak period using AI generation, ensuring the sustainability of advertising strategies and providing highly accurate advertising content that takes user emotions into consideration.
[0673] (Example 2)
[0674] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0675] Traditional advertising content generation systems struggled to effectively reproduce the visual, auditory, and thought patterns of celebrities during their peak, resulting in limited advertising effectiveness. Furthermore, there was a lack of means to analyze user emotions in real time and provide personalized advertising content accordingly. Additionally, centralized management of ownership and transaction history of generated data was a challenge.
[0676] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0677] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generative AI model based on the analyzed data; means for incorporating the AI talent into advertising content; means for storing the generated AI talent data as a non-fungible token and registering it in a distributed ledger; means for collecting and analyzing user emotion data; and means for adjusting the advertising content in real time based on the emotion analysis results. This makes it possible to effectively recreate the talent's peak performance and provide personalized advertising content that responds to the user's emotions. Furthermore, registering the ownership and transaction history of the generated data in a distributed ledger enhances the reliability and transparency of the data.
[0678] "Visual information" refers to visual data extracted from video files, such as the face, body movements, and facial expressions of the talent.
[0679] "Audio information" refers to acoustic data extracted from audio files, such as the characteristics of a talent's voice, tone, and speaking style.
[0680] "Thinking patterns" refer to the way of thinking, the content of what is said, and the choice of words that can be extracted from interview transcripts and speeches of celebrities.
[0681] A "generative AI model" is an artificial intelligence model used to generate AI talents based on data collected and analyzed using deep learning technology.
[0682] An "AI talent" is a virtual person created by a generative AI model that reproduces the visual information, auditory information, and thought patterns of a real talent.
[0683] "Advertising content" refers to media such as videos, images, audio, and text used for the purpose of promoting or marketing a company.
[0684] A "non-fungible token" is a unique token that uses blockchain technology to prove ownership of digital data.
[0685] A "distributed ledger" is a distributed management system that uses blockchain technology to centrally manage data ownership and transaction history.
[0686] "Emotional data" refers to data about a user's emotional state, collected from their facial expressions and tone of voice.
[0687] "Sentiment analysis" is the process of analyzing collected emotional data to identify the user's emotional state.
[0688] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak, aiming to solidify advertising strategies and brand image. Furthermore, by combining it with an emotion engine that recognizes user emotions, the effectiveness of advertising content is enhanced.
[0689] First, the device collects past video files, audio files, and interview records of the talent from a database. Specifically, videos are downloaded from the internet and saved to the local disk. Audio files are retrieved from cloud storage and saved in the same way. Interview records are collected as digital data from scanned documents using OCR technology (e.g., Tesseract).
[0690] Next, the device analyzes the collected video files using a video analysis algorithm (e.g., OpenCV) to extract visual data. Similarly, audio files are analyzed using a speech analysis algorithm (e.g., Librosa) to extract audio data. Interview recordings are analyzed using a natural language processing algorithm (e.g., NLTK) to extract thought data.
[0691] Subsequently, the server prepares a generative AI model (e.g., a GAN or Transformer model) using the analyzed visual, audio, and thought data. Specifically, it initializes the model using a deep learning framework (e.g., TensorFlow or PyTorch) and supplies the collected data as training data. The server then adjusts the model's hyperparameters and selects the optimal model.
[0692] The server generates AI talents using a trained generative AI model. This process involves inputting visual, auditory, and thought data to create realistic AI talents. The generated AI talents undergo filtering and enhancement to reproduce natural facial expressions and voice tones.
[0693] The device then retrieves an advertising template provided by the company. This template includes storyline and design specifications. For example, it retrieves a Google Slides template and downloads it from a cloud service (e.g., AWS S3). The server incorporates the generated AI talent into the advertising template and edits the advertising video according to the specified storyline. This process uses video editing software such as Adobe Premiere Pro.
[0694] The user (company representative) reviews the generated advertising content and makes corrections as needed. For example, they play the advertising video in Adobe Premiere Pro and mark the areas that need correction. The user then checks if it matches the brand image and provides feedback.
[0695] The server stores the generated AI talent data as non-fungible tokens (NFTs) and registers them on a distributed ledger. This process involves executing smart contracts to generate a unique digital ID for the AI talent and then creating the NFT based on that ID. Blockchain platforms such as Ethereum and Polygon are used as the distributed ledger.
[0696] Furthermore, by incorporating an emotion engine that recognizes user emotions, advertising content can be personalized. The device collects the user's facial expressions and tone of voice in real time using visual information (webcam) and audio information (microphone). This data is sent to a server, where an emotion analysis algorithm (e.g., DeepFace) analyzes the user's emotions. Based on the analysis results, the server adjusts the AI talent's facial expressions and speech in real time, tailoring the advertising content to best match the user's emotional state. For example, if the user shows a surprised reaction, a corresponding catchphrase is displayed to emphasize the visual impact.
[0697] Example of a prompt
[0698] An example of a prompt message is: "Generate an advertisement based on an interview with talent A during their heyday. Analyze user sentiment data in real time to personalize the ad content."
[0699] The above describes a specific embodiment for carrying out the present invention. This makes it possible to effectively recreate a talent's heyday and provide personalized advertising content that responds to the user's emotions. Furthermore, by registering the ownership and transaction history of the generated data in a distributed ledger, the reliability and transparency of the data can be enhanced.
[0700] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0701] Step 1: Data Collection
[0702] The device collects past video files, audio files, and interview records of the talent. Specifically, it downloads video files from the internet (input) and saves them to the local disk (output). It retrieves audio files from cloud storage (input) and saves them similarly (output). Interview records are collected as digital data from scanned documents using OCR technology (e.g., Tesseract) (input) and saved as text data (output).
[0703] Step 2: Analysis of the month
[0704] The terminal analyzes collected video files using a video analysis algorithm (e.g., OpenCV) (input) and extracts visual data (output). Specifically, it performs face detection frame by frame and extracts features of facial expressions and movements. Similarly, it analyzes audio files using a speech analysis algorithm (e.g., Librosa) (input) and extracts audio data (output). Furthermore, it analyzes interview recordings using a natural language processing algorithm (e.g., NLTK) (input) and extracts thought data (output). This allows the characteristics of the talent during their peak to be obtained as data.
[0705] Step 3: Preparing the Generative AI Model
[0706] The server prepares a generative AI model (e.g., a GAN or Transformer model) using collected and analyzed visual, audio, and thought data (input). Specifically, it initializes the model using a deep learning framework (e.g., TensorFlow or PyTorch) (input) and supplies the analyzed data as training data (output). The server then adjusts the hyperparameters of the model and selects the optimal model (output).
[0707] Step 4: Generating AI Talent
[0708] The server generates AI talent using a trained generative AI model (input). Based on visual data, audio data, and thought data, it creates realistic AI talent (output). During this process, the generated AI is fine-tuned to match the actual video and audio (specifically, this includes filtering and enhancement to naturally reproduce facial expressions and voice tone).
[0709] Step 5: Generate Ads
[0710] The device retrieves an advertising template provided by the company (input). This template includes storyline and design specifications. For example, a Google Slides template is downloaded from a cloud service (e.g., AWS S3) (output). The server incorporates the generated AI talent into the advertising template (input) and edits it according to the specified storyline (output). This process uses video editing software (e.g., Adobe Premiere Pro) to add movement and effects.
[0711] Step 6: Review and correct the generated ads
[0712] The user (company representative) reviews the generated advertising content (input) and makes corrections as needed (output). For example, they might play the advertising video in Adobe Premiere Pro and mark the areas that need correction. The user then checks for consistency with the brand image and provides feedback (specifically, by marking up the corrections on the screen).
[0713] Step 7: NFT creation and blockchain registration
[0714] The server stores the generated AI talent data as non-fungible tokens (NFTs) (input) and registers them on a distributed ledger (output). This process involves executing smart contracts to generate a unique digital ID for the AI talent (input) and then creating an NFT based on that ID (output). Blockchain platforms such as Ethereum and Polygon are used as the distributed ledger (specifically, this involves issuing tokens and writing them to the blockchain).
[0715] Step 8: Integrating the Emotion Engine
[0716] The device uses a webcam and microphone to collect the user's facial expressions and voice tone in real time (input). The server processes the collected data using an emotion analysis algorithm (e.g., DeepFace) to analyze the user's emotions (output). Based on the analysis results, the server adjusts the AI talent's facial expressions and speech in real time (input), and adjusts the advertisement content to best match the user's emotional state (output). For example, if the user shows a surprised reaction, a corresponding catchphrase is displayed to emphasize the visual impact (specific actions include real-time changes in facial expressions and adjustments to voice tone).
[0717] Example of a prompt
[0718] "Generate an advertisement based on an interview with talent A during their heyday. Analyze user sentiment data in real time and personalize the ad content."
[0719] (Application Example 2)
[0720] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0721] While using celebrities and characters is common in current advertising strategies, their effectiveness has certain limitations, and providing advertising content tailored to individual user emotions in real time is a particularly challenging task. Furthermore, there is a need to generate more effective advertisements by leveraging data from celebrities' peak periods. Against this backdrop, there is a need for methods to increase user engagement by providing more personalized advertising content.
[0722] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0723] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content; means for recognizing emotions from the user's visual information and audio information and adjusting the AI talent's facial expressions and statements in real time; and means for saving the generated AI talent data as an NFT and registering it on the blockchain. This makes it possible to effectively provide personalized advertising content that responds to the user's emotions.
[0724] A "talent" is a person or character used in advertising and entertainment, primarily based on visual and auditory information.
[0725] "Visual information" refers to information that can be recognized visually, such as videos and image data of celebrities.
[0726] "Audio information" refers to information that can be recognized as sound, such as the voice and statements of a celebrity.
[0727] "Thinking patterns" refer to the tendencies in thinking and statements that a celebrity has shown in past interviews and statements.
[0728] A "generative AI model" is an artificial intelligence model that uses deep learning to analyze data and generate new data. Specifically, this includes GANs (Generative Adversarial Networks) and Transformer models.
[0729] An "AI talent" is a virtual talent generated using a generative AI model, based on collected visual information, audio information, and thought patterns of real talents.
[0730] "Advertising content" refers to media content that conveys advertising messages through visual information, audio information, text information, etc.
[0731] "Emotion recognition" is a technology that analyzes a user's visual and auditory information to determine their emotional state.
[0732] "Real-time adjustment" refers to instantly reflecting the user's current emotional state and changing the content accordingly.
[0733] "NFT" stands for "Non-Fungible Token," which is a non-fungible token that uses blockchain technology to prove ownership of digital assets.
[0734] Blockchain is a technology that uses distributed ledger technology to prevent data tampering and ensure transparency.
[0735] This invention provides personalized advertising content that reflects the user's emotions based on the following procedure. Specific hardware and software configurations are shown for each step.
[0736] Data collection
[0737] To collect visual, auditory, and thought patterns of talent, the server retrieves past video files, audio files, and interview records from a database. The databases and storage used for this purpose include cloud-based storage services such as Google Cloud Storage and Amazon S3.
[0738] Data Analysis
[0739] To analyze the collected visual, auditory, and thought patterns, the terminal uses video analysis algorithms (such as OpenCV), audio analysis algorithms (such as Librosa), and natural language processing algorithms (such as the transformers library). This allows for the extraction of the talent's visual, auditory, and thought data.
[0740] Training of generative AI models
[0741] The server prepares a generative AI model using the analyzed visual, auditory, and thought data. It utilizes GANs (Generative Adversarial Networks) and Transformer models (e.g., GPT-3) as generative AI models. These models are based on deep learning and generate AI talents by supplying collected data as training data.
[0742] AI Talent Generate
[0743] The server uses a trained generative AI model to generate AI talents that replicate the visual information, auditory information, and thought patterns of the talent during their prime.
[0744] Creating advertising content
[0745] The device retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into this template. The advertising template includes storyline and design specifications, and the AI talent generates advertising content based on these.
[0746] User emotion recognition
[0747] When a user views advertising content, the device collects the user's visual and auditory information in real time through the smartphone's camera and microphone. Using this data, the server identifies the user's emotions using sentiment analysis algorithms (such as dlib or OpenCV).
[0748] Real-time adjustment of advertising content
[0749] The server adjusts the AI talent's facial expressions and speech in real time based on the results of sentiment analysis. This process uses GPT-2 and GPT-3 models to generate appropriate prompt sentences, which are then reflected in the advertising content.
[0750] Storage and management of AI talent data
[0751] The generated AI talent data is stored as an NFT (Non-Fungible Token) and registered using blockchain technology (such as Ethereum or Hyperledger). This generates a unique digital ID for the AI talent, and ownership and transaction history are recorded on a distributed ledger.
[0752] Specific example
[0753] If a user is watching a video ad on their smartphone, and the camera detects their facial expression and recognizes that they are "surprised," the ad message will automatically adjust to a tone such as "Amazing product!". An example of a prompt message might be, "When the user looks happy, our AI talent will generate the most appropriate ad message."
[0754] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0755] Step 1:
[0756] Collect visual information, auditory information, and thought patterns of the talent.
[0757] The server retrieves past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak. Specifically, it downloads data from cloud-based storage services (e.g., Google Cloud Storage, Amazon S3).
[0758] Step 2:
[0759] The collected visual information, auditory information, and thought patterns are analyzed.
[0760] The device analyzes data using video analysis algorithms (OpenCV), audio analysis algorithms (Librosa), and natural language processing algorithms (transformers library). Visual data, audio data, and thought data are extracted. Specifically, it detects face regions frame by frame from video files, extracts audio features from audio files, and performs semantic analysis on interview recordings.
[0761] Step 3:
[0762] Train a generative AI model.
[0763] The server prepares a generative AI model using the analyzed visual, audio, and thought data. This generative AI model (GAN or Transformer model) is trained based on deep learning. Specifically, it uses a Python deep learning framework (e.g., TensorFlow, PyTorch) to feed the dataset into the training model and optimize the model parameters.
[0764] Step 4:
[0765] Generate AI talent.
[0766] The server uses a trained generative AI model to generate AI talents that recreate the visual, auditory, and thought patterns of the talent during their prime. Specifically, when a user provides input data, the model generates realistic video and audio of the talent based on that data. This generated data is then used in the next step.
[0767] Step 5:
[0768] Create advertising content.
[0769] The device retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into this template. Specifically, the template includes storyline and design specifications, and the AI talent generates advertising content based on these.
[0770] Step 6:
[0771] Collect and analyze user sentiment data.
[0772] The device collects the user's visual and audio information in real time through the smartphone's camera and microphone. The server processes the collected data using emotion analysis algorithms (such as dlib or OpenCV) to detect emotions. Specifically, it identifies the user's emotions by analyzing facial expressions from the camera image and voice tone from the audio.
[0773] Step 7:
[0774] Adjust advertising content in real time.
[0775] Based on the results of sentiment analysis, the server adjusts the facial expressions and dialogue of the AI talent in the generated advertising content in real time. Specifically, it takes specific prompt text as input to the generating AI model and reflects the resulting text and video in the advertisement, thereby customizing it according to the user's emotions.
[0776] Step 8:
[0777] Store and manage AI talent data.
[0778] The server stores the generated AI talent data as an NFT and registers it using blockchain technology (e.g., Ethereum, Hyperledger). This generates a unique digital ID for the AI talent, and ownership and transaction history are recorded on a distributed ledger. Specifically, it sends transactions to the blockchain network, ensuring that the data is permanently recorded on the blockchain.
[0779] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0780] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0781] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0782] [Third Embodiment]
[0783] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0784] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0785] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0786] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0787] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0788] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0789] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0790] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0791] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0792] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0793] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0794] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0795] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak period, with the aim of establishing advertising strategies and brand image. This system is implemented as follows.
[0796] 1. Data collection:
[0797] The terminal collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak period.
[0798] 2. Data Analysis:
[0799] The device extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm.
[0800] 3. Preparing the Generative AI Model:
[0801] The server prepares a generative AI model based on the analyzed data. The generative AI model is trained using deep learning to recreate the celebrity's appearance during their prime, and processes visual information, auditory information, and thought patterns as input data.
[0802] 4. AI Talent Generation:
[0803] The server uses a prepared generative AI model to generate realistic AI talents. These AI talents represent the real-life talents in their prime and possess high fidelity in terms of visual, auditory, and cognitive aspects.
[0804] 5. Ad generation:
[0805] The terminal retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into the advertising template. This generates a finished advertisement with visual consistency and audio adjustments. For example, the generated AI talent can be used to create new commercial content and promote products.
[0806] 6. NFT creation and blockchain registration:
[0807] The server stores the generated AI talent data as an NFT. This NFT is assigned a unique digital ID and registered on the blockchain. As a result, the generated AI talent data is recognized as a new digital asset and becomes tradable on the market.
[0808] Specific example
[0809] For example, consider a case where past data of a famous celebrity A is collected and analyzed, and then an AI celebrity is generated based on that data.
[0810] 1. The device collects past video files, audio files, and interview records of talent A.
[0811] 2. The device analyzes collected video files to extract visual data, and analyzes audio files to extract audio data. Furthermore, it analyzes interview recordings to extract thought data.
[0812] 3. The server trains an AI model based on this data to generate an AI talent that represents Talent A in their prime.
[0813] 4. The server incorporates the generated AI talent into an advertising template provided by the company to produce the final advertisement. This advertisement features an AI talent that represents Talent A in their prime.
[0814] 5. The server stores the generated AI talent data as an NFT and registers it on the blockchain. This NFT can be traded on the market as a digital asset.
[0815] In this way, the present invention uses AI to recreate a talent's peak period, enabling the maintenance of advertising strategies and brand image, and the creation of new business models.
[0816] The following describes the processing flow.
[0817] Program processing flow: Detailed step-by-step explanation
[0818] Step 1: Data Collection
[0819] The terminal collects past video files, audio files, and interview recordings of the talent from a database. This includes searching for and retrieving all related media files using a specified talent ID.
[0820] Step 2: Analysis of the month
[0821] The device processes the collected video files using a video analysis algorithm to extract visual data. This analysis captures the talent's facial features, movements, and expressions in high definition.
[0822] The device extracts audio data from the collected audio files by running them through an audio analysis algorithm. This analysis examines the tone of voice, intonation, and speaking style characteristics.
[0823] The device processes the collected interview records using a natural language processing algorithm to extract thought data. This analysis examines the content of the talent's statements and logical patterns.
[0824] Step 3: Preparing the Generative AI Model
[0825] The server prepares a generative AI model using the analyzed visual, audio, and thought data. Specifically, it initializes a deep learning model (e.g., a GAN or Transformer model) and supplies the collected data as training data.
[0826] Step 4: Generating AI Talent
[0827] The server generates AI talent using a pre-trained generative AI model. This generation process creates realistic AI talent based on input visual data, audio data, and thought data.
[0828] Step 5: Generate Ads
[0829] The device retrieves an advertising template provided by the company. This template includes storyline and design specifications.
[0830] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency.
[0831] Step 6: Review and correct the generated ads
[0832] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[0833] Step 7: NFT conversion
[0834] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[0835] Step 8: Register on the blockchain
[0836] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[0837] This allows for the recreation of a talent's peak period using AI, ensuring the sustainability of advertising strategies and enabling the utilization of talent value as a new digital asset.
[0838] (Example 1)
[0839] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0840] Currently, leveraging the influence of popular celebrities is common for advertising strategies and brand image building. However, once a celebrity's peak is over, or due to their constantly busy schedule, it often becomes difficult to secure their appearances in advertisements. To solve this problem, a system is needed that can recreate a celebrity's peak performance based on visual, auditory, and thought patterns, and effectively utilize this in advertising.
[0841] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0842] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for storing the generated AI talent data as an NFT and registering it on a blockchain; and means for obtaining an advertising template and generating an advertisement with visual and audio consistency. This makes it possible to recreate the talent's appearance during their prime while establishing an appropriate advertising strategy and brand image.
[0843] A "talent" is a famous or well-known person who is active in the media or advertising.
[0844] "Visual information" refers to digital data such as videos and images that pertain to the appearance and actions of a talent.
[0845] "Audio information" refers to digital audio data related to the voice and speaking style of a talent.
[0846] "Thinking patterns" refer to data about a celebrity's tendencies in thinking and emotions, extracted from their statements, interview content, and other sources.
[0847] "Means of collection" refers to the methods and technologies used to obtain necessary data from databases, the internet, and other sources.
[0848] "Means of analysis" refers to methods of analyzing collected data using technologies such as video analysis algorithms, audio analysis algorithms, and natural language processing algorithms.
[0849] A "generative AI model" is an artificial intelligence model that can reproduce the characteristics of a talent with high accuracy, based on their visual information, auditory information, and thought patterns during their peak.
[0850] An "AI talent" is a virtual talent created using a generative AI model.
[0851] An "ad template" refers to a basic format or design layout used to create advertising content.
[0852] "NFT" stands for Non-Fungible Token, and it refers to a unique token that uses blockchain technology to prove ownership of a digital item.
[0853] Blockchain is a technology that securely manages and trades digital data on a decentralized network.
[0854] Modes for carrying out the invention
[0855] This invention is a system that analyzes the visual information, auditory information, and thought patterns of a talent during their peak period and reproduces them using a generative AI, and is used to improve advertising strategies and establish brand image. Specific embodiments are described in detail below.
[0856] Program Description
[0857] 1. Data collection:
[0858] The terminal connects to a database and collects past video files, audio files, and interview recordings of the talent. This uses automated scripts to download necessary data from sources such as YouTube and television station archives.
[0859] 2. Data Analysis:
[0860] The device uses OpenCV to analyze video files and extract facial features and movements as visual data. For example, it can analyze the microexpressions and body movements of a celebrity in detail.
[0861] The device uses LibROSA to analyze audio files and extract audio features (pitch, tone, volume, etc.). For example, it can analyze the intonation and tone of a specific voice of a talent.
[0862] The device uses Spacy and BERT to analyze interview recordings and extract the talent's thought patterns and frequently occurring words as natural language data. For example, it analyzes characteristics such as "frequently using positive words" and "having a tendency to respond to stress."
[0863] 3. Preparing the Generative AI Model:
[0864] The server integrates the analyzed visual, audio, and thought data to train a generative AI model. This model utilizes deep learning frameworks such as TensorFlow and PyTorch. For example, a large dataset containing tens of thousands of image and audio data is used to recreate a celebrity's appearance during their prime.
[0865] 4. AI Talent Generation:
[0866] The server uses a trained generative AI model to generate an AI talent that looks like the real talent in their prime. This AI talent has a very realistic appearance and voice based on actual video and audio data. For example, it can generate a scene in which the talent introduces a specific product.
[0867] 5. Ad generation:
[0868] The device retrieves advertising templates provided by companies, and the server integrates AI talent into these templates. This process utilizes video editing software such as After Effects or Premiere Pro. For example, the AI talent could create a commercial explaining the features of a new smartphone.
[0869] 6. NFT creation and blockchain registration:
[0870] The server stores the generated AI talent data as an NFT using platforms such as OpenSea and Rarible, and registers it on a blockchain such as Ethereum. This makes the AI talent tradable on the market as a unique digital asset. For example, the AI talent could be transformed into an NFT as a signed digital poster.
[0871] Specific example
[0872] For example, consider a case where past data of famous celebrities is collected and analyzed, and then AI celebrities are generated based on that data. Specifically, the following procedure is used:
[0873] 1. The device collects past video files, audio files, and interview recordings of famous celebrities.
[0874] 2. The device analyzes this data and extracts visual data, audio data, and thought data.
[0875] 3. The server uses this data to train an AI model and generate an AI talent that represents the talent in their prime.
[0876] 4. The server incorporates the generated AI talent into the ad template and generates the finished ad. This ad features the AI talent in their prime.
[0877] 5. The server stores the generated AI talent data as an NFT and registers it on the blockchain. This NFT can be traded on the market as a digital asset.
[0878] Example of a prompt
[0879] "Train an AI model to recreate the peak appearance of famous celebrities based on their past data. Provide advertising templates using the generated AI celebrity and create completed advertisements with visual and audio consistency. Also, save the data of this AI celebrity as an NFT and register it on the blockchain."
[0880] The above describes specific embodiments for carrying out the present invention.
[0881] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0882] Step 1:
[0883] Data collection
[0884] The terminal connects to a database to search for and download past video files, audio files, and interview recordings of a specified talent. Input is information such as the talent's name and the URL of the data source, and output is a set of downloaded video files, audio files, and interview recordings. For example, it can run a script that automatically collects past appearance videos using the YouTube API.
[0885] Step 2:
[0886] Analysis of video data
[0887] The device analyzes video files collected using OpenCV frame by frame. It extracts facial features and movements and saves them as visual data. The input is a video file, and the output is visual data including facial features and movements. Specifically, it uses a face detection algorithm to extract facial contours, microexpressions, and body movements.
[0888] Step 3:
[0889] Analysis of audio data
[0890] The device uses LibROSA to analyze collected audio files and extract audio features (pitch, tone, volume, etc.). The input is an audio file, and the output is audio feature data. Specifically, it converts the audio file into waveform data and analyzes the pitch range and intonation.
[0891] Step 4:
[0892] Analysis of interview data
[0893] The device uses Spacy and BERT to process interview records collected using natural language processing, extracting the talent's thought patterns and frequently occurring words. The input is the interview records, and the output is a list of thought patterns and frequently occurring words. Specifically, the text data is tokenized, and sentiment analysis and theme extraction are performed.
[0894] Step 5:
[0895] Preparing a Generative AI Model
[0896] The server integrates the visual, auditory, and thought data obtained in steps 2 through 4 to train a generative AI model. Deep learning frameworks such as TensorFlow and PyTorch are used. The input is visual, auditory, and thought data, and the output is the generative AI model. Specifically, a deep neural network is designed and trained on these datasets.
[0897] Step 6:
[0898] AI Talent Generate
[0899] The server uses a trained generative AI model to generate an AI talent that represents the talent in their prime. The input is the generative AI model, and the output is the generated AI talent. Specifically, it generates realistic video and audio from the input data to enhance the accuracy of the talent's representation.
[0900] Step 7:
[0901] Ad generation
[0902] The terminal retrieves advertising templates provided by companies, and the server integrates the generated AI talent into the advertising templates. The input is the advertising template and the generated AI talent, and the output is the completed advertising video. Specifically, video editing software such as After Effects or Premiere Pro is used to create a commercial in which the AI talent introduces a new product.
[0903] Step 8:
[0904] NFT creation and blockchain registration
[0905] The server stores the generated AI talent data as an NFT and registers it on a blockchain such as Ethereum. The input is the generated AI talent data, and the output is the NFT and its registration information on the blockchain. Specifically, it uses platforms such as OpenSea and Rarible to create the digital assets of the AI talent and records them on the blockchain.
[0906] Through the steps described above, this system can recreate a talent's peak period with high fidelity, contributing to advertising strategies and the establishment of a brand image.
[0907] (Application Example 1)
[0908] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0909] In today's advertising industry, there is a growing demand to effectively leverage the image of a celebrity during their peak to promote products and services. However, integrating the latest advertising content while recreating the video, audio, and thought patterns of a celebrity from their prime is difficult. Furthermore, there is a lack of mechanisms to manage this generated data and preserve its value as a digital asset. Therefore, there is a need for an advertising generation system that utilizes celebrity data.
[0910] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0911] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content automatically generated based on user input; and means for storing the generated AI talent data as an NFT and registering it on the blockchain. This makes it possible to automatically generate highly accurate advertising content using data from the talent's peak period and guarantee its asset value.
[0912] "Visual information" refers to video data related to the talent's appearance, movements, posture, facial expressions, etc.
[0913] "Audio information" refers to audio data related to a talent's voice quality, pronunciation, intonation, speaking style, etc.
[0914] "Thinking patterns" refer to tendencies in a celebrity's thinking, opinions, and ideas, based on their past interviews and statements.
[0915] A "generative AI model" refers to an artificial intelligence model that generates AI talent based on collected and analyzed visual information, audio information, and thought patterns.
[0916] "AI talent" refers to a virtual character created based on collected and analyzed data, possessing the appearance, voice, and thought patterns of a talent during their prime.
[0917] "Advertising content" refers to digital media, including videos and audio, used to promote products and services.
[0918] An "ad template" refers to a format or design template used when generating advertising content.
[0919] "NFT" stands for Non-Fungible Token, and refers to a token used to prove ownership or uniqueness of a digital asset.
[0920] "Blockchain" refers to a database technology that manages digital data in a decentralized manner and is resistant to tampering.
[0921] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak period, with the aim of establishing advertising strategies and brand image. This system is implemented as follows.
[0922] 1. Data collection:
[0923] Users collect past video files, audio files, and interview records of talents from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak. To do this, they extract the necessary data from the database using programming languages such as Python.
[0924] 2. Data Analysis:
[0925] The server extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm. Specifically, it uses PIL (Python Imaging Library) for video analysis, the TextToSpeech library for speech analysis, and spaCy and transformers for natural language processing.
[0926] 3. Preparing the Generative AI Model:
[0927] The server prepares a generative AI model based on the analyzed data. This generative AI model is trained using deep learning to recreate the talent's peak performance, processing visual information, audio information, and thought patterns as input data. Frameworks such as TensorFlow and PyTorch are used to train the deep learning model.
[0928] 4. AI Talent Generation:
[0929] The server generates realistic AI talents using a prepared generative AI model. These AI talents represent real-life talents in their prime and possess high fidelity in terms of visual, auditory, and cognitive aspects. The generated AI talents are then used in advertising content for products and services.
[0930] 5. Ad generation:
[0931] The server retrieves the ad template provided by the user and incorporates the generated AI talent into the ad template. It then adjusts the visual consistency and audio to generate the finished ad. The generated ad is converted into a video file as visual data and an audio file as audio data.
[0932] 6. NFT creation and blockchain registration:
[0933] The server stores the generated AI talent data as NFTs. Each NFT is assigned a unique digital ID and registered using blockchain technology. This process makes the generated AI talent data tradable on the market as a new digital asset.
[0934] Specific example
[0935] For example, suppose a company wants to create a commercial for a new product launch that uses the image of a famous celebrity in their prime. The user collects past data on the celebrity and inputs it into the system. Next, the server generates an AI celebrity based on the analyzed data. Then, the AI celebrity is incorporated into a provided advertising template to generate the completed advertisement. The generated advertisement is used to promote the sale of the product or service.
[0936] Example of a prompt
[0937] "Based on data from Talent A's heyday, recreate the product's visual characteristics, voice, and thought patterns to create a 30-second advertisement promoting the new BeeWatch product. The advertisement features Talent A discussing the new product's features and explaining its ease of use. The background should feature a city nightscape and a luxurious set."
[0938] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0939] Step 1:
[0940] The user inputs prompt text and advertising content details to recreate the talent's heyday. This provides the system with specific target talent information and advertising direction.
[0941] Input: Talent identification information, advertising content details, prompt text
[0942] Output: List of data to be collected
[0943] Specific operation: The user enters information into an input form via a smartphone application and sends it to the server.
[0944] Step 2:
[0945] The device collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information to recreate the talent's heyday.
[0946] Input: List of data to be collected
[0947] Output: Collected video files, audio files, interview recordings
[0948] Specific operation: The device accesses the database and downloads the necessary video, audio, and interview recordings.
[0949] Step 3:
[0950] The server extracts visual data from collected video files by applying a video analysis algorithm, and extracts audio data from audio files by analyzing them with an audio analysis algorithm. Similarly, it extracts thought data from interview recordings by applying a natural language processing algorithm.
[0951] Input: Collected video files, audio files, interview recordings
[0952] Output: Visual data, audio data, thought data
[0953] Specific operation: The server analyzes video data using PIL (Python Imaging Library), analyzes audio data using the TextToSpeech library, and analyzes interview recordings using spaCy and transformers.
[0954] Step 4:
[0955] The server trains a generative AI model based on analyzed visual, audio, and thought data. This generative AI model is used to recreate the talent's peak performance. TensorFlow and PyTorch are used to train the deep learning model.
[0956] Input: Visual data, audio data, thought data
[0957] Output: Trained generative AI model
[0958] Specific operation: The server uses a deep learning framework to train an AI model and optimize the model parameters.
[0959] Step 5:
[0960] The server uses a trained generative AI model to generate AI talents that represent the talents in their prime. The AI talents are generated as virtual characters with high visual, auditory, and cognitive accuracy.
[0961] Input: Trained generative AI model
[0962] Output: Generated AI talents
[0963] Specific operation: The server runs the generated AI model to produce 3D models and voice samples that recreate the appearance of the talent.
[0964] Step 6:
[0965] The server retrieves the ad template provided by the user and incorporates the generated AI talent into the ad template. This results in the creation of a finished ad with visual consistency and audio adjustments.
[0966] Input: Generated AI talent, ad template
[0967] Output: Completed advertising content
[0968] Specific operation: The server inserts AI talent data into the ad template and uses a video editing library (e.g., moviepy) to create the final ad content.
[0969] Step 7:
[0970] The server stores the generated AI talent data as an NFT, assigns a unique digital ID to it, and registers it on the blockchain. This allows the generated AI talent data to be recognized as a digital asset and traded on the market.
[0971] Input: Generated AI talent data
[0972] Output: AI talent data as NFTs, digital ID registered on the blockchain
[0973] Specific operation: The server uses an NFT generation tool to assign a digital ID and register the data on the blockchain network.
[0974] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0975] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak, aiming to solidify advertising strategies and brand image. Furthermore, by combining it with an emotion engine that recognizes user emotions, the effectiveness of advertising content is enhanced.
[0976] 1. Data collection:
[0977] The terminal collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak period.
[0978] 2. Data Analysis:
[0979] The device extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm.
[0980] 3. Preparing the Generative AI Model:
[0981] The server prepares a generative AI model using the analyzed visual, audio, and thought data. The generative AI model initializes a deep learning model (e.g., GAN or Transformer model) and supplies the collected data as training data.
[0982] 4. AI Talent Generation:
[0983] The server generates AI talent using a pre-trained generative AI model. This generation process creates realistic AI talent based on input visual data, audio data, and thought data.
[0984] 5. Ad generation:
[0985] The device retrieves an advertising template provided by the company. This template includes storyline and design specifications.
[0986] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency.
[0987] 6. Review and correct the generated ads:
[0988] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[0989] 7. NFT creation and blockchain registration:
[0990] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[0991] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[0992] Embedding an emotion engine
[0993] Furthermore, by incorporating an emotion engine that recognizes user emotions, advertising content will be personalized.
[0994] 1. Collecting user sentiment data:
[0995] The device collects the user's facial expressions and tone of voice in real time from visual and auditory information. This captures the user's emotional state while they are watching the advertisement.
[0996] 2. Emotion analysis:
[0997] The server processes the collected visual and audio information of the user through an emotion analysis algorithm to detect the user's emotions. This analysis identifies emotions such as happiness, sadness, surprise, and excitement.
[0998] 3. Adjusting advertising content:
[0999] Based on the results of emotion analysis, the server adjusts the facial expressions and dialogue of the generated AI talent in real time. For example, if the user is excited, the server enhances the energetic expressions of the AI talent.
[1000] Specific example
[1001] For example, if you collect data on talent A during their peak, analyze it to generate an AI talent, and then use an emotion engine to personalize advertising content:
[1002] 1. The device collects past video files, audio files, and interview records of talent A.
[1003] 2. The device analyzes the collected data and extracts visual data, audio data, and thought data.
[1004] 3. The server trains a generation AI model to create an AI talent that resembles Talent A in his prime.
[1005] 4. The terminal retrieves the advertising template provided by the company, and the server incorporates the AI talent into the advertising template.
[1006] 5. The user (company representative) reviews the generated advertising content and makes corrections as needed.
[1007] 6. The server converts the generated AI talent data into an NFT and registers it on the blockchain.
[1008] 7. The terminal collects the user's visual and auditory information, and the server detects the user's emotions using an emotion analysis algorithm.
[1009] 8. Based on the results of the emotion analysis, the server adjusts the facial expressions and speech content of the AI talent to provide personalized advertising content.
[1010] This invention enables the reproduction of a talent's peak period using AI generation, ensuring the sustainability of advertising strategies and providing highly accurate advertising content that takes user emotions into consideration.
[1011] The following describes the processing flow.
[1012] Program processing flow: Detailed step-by-step explanation
[1013] Step 1: Data Collection
[1014] The terminal collects past video files, audio files, and interview recordings of the talent from a database. This includes searching for and retrieving all related media files using a specified talent ID.
[1015] Step 2: Visual Data Analysis
[1016] The device processes the collected video files using a video analysis algorithm to extract visual data. This analysis captures the talent's facial features, movements, and expressions in high definition. Specifically, it uses facial recognition and motion analysis technologies.
[1017] Step 3: Analyzing the audio data
[1018] The device extracts audio data from collected audio files by running them through an audio analysis algorithm. This analysis examines voice tone, intonation, and speaking style characteristics. For example, it may use speech recognition technology or emotion analysis technology.
[1019] Step 4: Analysis of thought data
[1020] The device processes the collected interview records using natural language processing algorithms to extract thought data. This analysis examines the content and logical patterns of the talent's statements. For example, it employs text mining and semantic analysis techniques.
[1021] Step 5: Preparing the Generative AI Model
[1022] The server prepares generative AI models using the analyzed visual, audio, and thought data. Specifically, it initializes deep learning models (such as GANs and Transformer models) and supplies the collected data as training data.
[1023] Step 6: Generating AI Talents
[1024] The server generates AI talents using a pre-trained generative AI model. The generation process creates realistic AI talents based on input visual, audio, and thought data. For example, it utilizes technologies such as synthetic facial image generation and speech synthesis.
[1025] Step 7: Obtain an ad template
[1026] The device retrieves an ad template provided by the company. This template includes the ad's storyline and design specifications.
[1027] Step 8: Generate Ads
[1028] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency. For example, it might use composite video editing technology.
[1029] Step 9: Review and correct the generated ads
[1030] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[1031] Step 10: NFT conversion
[1032] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[1033] Step 11: Register on the blockchain
[1034] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[1035] Step 12: Collecting user sentiment data
[1036] The device collects the user's facial expressions and tone of voice in real time from visual and auditory information. Specifically, it uses a camera and microphone to capture the user's reactions.
[1037] Step 13: Emotion Analysis
[1038] The server uses an emotion analysis algorithm to process the collected visual and audio information of the user to detect their emotions. This analysis identifies emotions such as happiness, sadness, surprise, and excitement.
[1039] Step 14: Adjusting ad content
[1040] Based on the results of emotion analysis, the server adjusts the facial expressions and dialogue of the generated AI talent in real time. For example, if the user is excited, the server enhances the energetic expressions of the AI talent.
[1041] Specific example
[1042] For example, consider a case where past data of a famous celebrity A is collected and analyzed to generate an AI talent, while simultaneously analyzing user emotions to provide appropriate advertising content:
[1043] 1. The device collects past video files, audio files, and interview records of talent A.
[1044] 2. The device analyzes collected video files to extract visual data, and analyzes audio files to extract audio data. Furthermore, it analyzes interview recordings to extract thought data.
[1045] 3. The server trains a generation AI model based on this data and generates an AI talent that resembles Talent A in his prime.
[1046] 4. The terminal retrieves the advertising template provided by the company, and the server incorporates the generated AI talent into the template.
[1047] 5. The user (company representative) reviews the generated advertising content and makes corrections as necessary.
[1048] 6. The server converts the generated AI talent data into an NFT and registers it on the blockchain.
[1049] 7. The terminal collects the user's visual and auditory information, and the server uses an emotion analysis algorithm to detect the user's emotions.
[1050] 8. The server adjusts the AI talent's facial expressions and dialogue based on the results of the emotion analysis, providing advertising content optimized for the user's emotions.
[1051] This invention enables the reproduction of a talent's peak period using AI generation, ensuring the sustainability of advertising strategies and providing highly accurate advertising content that takes user emotions into consideration.
[1052] (Example 2)
[1053] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1054] Traditional advertising content generation systems struggled to effectively reproduce the visual, auditory, and thought patterns of celebrities during their peak, resulting in limited advertising effectiveness. Furthermore, there was a lack of means to analyze user emotions in real time and provide personalized advertising content accordingly. Additionally, centralized management of ownership and transaction history of generated data was a challenge.
[1055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1056] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generative AI model based on the analyzed data; means for incorporating the AI talent into advertising content; means for storing the generated AI talent data as a non-fungible token and registering it in a distributed ledger; means for collecting and analyzing user emotion data; and means for adjusting the advertising content in real time based on the emotion analysis results. This makes it possible to effectively recreate the talent's peak performance and provide personalized advertising content that responds to the user's emotions. Furthermore, registering the ownership and transaction history of the generated data in a distributed ledger enhances the reliability and transparency of the data.
[1057] "Visual information" refers to visual data extracted from video files, such as the face, body movements, and facial expressions of the talent.
[1058] "Audio information" refers to acoustic data extracted from audio files, such as the characteristics of a talent's voice, tone, and speaking style.
[1059] "Thinking patterns" refer to the way of thinking, the content of what is said, and the choice of words that can be extracted from interview transcripts and speeches of celebrities.
[1060] A "generative AI model" is an artificial intelligence model used to generate AI talents based on data collected and analyzed using deep learning technology.
[1061] An "AI talent" is a virtual person created by a generative AI model that reproduces the visual information, auditory information, and thought patterns of a real talent.
[1062] "Advertising content" refers to media such as videos, images, audio, and text used for the purpose of promoting or marketing a company.
[1063] A "non-fungible token" is a unique token that uses blockchain technology to prove ownership of digital data.
[1064] A "distributed ledger" is a distributed management system that uses blockchain technology to centrally manage data ownership and transaction history.
[1065] "Emotional data" refers to data about a user's emotional state, collected from their facial expressions and tone of voice.
[1066] "Sentiment analysis" is the process of analyzing collected emotional data to identify the user's emotional state.
[1067] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak, aiming to solidify advertising strategies and brand image. Furthermore, by combining it with an emotion engine that recognizes user emotions, the effectiveness of advertising content is enhanced.
[1068] First, the device collects past video files, audio files, and interview records of the talent from a database. Specifically, videos are downloaded from the internet and saved to the local disk. Audio files are retrieved from cloud storage and saved in the same way. Interview records are collected as digital data from scanned documents using OCR technology (e.g., Tesseract).
[1069] Next, the device analyzes the collected video files using a video analysis algorithm (e.g., OpenCV) to extract visual data. Similarly, audio files are analyzed using a speech analysis algorithm (e.g., Librosa) to extract audio data. Interview recordings are analyzed using a natural language processing algorithm (e.g., NLTK) to extract thought data.
[1070] Subsequently, the server prepares a generative AI model (e.g., a GAN or Transformer model) using the analyzed visual, audio, and thought data. Specifically, it initializes the model using a deep learning framework (e.g., TensorFlow or PyTorch) and supplies the collected data as training data. The server then adjusts the model's hyperparameters and selects the optimal model.
[1071] The server generates AI talents using a trained generative AI model. This process involves inputting visual, auditory, and thought data to create realistic AI talents. The generated AI talents undergo filtering and enhancement to reproduce natural facial expressions and voice tones.
[1072] The device then retrieves an advertising template provided by the company. This template includes storyline and design specifications. For example, it retrieves a Google Slides template and downloads it from a cloud service (e.g., AWS S3). The server incorporates the generated AI talent into the advertising template and edits the advertising video according to the specified storyline. This process uses video editing software such as Adobe Premiere Pro.
[1073] The user (company representative) reviews the generated advertising content and makes corrections as needed. For example, they play the advertising video in Adobe Premiere Pro and mark the areas that need correction. The user then checks if it matches the brand image and provides feedback.
[1074] The server stores the generated AI talent data as non-fungible tokens (NFTs) and registers them on a distributed ledger. This process involves executing smart contracts to generate a unique digital ID for the AI talent and then creating the NFT based on that ID. Blockchain platforms such as Ethereum and Polygon are used as the distributed ledger.
[1075] Furthermore, by incorporating an emotion engine that recognizes user emotions, advertising content can be personalized. The device collects the user's facial expressions and tone of voice in real time using visual information (webcam) and audio information (microphone). This data is sent to a server, where an emotion analysis algorithm (e.g., DeepFace) analyzes the user's emotions. Based on the analysis results, the server adjusts the AI talent's facial expressions and speech in real time, tailoring the advertising content to best match the user's emotional state. For example, if the user shows a surprised reaction, a corresponding catchphrase is displayed to emphasize the visual impact.
[1076] Example of a prompt
[1077] An example of a prompt message is: "Generate an advertisement based on an interview with talent A during their heyday. Analyze user sentiment data in real time to personalize the ad content."
[1078] The above describes a specific embodiment for carrying out the present invention. This makes it possible to effectively recreate a talent's heyday and provide personalized advertising content that responds to the user's emotions. Furthermore, by registering the ownership and transaction history of the generated data in a distributed ledger, the reliability and transparency of the data can be enhanced.
[1079] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1080] Step 1: Data Collection
[1081] The device collects past video files, audio files, and interview records of the talent. Specifically, it downloads video files from the internet (input) and saves them to the local disk (output). It retrieves audio files from cloud storage (input) and saves them similarly (output). Interview records are collected as digital data from scanned documents using OCR technology (e.g., Tesseract) (input) and saved as text data (output).
[1082] Step 2: Analysis of the month
[1083] The terminal analyzes collected video files using a video analysis algorithm (e.g., OpenCV) (input) and extracts visual data (output). Specifically, it performs face detection frame by frame and extracts features of facial expressions and movements. Similarly, it analyzes audio files using a speech analysis algorithm (e.g., Librosa) (input) and extracts audio data (output). Furthermore, it analyzes interview recordings using a natural language processing algorithm (e.g., NLTK) (input) and extracts thought data (output). This allows the characteristics of the talent during their peak to be obtained as data.
[1084] Step 3: Preparing the Generative AI Model
[1085] The server prepares a generative AI model (e.g., a GAN or Transformer model) using collected and analyzed visual, audio, and thought data (input). Specifically, it initializes the model using a deep learning framework (e.g., TensorFlow or PyTorch) (input) and supplies the analyzed data as training data (output). The server then adjusts the hyperparameters of the model and selects the optimal model (output).
[1086] Step 4: Generating AI Talent
[1087] The server generates AI talent using a trained generative AI model (input). Based on visual data, audio data, and thought data, it creates realistic AI talent (output). During this process, the generated AI is fine-tuned to match the actual video and audio (specifically, this includes filtering and enhancement to naturally reproduce facial expressions and voice tone).
[1088] Step 5: Generate Ads
[1089] The device retrieves an advertising template provided by the company (input). This template includes storyline and design specifications. For example, a Google Slides template is downloaded from a cloud service (e.g., AWS S3) (output). The server incorporates the generated AI talent into the advertising template (input) and edits it according to the specified storyline (output). This process uses video editing software (e.g., Adobe Premiere Pro) to add movement and effects.
[1090] Step 6: Review and correct the generated ads
[1091] The user (company representative) reviews the generated advertising content (input) and makes corrections as needed (output). For example, they might play the advertising video in Adobe Premiere Pro and mark the areas that need correction. The user then checks for consistency with the brand image and provides feedback (specifically, by marking up the corrections on the screen).
[1092] Step 7: NFT creation and blockchain registration
[1093] The server stores the generated AI talent data as non-fungible tokens (NFTs) (input) and registers them on a distributed ledger (output). This process involves executing smart contracts to generate a unique digital ID for the AI talent (input) and then creating an NFT based on that ID (output). Blockchain platforms such as Ethereum and Polygon are used as the distributed ledger (specifically, this involves issuing tokens and writing them to the blockchain).
[1094] Step 8: Integrating the Emotion Engine
[1095] The device uses a webcam and microphone to collect the user's facial expressions and voice tone in real time (input). The server processes the collected data using an emotion analysis algorithm (e.g., DeepFace) to analyze the user's emotions (output). Based on the analysis results, the server adjusts the AI talent's facial expressions and speech in real time (input), and adjusts the advertisement content to best match the user's emotional state (output). For example, if the user shows a surprised reaction, a corresponding catchphrase is displayed to emphasize the visual impact (specific actions include real-time changes in facial expressions and adjustments to voice tone).
[1096] Example of a prompt
[1097] "Generate an advertisement based on an interview with talent A during their heyday. Analyze user sentiment data in real time and personalize the ad content."
[1098] (Application Example 2)
[1099] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1100] While using celebrities and characters is common in current advertising strategies, their effectiveness has certain limitations, and providing advertising content tailored to individual user emotions in real time is a particularly challenging task. Furthermore, there is a need to generate more effective advertisements by leveraging data from celebrities' peak periods. Against this backdrop, there is a need for methods to increase user engagement by providing more personalized advertising content.
[1101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1102] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content; means for recognizing emotions from the user's visual information and audio information and adjusting the AI talent's facial expressions and statements in real time; and means for saving the generated AI talent data as an NFT and registering it on the blockchain. This makes it possible to effectively provide personalized advertising content that responds to the user's emotions.
[1103] A "talent" is a person or character used in advertising and entertainment, primarily based on visual and auditory information.
[1104] "Visual information" refers to information that can be recognized visually, such as videos and image data of celebrities.
[1105] "Audio information" refers to information that can be recognized as sound, such as the voice and statements of a celebrity.
[1106] "Thinking patterns" refer to the tendencies in thinking and statements that a celebrity has shown in past interviews and statements.
[1107] A "generative AI model" is an artificial intelligence model that uses deep learning to analyze data and generate new data. Specifically, this includes GANs (Generative Adversarial Networks) and Transformer models.
[1108] An "AI talent" is a virtual talent generated using a generative AI model, based on collected visual information, audio information, and thought patterns of real talents.
[1109] "Advertising content" refers to media content that conveys advertising messages through visual information, audio information, text information, etc.
[1110] "Emotion recognition" is a technology that analyzes a user's visual and auditory information to determine their emotional state.
[1111] "Real-time adjustment" refers to instantly reflecting the user's current emotional state and changing the content accordingly.
[1112] "NFT" stands for "Non-Fungible Token," which is a non-fungible token that uses blockchain technology to prove ownership of digital assets.
[1113] Blockchain is a technology that uses distributed ledger technology to prevent data tampering and ensure transparency.
[1114] This invention provides personalized advertising content that reflects the user's emotions based on the following procedure. Specific hardware and software configurations are shown for each step.
[1115] Data collection
[1116] To collect visual, auditory, and thought patterns of talent, the server retrieves past video files, audio files, and interview records from a database. The databases and storage used for this purpose include cloud-based storage services such as Google Cloud Storage and Amazon S3.
[1117] Data Analysis
[1118] To analyze the collected visual, auditory, and thought patterns, the terminal uses video analysis algorithms (such as OpenCV), audio analysis algorithms (such as Librosa), and natural language processing algorithms (such as the transformers library). This allows for the extraction of the talent's visual, auditory, and thought data.
[1119] Training of generative AI models
[1120] The server prepares a generative AI model using the analyzed visual, auditory, and thought data. It utilizes GANs (Generative Adversarial Networks) and Transformer models (e.g., GPT-3) as generative AI models. These models are based on deep learning and generate AI talents by supplying collected data as training data.
[1121] AI Talent Generate
[1122] The server uses a trained generative AI model to generate AI talents that replicate the visual information, auditory information, and thought patterns of the talent during their prime.
[1123] Creating advertising content
[1124] The device retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into this template. The advertising template includes storyline and design specifications, and the AI talent generates advertising content based on these.
[1125] User emotion recognition
[1126] When a user views advertising content, the device collects the user's visual and auditory information in real time through the smartphone's camera and microphone. Using this data, the server identifies the user's emotions using sentiment analysis algorithms (such as dlib or OpenCV).
[1127] Real-time adjustment of advertising content
[1128] The server adjusts the AI talent's facial expressions and speech in real time based on the results of sentiment analysis. This process uses GPT-2 and GPT-3 models to generate appropriate prompt sentences, which are then reflected in the advertising content.
[1129] Storage and management of AI talent data
[1130] The generated AI talent data is stored as an NFT (Non-Fungible Token) and registered using blockchain technology (such as Ethereum or Hyperledger). This generates a unique digital ID for the AI talent, and ownership and transaction history are recorded on a distributed ledger.
[1131] Specific example
[1132] If a user is watching a video ad on their smartphone, and the camera detects their facial expression and recognizes that they are "surprised," the ad message will automatically adjust to a tone such as "Amazing product!". An example of a prompt message might be, "When the user looks happy, our AI talent will generate the most appropriate ad message."
[1133] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1134] Step 1:
[1135] Collect visual information, auditory information, and thought patterns of the talent.
[1136] The server retrieves past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak. Specifically, it downloads data from cloud-based storage services (e.g., Google Cloud Storage, Amazon S3).
[1137] Step 2:
[1138] The collected visual information, auditory information, and thought patterns are analyzed.
[1139] The device analyzes data using video analysis algorithms (OpenCV), audio analysis algorithms (Librosa), and natural language processing algorithms (transformers library). Visual data, audio data, and thought data are extracted. Specifically, it detects face regions frame by frame from video files, extracts audio features from audio files, and performs semantic analysis on interview recordings.
[1140] Step 3:
[1141] Train a generative AI model.
[1142] The server prepares a generative AI model using the analyzed visual, audio, and thought data. This generative AI model (GAN or Transformer model) is trained based on deep learning. Specifically, it uses a Python deep learning framework (e.g., TensorFlow, PyTorch) to feed the dataset into the training model and optimize the model parameters.
[1143] Step 4:
[1144] Generate AI talent.
[1145] The server uses a trained generative AI model to generate AI talents that recreate the visual, auditory, and thought patterns of the talent during their prime. Specifically, when a user provides input data, the model generates realistic video and audio of the talent based on that data. This generated data is then used in the next step.
[1146] Step 5:
[1147] Create advertising content.
[1148] The device retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into this template. Specifically, the template includes storyline and design specifications, and the AI talent generates advertising content based on these.
[1149] Step 6:
[1150] Collect and analyze user sentiment data.
[1151] The device collects the user's visual and audio information in real time through the smartphone's camera and microphone. The server processes the collected data using emotion analysis algorithms (such as dlib or OpenCV) to detect emotions. Specifically, it identifies the user's emotions by analyzing facial expressions from the camera image and voice tone from the audio.
[1152] Step 7:
[1153] Adjust advertising content in real time.
[1154] Based on the results of sentiment analysis, the server adjusts the facial expressions and dialogue of the AI talent in the generated advertising content in real time. Specifically, it takes specific prompt text as input to the generating AI model and reflects the resulting text and video in the advertisement, thereby customizing it according to the user's emotions.
[1155] Step 8:
[1156] Store and manage AI talent data.
[1157] The server stores the generated AI talent data as an NFT and registers it using blockchain technology (e.g., Ethereum, Hyperledger). This generates a unique digital ID for the AI talent, and ownership and transaction history are recorded on a distributed ledger. Specifically, it sends transactions to the blockchain network, ensuring that the data is permanently recorded on the blockchain.
[1158] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1159] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1160] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1161] [Fourth Embodiment]
[1162] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1163] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1164] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1165] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1166] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1167] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1168] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1169] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1170] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1171] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1172] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1173] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1174] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1175] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak period, with the aim of establishing advertising strategies and brand image. This system is implemented as follows.
[1176] 1. Data collection:
[1177] The terminal collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak period.
[1178] 2. Data Analysis:
[1179] The device extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm.
[1180] 3. Preparing the Generative AI Model:
[1181] The server prepares a generative AI model based on the analyzed data. The generative AI model is trained using deep learning to recreate the celebrity's appearance during their prime, and processes visual information, auditory information, and thought patterns as input data.
[1182] 4. AI Talent Generation:
[1183] The server uses a prepared generative AI model to generate realistic AI talents. These AI talents represent the real-life talents in their prime and possess high fidelity in terms of visual, auditory, and cognitive aspects.
[1184] 5. Ad generation:
[1185] The terminal retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into the advertising template. This generates a finished advertisement with visual consistency and audio adjustments. For example, the generated AI talent can be used to create new commercial content and promote products.
[1186] 6. NFT creation and blockchain registration:
[1187] The server stores the generated AI talent data as an NFT. This NFT is assigned a unique digital ID and registered on the blockchain. As a result, the generated AI talent data is recognized as a new digital asset and becomes tradable on the market.
[1188] Specific example
[1189] For example, consider a case where past data of a famous celebrity A is collected and analyzed, and then an AI celebrity is generated based on that data.
[1190] 1. The device collects past video files, audio files, and interview records of talent A.
[1191] 2. The device analyzes collected video files to extract visual data, and analyzes audio files to extract audio data. Furthermore, it analyzes interview recordings to extract thought data.
[1192] 3. The server trains an AI model based on this data to generate an AI talent that represents Talent A in their prime.
[1193] 4. The server incorporates the generated AI talent into an advertising template provided by the company to produce the final advertisement. This advertisement features an AI talent that represents Talent A in their prime.
[1194] 5. The server stores the generated AI talent data as an NFT and registers it on the blockchain. This NFT can be traded on the market as a digital asset.
[1195] In this way, the present invention uses AI to recreate a talent's peak period, enabling the maintenance of advertising strategies and brand image, and the creation of new business models.
[1196] The following describes the processing flow.
[1197] Program processing flow: Detailed step-by-step explanation
[1198] Step 1: Data Collection
[1199] The terminal collects past video files, audio files, and interview recordings of the talent from a database. This includes searching for and retrieving all related media files using a specified talent ID.
[1200] Step 2: Analysis of the month
[1201] The device processes the collected video files using a video analysis algorithm to extract visual data. This analysis captures the talent's facial features, movements, and expressions in high definition.
[1202] The device extracts audio data from the collected audio files by running them through an audio analysis algorithm. This analysis examines the tone of voice, intonation, and speaking style characteristics.
[1203] The device processes the collected interview records using a natural language processing algorithm to extract thought data. This analysis examines the content of the talent's statements and logical patterns.
[1204] Step 3: Preparing the Generative AI Model
[1205] The server prepares a generative AI model using the analyzed visual, audio, and thought data. Specifically, it initializes a deep learning model (e.g., a GAN or Transformer model) and supplies the collected data as training data.
[1206] Step 4: Generating AI Talent
[1207] The server generates AI talent using a pre-trained generative AI model. This generation process creates realistic AI talent based on input visual data, audio data, and thought data.
[1208] Step 5: Generate Ads
[1209] The device retrieves an advertising template provided by the company. This template includes storyline and design specifications.
[1210] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency.
[1211] Step 6: Review and correct the generated ads
[1212] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[1213] Step 7: NFT conversion
[1214] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[1215] Step 8: Register on the blockchain
[1216] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[1217] This allows for the recreation of a talent's peak period using AI, ensuring the sustainability of advertising strategies and enabling the utilization of talent value as a new digital asset.
[1218] (Example 1)
[1219] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1220] Currently, leveraging the influence of popular celebrities is common for advertising strategies and brand image building. However, once a celebrity's peak is over, or due to their constantly busy schedule, it often becomes difficult to secure their appearances in advertisements. To solve this problem, a system is needed that can recreate a celebrity's peak performance based on visual, auditory, and thought patterns, and effectively utilize this in advertising.
[1221] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1222] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for storing the generated AI talent data as an NFT and registering it on a blockchain; and means for obtaining an advertising template and generating an advertisement with visual and audio consistency. This makes it possible to recreate the talent's appearance during their prime while establishing an appropriate advertising strategy and brand image.
[1223] A "talent" is a famous or well-known person who is active in the media or advertising.
[1224] "Visual information" refers to digital data such as videos and images that pertain to the appearance and actions of a talent.
[1225] "Audio information" refers to digital audio data related to the voice and speaking style of a talent.
[1226] "Thinking patterns" refer to data about a celebrity's tendencies in thinking and emotions, extracted from their statements, interview content, and other sources.
[1227] "Means of collection" refers to the methods and technologies used to obtain necessary data from databases, the internet, and other sources.
[1228] "Means of analysis" refers to methods of analyzing collected data using technologies such as video analysis algorithms, audio analysis algorithms, and natural language processing algorithms.
[1229] A "generative AI model" is an artificial intelligence model that can reproduce the characteristics of a talent with high accuracy, based on their visual information, auditory information, and thought patterns during their peak.
[1230] An "AI talent" is a virtual talent created using a generative AI model.
[1231] An "ad template" refers to a basic format or design layout used to create advertising content.
[1232] "NFT" stands for Non-Fungible Token, and it refers to a unique token that uses blockchain technology to prove ownership of a digital item.
[1233] Blockchain is a technology that securely manages and trades digital data on a decentralized network.
[1234] Modes for carrying out the invention
[1235] This invention is a system that analyzes the visual information, auditory information, and thought patterns of a talent during their peak period and reproduces them using a generative AI, and is used to improve advertising strategies and establish brand image. Specific embodiments are described in detail below.
[1236] Program Description
[1237] 1. Data collection:
[1238] The terminal connects to a database and collects past video files, audio files, and interview recordings of the talent. This uses automated scripts to download necessary data from sources such as YouTube and television station archives.
[1239] 2. Data Analysis:
[1240] The device uses OpenCV to analyze video files and extract facial features and movements as visual data. For example, it can analyze the microexpressions and body movements of a celebrity in detail.
[1241] The device uses LibROSA to analyze audio files and extract audio features (pitch, tone, volume, etc.). For example, it can analyze the intonation and tone of a specific voice of a talent.
[1242] The device uses Spacy and BERT to analyze interview recordings and extract the talent's thought patterns and frequently occurring words as natural language data. For example, it analyzes characteristics such as "frequently using positive words" and "having a tendency to respond to stress."
[1243] 3. Preparing the Generative AI Model:
[1244] The server integrates the analyzed visual, audio, and thought data to train a generative AI model. This model utilizes deep learning frameworks such as TensorFlow and PyTorch. For example, a large dataset containing tens of thousands of image and audio data is used to recreate a celebrity's appearance during their prime.
[1245] 4. AI Talent Generation:
[1246] The server uses a trained generative AI model to generate an AI talent that looks like the real talent in their prime. This AI talent has a very realistic appearance and voice based on actual video and audio data. For example, it can generate a scene in which the talent introduces a specific product.
[1247] 5. Ad generation:
[1248] The device retrieves advertising templates provided by companies, and the server integrates AI talent into these templates. This process utilizes video editing software such as After Effects or Premiere Pro. For example, the AI talent could create a commercial explaining the features of a new smartphone.
[1249] 6. NFT creation and blockchain registration:
[1250] The server stores the generated AI talent data as an NFT using platforms such as OpenSea and Rarible, and registers it on a blockchain such as Ethereum. This makes the AI talent tradable on the market as a unique digital asset. For example, the AI talent could be transformed into an NFT as a signed digital poster.
[1251] Specific example
[1252] For example, consider a case where past data of famous celebrities is collected and analyzed, and then AI celebrities are generated based on that data. Specifically, the following procedure is used:
[1253] 1. The device collects past video files, audio files, and interview recordings of famous celebrities.
[1254] 2. The device analyzes this data and extracts visual data, audio data, and thought data.
[1255] 3. The server uses this data to train an AI model and generate an AI talent that represents the talent in their prime.
[1256] 4. The server incorporates the generated AI talent into the ad template and generates the finished ad. This ad features the AI talent in their prime.
[1257] 5. The server stores the generated AI talent data as an NFT and registers it on the blockchain. This NFT can be traded on the market as a digital asset.
[1258] Example of a prompt
[1259] "Train an AI model to recreate the peak appearance of famous celebrities based on their past data. Provide advertising templates using the generated AI celebrity and create completed advertisements with visual and audio consistency. Also, save the data of this AI celebrity as an NFT and register it on the blockchain."
[1260] The above describes specific embodiments for carrying out the present invention.
[1261] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1262] Step 1:
[1263] Data collection
[1264] The terminal connects to a database to search for and download past video files, audio files, and interview recordings of a specified talent. Input is information such as the talent's name and the URL of the data source, and output is a set of downloaded video files, audio files, and interview recordings. For example, it can run a script that automatically collects past appearance videos using the YouTube API.
[1265] Step 2:
[1266] Analysis of video data
[1267] The device analyzes video files collected using OpenCV frame by frame. It extracts facial features and movements and saves them as visual data. The input is a video file, and the output is visual data including facial features and movements. Specifically, it uses a face detection algorithm to extract facial contours, microexpressions, and body movements.
[1268] Step 3:
[1269] Analysis of audio data
[1270] The device uses LibROSA to analyze collected audio files and extract audio features (pitch, tone, volume, etc.). The input is an audio file, and the output is audio feature data. Specifically, it converts the audio file into waveform data and analyzes the pitch range and intonation.
[1271] Step 4:
[1272] Analysis of interview data
[1273] The device uses Spacy and BERT to process interview records collected using natural language processing, extracting the talent's thought patterns and frequently occurring words. The input is the interview records, and the output is a list of thought patterns and frequently occurring words. Specifically, the text data is tokenized, and sentiment analysis and theme extraction are performed.
[1274] Step 5:
[1275] Preparing a Generative AI Model
[1276] The server integrates the visual, auditory, and thought data obtained in steps 2 through 4 to train a generative AI model. Deep learning frameworks such as TensorFlow and PyTorch are used. The input is visual, auditory, and thought data, and the output is the generative AI model. Specifically, a deep neural network is designed and trained on these datasets.
[1277] Step 6:
[1278] AI Talent Generate
[1279] The server uses a trained generative AI model to generate an AI talent that represents the talent in their prime. The input is the generative AI model, and the output is the generated AI talent. Specifically, it generates realistic video and audio from the input data to enhance the accuracy of the talent's representation.
[1280] Step 7:
[1281] Ad generation
[1282] The terminal retrieves advertising templates provided by companies, and the server integrates the generated AI talent into the advertising templates. The input is the advertising template and the generated AI talent, and the output is the completed advertising video. Specifically, video editing software such as After Effects or Premiere Pro is used to create a commercial in which the AI talent introduces a new product.
[1283] Step 8:
[1284] NFT creation and blockchain registration
[1285] The server stores the generated AI talent data as an NFT and registers it on a blockchain such as Ethereum. The input is the generated AI talent data, and the output is the NFT and its registration information on the blockchain. Specifically, it uses platforms such as OpenSea and Rarible to create the digital assets of the AI talent and records them on the blockchain.
[1286] Through the steps described above, this system can recreate a talent's peak period with high fidelity, contributing to advertising strategies and the establishment of a brand image.
[1287] (Application Example 1)
[1288] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1289] In today's advertising industry, there is a growing demand to effectively leverage the image of a celebrity during their peak to promote products and services. However, integrating the latest advertising content while recreating the video, audio, and thought patterns of a celebrity from their prime is difficult. Furthermore, there is a lack of mechanisms to manage this generated data and preserve its value as a digital asset. Therefore, there is a need for an advertising generation system that utilizes celebrity data.
[1290] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1291] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content automatically generated based on user input; and means for storing the generated AI talent data as an NFT and registering it on the blockchain. This makes it possible to automatically generate highly accurate advertising content using data from the talent's peak period and guarantee its asset value.
[1292] "Visual information" refers to video data related to the talent's appearance, movements, posture, facial expressions, etc.
[1293] "Audio information" refers to audio data related to a talent's voice quality, pronunciation, intonation, speaking style, etc.
[1294] "Thinking patterns" refer to tendencies in a celebrity's thinking, opinions, and ideas, based on their past interviews and statements.
[1295] A "generative AI model" refers to an artificial intelligence model that generates AI talent based on collected and analyzed visual information, audio information, and thought patterns.
[1296] "AI talent" refers to a virtual character created based on collected and analyzed data, possessing the appearance, voice, and thought patterns of a talent during their prime.
[1297] "Advertising content" refers to digital media, including videos and audio, used to promote products and services.
[1298] An "ad template" refers to a format or design template used when generating advertising content.
[1299] "NFT" stands for Non-Fungible Token, and refers to a token used to prove ownership or uniqueness of a digital asset.
[1300] "Blockchain" refers to a database technology that manages digital data in a decentralized manner and is resistant to tampering.
[1301] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak period, with the aim of establishing advertising strategies and brand image. This system is implemented as follows.
[1302] 1. Data collection:
[1303] Users collect past video files, audio files, and interview records of talents from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak. To do this, they extract the necessary data from the database using programming languages such as Python.
[1304] 2. Data Analysis:
[1305] The server extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm. Specifically, it uses PIL (Python Imaging Library) for video analysis, the TextToSpeech library for speech analysis, and spaCy and transformers for natural language processing.
[1306] 3. Preparing the Generative AI Model:
[1307] The server prepares a generative AI model based on the analyzed data. This generative AI model is trained using deep learning to recreate the talent's peak performance, processing visual information, audio information, and thought patterns as input data. Frameworks such as TensorFlow and PyTorch are used to train the deep learning model.
[1308] 4. AI Talent Generation:
[1309] The server generates realistic AI talents using a prepared generative AI model. These AI talents represent real-life talents in their prime and possess high fidelity in terms of visual, auditory, and cognitive aspects. The generated AI talents are then used in advertising content for products and services.
[1310] 5. Ad generation:
[1311] The server retrieves the ad template provided by the user and incorporates the generated AI talent into the ad template. It then adjusts the visual consistency and audio to generate the finished ad. The generated ad is converted into a video file as visual data and an audio file as audio data.
[1312] 6. NFT creation and blockchain registration:
[1313] The server stores the generated AI talent data as NFTs. Each NFT is assigned a unique digital ID and registered using blockchain technology. This process makes the generated AI talent data tradable on the market as a new digital asset.
[1314] Specific example
[1315] For example, suppose a company wants to create a commercial for a new product launch that uses the image of a famous celebrity in their prime. The user collects past data on the celebrity and inputs it into the system. Next, the server generates an AI celebrity based on the analyzed data. Then, the AI celebrity is incorporated into a provided advertising template to generate the completed advertisement. The generated advertisement is used to promote the sale of the product or service.
[1316] Example of a prompt
[1317] "Based on data from Talent A's heyday, recreate the product's visual characteristics, voice, and thought patterns to create a 30-second advertisement promoting the new BeeWatch product. The advertisement features Talent A discussing the new product's features and explaining its ease of use. The background should feature a city nightscape and a luxurious set."
[1318] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1319] Step 1:
[1320] The user inputs prompt text and advertising content details to recreate the talent's heyday. This provides the system with specific target talent information and advertising direction.
[1321] Input: Talent identification information, advertising content details, prompt text
[1322] Output: List of data to be collected
[1323] Specific operation: The user enters information into an input form via a smartphone application and sends it to the server.
[1324] Step 2:
[1325] The device collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information to recreate the talent's heyday.
[1326] Input: List of data to be collected
[1327] Output: Collected video files, audio files, interview recordings
[1328] Specific operation: The device accesses the database and downloads the necessary video, audio, and interview recordings.
[1329] Step 3:
[1330] The server extracts visual data from collected video files by applying a video analysis algorithm, and extracts audio data from audio files by analyzing them with an audio analysis algorithm. Similarly, it extracts thought data from interview recordings by applying a natural language processing algorithm.
[1331] Input: Collected video files, audio files, interview recordings
[1332] Output: Visual data, audio data, thought data
[1333] Specific operation: The server analyzes video data using PIL (Python Imaging Library), analyzes audio data using the TextToSpeech library, and analyzes interview recordings using spaCy and transformers.
[1334] Step 4:
[1335] The server trains a generative AI model based on analyzed visual, audio, and thought data. This generative AI model is used to recreate the talent's peak performance. TensorFlow and PyTorch are used to train the deep learning model.
[1336] Input: Visual data, audio data, thought data
[1337] Output: Trained generative AI model
[1338] Specific operation: The server uses a deep learning framework to train an AI model and optimize the model parameters.
[1339] Step 5:
[1340] The server uses a trained generative AI model to generate AI talents that represent the talents in their prime. The AI talents are generated as virtual characters with high visual, auditory, and cognitive accuracy.
[1341] Input: Trained generative AI model
[1342] Output: Generated AI talents
[1343] Specific operation: The server runs the generated AI model to produce 3D models and voice samples that recreate the appearance of the talent.
[1344] Step 6:
[1345] The server retrieves the ad template provided by the user and incorporates the generated AI talent into the ad template. This results in the creation of a finished ad with visual consistency and audio adjustments.
[1346] Input: Generated AI talent, ad template
[1347] Output: Completed advertising content
[1348] Specific operation: The server inserts AI talent data into the ad template and uses a video editing library (e.g., moviepy) to create the final ad content.
[1349] Step 7:
[1350] The server stores the generated AI talent data as an NFT, assigns a unique digital ID to it, and registers it on the blockchain. This allows the generated AI talent data to be recognized as a digital asset and traded on the market.
[1351] Input: Generated AI talent data
[1352] Output: AI talent data as NFTs, digital ID registered on the blockchain
[1353] Specific operation: The server uses an NFT generation tool to assign a digital ID and register the data on the blockchain network.
[1354] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1355] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak, aiming to solidify advertising strategies and brand image. Furthermore, by combining it with an emotion engine that recognizes user emotions, the effectiveness of advertising content is enhanced.
[1356] 1. Data collection:
[1357] The terminal collects past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak period.
[1358] 2. Data Analysis:
[1359] The device extracts visual data from collected video files using a video analysis algorithm, and extracts audio data from audio files using a speech analysis algorithm. Similarly, it extracts thought data from interview recordings using a natural language processing algorithm.
[1360] 3. Preparing the Generative AI Model:
[1361] The server prepares a generative AI model using the analyzed visual, audio, and thought data. The generative AI model initializes a deep learning model (e.g., GAN or Transformer model) and supplies the collected data as training data.
[1362] 4. AI Talent Generation:
[1363] The server generates AI talent using a pre-trained generative AI model. This generation process creates realistic AI talent based on input visual data, audio data, and thought data.
[1364] 5. Ad generation:
[1365] The device retrieves an advertising template provided by the company. This template includes storyline and design specifications.
[1366] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency.
[1367] 6. Review and correct the generated ads:
[1368] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[1369] 7. NFT creation and blockchain registration:
[1370] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[1371] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[1372] Embedding an emotion engine
[1373] Furthermore, by incorporating an emotion engine that recognizes user emotions, advertising content will be personalized.
[1374] 1. Collecting user sentiment data:
[1375] The device collects the user's facial expressions and tone of voice in real time from visual and auditory information. This captures the user's emotional state while they are watching the advertisement.
[1376] 2. Emotion analysis:
[1377] The server processes the collected visual and audio information of the user through an emotion analysis algorithm to detect the user's emotions. This analysis identifies emotions such as happiness, sadness, surprise, and excitement.
[1378] 3. Adjusting advertising content:
[1379] Based on the results of emotion analysis, the server adjusts the facial expressions and dialogue of the generated AI talent in real time. For example, if the user is excited, the server enhances the energetic expressions of the AI talent.
[1380] Specific example
[1381] For example, if you collect data on talent A during their peak, analyze it to generate an AI talent, and then use an emotion engine to personalize advertising content:
[1382] 1. The device collects past video files, audio files, and interview records of talent A.
[1383] 2. The device analyzes the collected data and extracts visual data, audio data, and thought data.
[1384] 3. The server trains a generation AI model to create an AI talent that resembles Talent A in his prime.
[1385] 4. The terminal retrieves the advertising template provided by the company, and the server incorporates the AI talent into the advertising template.
[1386] 5. The user (company representative) reviews the generated advertising content and makes corrections as needed.
[1387] 6. The server converts the generated AI talent data into an NFT and registers it on the blockchain.
[1388] 7. The terminal collects the user's visual and auditory information, and the server detects the user's emotions using an emotion analysis algorithm.
[1389] 8. Based on the results of the emotion analysis, the server adjusts the facial expressions and speech content of the AI talent to provide personalized advertising content.
[1390] This invention enables the reproduction of a talent's peak period using AI generation, ensuring the sustainability of advertising strategies and providing highly accurate advertising content that takes user emotions into consideration.
[1391] The following describes the processing flow.
[1392] Program processing flow: Detailed step-by-step explanation
[1393] Step 1: Data Collection
[1394] The terminal collects past video files, audio files, and interview recordings of the talent from a database. This includes searching for and retrieving all related media files using a specified talent ID.
[1395] Step 2: Visual Data Analysis
[1396] The device processes the collected video files using a video analysis algorithm to extract visual data. This analysis captures the talent's facial features, movements, and expressions in high definition. Specifically, it uses facial recognition and motion analysis technologies.
[1397] Step 3: Analyzing the audio data
[1398] The device extracts audio data from collected audio files by running them through an audio analysis algorithm. This analysis examines voice tone, intonation, and speaking style characteristics. For example, it may use speech recognition technology or emotion analysis technology.
[1399] Step 4: Analysis of thought data
[1400] The device processes the collected interview records using natural language processing algorithms to extract thought data. This analysis examines the content and logical patterns of the talent's statements. For example, it employs text mining and semantic analysis techniques.
[1401] Step 5: Preparing the Generative AI Model
[1402] The server prepares generative AI models using the analyzed visual, audio, and thought data. Specifically, it initializes deep learning models (such as GANs and Transformer models) and supplies the collected data as training data.
[1403] Step 6: Generating AI Talents
[1404] The server generates AI talents using a pre-trained generative AI model. The generation process creates realistic AI talents based on input visual, audio, and thought data. For example, it utilizes technologies such as synthetic facial image generation and speech synthesis.
[1405] Step 7: Obtain an ad template
[1406] The device retrieves an ad template provided by the company. This template includes the ad's storyline and design specifications.
[1407] Step 8: Generate Ads
[1408] The server integrates the generated AI talent into the ad template. This involves placing the AI talent in designated locations within the template and generating the ad while maintaining visual and auditory consistency. For example, it might use composite video editing technology.
[1409] Step 9: Review and correct the generated ads
[1410] The user (company representative) reviews the generated ad content and makes revisions as needed. For example, they check whether the ad's tone and message match the brand image.
[1411] Step 10: NFT conversion
[1412] The server stores the generated AI talent data as an NFT. This process generates a unique digital ID for the AI talent and uses it to create the NFT.
[1413] Step 11: Register on the blockchain
[1414] The server registers the generated NFTs on the blockchain. This registration process records the ownership and transaction history of the NFTs on a distributed ledger, guaranteeing their asset value.
[1415] Step 12: Collecting user sentiment data
[1416] The device collects the user's facial expressions and tone of voice in real time from visual and auditory information. Specifically, it uses a camera and microphone to capture the user's reactions.
[1417] Step 13: Emotion Analysis
[1418] The server uses an emotion analysis algorithm to process the collected visual and audio information of the user to detect their emotions. This analysis identifies emotions such as happiness, sadness, surprise, and excitement.
[1419] Step 14: Adjusting ad content
[1420] Based on the results of emotion analysis, the server adjusts the facial expressions and dialogue of the generated AI talent in real time. For example, if the user is excited, the server enhances the energetic expressions of the AI talent.
[1421] Specific example
[1422] For example, consider a case where past data of a famous celebrity A is collected and analyzed to generate an AI talent, while simultaneously analyzing user emotions to provide appropriate advertising content:
[1423] 1. The device collects past video files, audio files, and interview records of talent A.
[1424] 2. The device analyzes collected video files to extract visual data, and analyzes audio files to extract audio data. Furthermore, it analyzes interview recordings to extract thought data.
[1425] 3. The server trains a generation AI model based on this data and generates an AI talent that resembles Talent A in his prime.
[1426] 4. The terminal retrieves the advertising template provided by the company, and the server incorporates the generated AI talent into the template.
[1427] 5. The user (company representative) reviews the generated advertising content and makes corrections as necessary.
[1428] 6. The server converts the generated AI talent data into an NFT and registers it on the blockchain.
[1429] 7. The terminal collects the user's visual and auditory information, and the server uses an emotion analysis algorithm to detect the user's emotions.
[1430] 8. The server adjusts the AI talent's facial expressions and dialogue based on the results of the emotion analysis, providing advertising content optimized for the user's emotions.
[1431] This invention enables the reproduction of a talent's peak period using AI generation, ensuring the sustainability of advertising strategies and providing highly accurate advertising content that takes user emotions into consideration.
[1432] (Example 2)
[1433] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1434] Traditional advertising content generation systems struggled to effectively reproduce the visual, auditory, and thought patterns of celebrities during their peak, resulting in limited advertising effectiveness. Furthermore, there was a lack of means to analyze user emotions in real time and provide personalized advertising content accordingly. Additionally, centralized management of ownership and transaction history of generated data was a challenge.
[1435] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1436] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generative AI model based on the analyzed data; means for incorporating the AI talent into advertising content; means for storing the generated AI talent data as a non-fungible token and registering it in a distributed ledger; means for collecting and analyzing user emotion data; and means for adjusting the advertising content in real time based on the emotion analysis results. This makes it possible to effectively recreate the talent's peak performance and provide personalized advertising content that responds to the user's emotions. Furthermore, registering the ownership and transaction history of the generated data in a distributed ledger enhances the reliability and transparency of the data.
[1437] "Visual information" refers to visual data extracted from video files, such as the face, body movements, and facial expressions of the talent.
[1438] "Audio information" refers to acoustic data extracted from audio files, such as the characteristics of a talent's voice, tone, and speaking style.
[1439] "Thinking patterns" refer to the way of thinking, the content of what is said, and the choice of words that can be extracted from interview transcripts and speeches of celebrities.
[1440] A "generative AI model" is an artificial intelligence model used to generate AI talents based on data collected and analyzed using deep learning technology.
[1441] An "AI talent" is a virtual person created by a generative AI model that reproduces the visual information, auditory information, and thought patterns of a real talent.
[1442] "Advertising content" refers to media such as videos, images, audio, and text used for the purpose of promoting or marketing a company.
[1443] A "non-fungible token" is a unique token that uses blockchain technology to prove ownership of digital data.
[1444] A "distributed ledger" is a distributed management system that uses blockchain technology to centrally manage data ownership and transaction history.
[1445] "Emotional data" refers to data about a user's emotional state, collected from their facial expressions and tone of voice.
[1446] "Sentiment analysis" is the process of analyzing collected emotional data to identify the user's emotional state.
[1447] This invention is a system that uses AI to reproduce the visual information, audio information, and thought patterns of a talent during their peak, aiming to solidify advertising strategies and brand image. Furthermore, by combining it with an emotion engine that recognizes user emotions, the effectiveness of advertising content is enhanced.
[1448] First, the device collects past video files, audio files, and interview records of the talent from a database. Specifically, videos are downloaded from the internet and saved to the local disk. Audio files are retrieved from cloud storage and saved in the same way. Interview records are collected as digital data from scanned documents using OCR technology (e.g., Tesseract).
[1449] Next, the device analyzes the collected video files using a video analysis algorithm (e.g., OpenCV) to extract visual data. Similarly, audio files are analyzed using a speech analysis algorithm (e.g., Librosa) to extract audio data. Interview recordings are analyzed using a natural language processing algorithm (e.g., NLTK) to extract thought data.
[1450] Subsequently, the server prepares a generative AI model (e.g., a GAN or Transformer model) using the analyzed visual, audio, and thought data. Specifically, it initializes the model using a deep learning framework (e.g., TensorFlow or PyTorch) and supplies the collected data as training data. The server then adjusts the model's hyperparameters and selects the optimal model.
[1451] The server generates AI talents using a trained generative AI model. This process involves inputting visual, auditory, and thought data to create realistic AI talents. The generated AI talents undergo filtering and enhancement to reproduce natural facial expressions and voice tones.
[1452] The device then retrieves an advertising template provided by the company. This template includes storyline and design specifications. For example, it retrieves a Google Slides template and downloads it from a cloud service (e.g., AWS S3). The server incorporates the generated AI talent into the advertising template and edits the advertising video according to the specified storyline. This process uses video editing software such as Adobe Premiere Pro.
[1453] The user (company representative) reviews the generated advertising content and makes corrections as needed. For example, they play the advertising video in Adobe Premiere Pro and mark the areas that need correction. The user then checks if it matches the brand image and provides feedback.
[1454] The server stores the generated AI talent data as non-fungible tokens (NFTs) and registers them on a distributed ledger. This process involves executing smart contracts to generate a unique digital ID for the AI talent and then creating the NFT based on that ID. Blockchain platforms such as Ethereum and Polygon are used as the distributed ledger.
[1455] Furthermore, by incorporating an emotion engine that recognizes user emotions, advertising content can be personalized. The device collects the user's facial expressions and tone of voice in real time using visual information (webcam) and audio information (microphone). This data is sent to a server, where an emotion analysis algorithm (e.g., DeepFace) analyzes the user's emotions. Based on the analysis results, the server adjusts the AI talent's facial expressions and speech in real time, tailoring the advertising content to best match the user's emotional state. For example, if the user shows a surprised reaction, a corresponding catchphrase is displayed to emphasize the visual impact.
[1456] Example of a prompt
[1457] An example of a prompt message is: "Generate an advertisement based on an interview with talent A during their heyday. Analyze user sentiment data in real time to personalize the ad content."
[1458] The above describes a specific embodiment for carrying out the present invention. This makes it possible to effectively recreate a talent's heyday and provide personalized advertising content that responds to the user's emotions. Furthermore, by registering the ownership and transaction history of the generated data in a distributed ledger, the reliability and transparency of the data can be enhanced.
[1459] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1460] Step 1: Data Collection
[1461] The device collects past video files, audio files, and interview records of the talent. Specifically, it downloads video files from the internet (input) and saves them to the local disk (output). It retrieves audio files from cloud storage (input) and saves them similarly (output). Interview records are collected as digital data from scanned documents using OCR technology (e.g., Tesseract) (input) and saved as text data (output).
[1462] Step 2: Analysis of the month
[1463] The terminal analyzes collected video files using a video analysis algorithm (e.g., OpenCV) (input) and extracts visual data (output). Specifically, it performs face detection frame by frame and extracts features of facial expressions and movements. Similarly, it analyzes audio files using a speech analysis algorithm (e.g., Librosa) (input) and extracts audio data (output). Furthermore, it analyzes interview recordings using a natural language processing algorithm (e.g., NLTK) (input) and extracts thought data (output). This allows the characteristics of the talent during their peak to be obtained as data.
[1464] Step 3: Preparing the Generative AI Model
[1465] The server prepares a generative AI model (e.g., a GAN or Transformer model) using collected and analyzed visual, audio, and thought data (input). Specifically, it initializes the model using a deep learning framework (e.g., TensorFlow or PyTorch) (input) and supplies the analyzed data as training data (output). The server then adjusts the hyperparameters of the model and selects the optimal model (output).
[1466] Step 4: Generating AI Talent
[1467] The server generates AI talent using a trained generative AI model (input). Based on visual data, audio data, and thought data, it creates realistic AI talent (output). During this process, the generated AI is fine-tuned to match the actual video and audio (specifically, this includes filtering and enhancement to naturally reproduce facial expressions and voice tone).
[1468] Step 5: Generate Ads
[1469] The device retrieves an advertising template provided by the company (input). This template includes storyline and design specifications. For example, a Google Slides template is downloaded from a cloud service (e.g., AWS S3) (output). The server incorporates the generated AI talent into the advertising template (input) and edits it according to the specified storyline (output). This process uses video editing software (e.g., Adobe Premiere Pro) to add movement and effects.
[1470] Step 6: Review and correct the generated ads
[1471] The user (company representative) reviews the generated advertising content (input) and makes corrections as needed (output). For example, they might play the advertising video in Adobe Premiere Pro and mark the areas that need correction. The user then checks for consistency with the brand image and provides feedback (specifically, by marking up the corrections on the screen).
[1472] Step 7: NFT creation and blockchain registration
[1473] The server stores the generated AI talent data as non-fungible tokens (NFTs) (input) and registers them on a distributed ledger (output). This process involves executing smart contracts to generate a unique digital ID for the AI talent (input) and then creating an NFT based on that ID (output). Blockchain platforms such as Ethereum and Polygon are used as the distributed ledger (specifically, this involves issuing tokens and writing them to the blockchain).
[1474] Step 8: Integrating the Emotion Engine
[1475] The device uses a webcam and microphone to collect the user's facial expressions and voice tone in real time (input). The server processes the collected data using an emotion analysis algorithm (e.g., DeepFace) to analyze the user's emotions (output). Based on the analysis results, the server adjusts the AI talent's facial expressions and speech in real time (input), and adjusts the advertisement content to best match the user's emotional state (output). For example, if the user shows a surprised reaction, a corresponding catchphrase is displayed to emphasize the visual impact (specific actions include real-time changes in facial expressions and adjustments to voice tone).
[1476] Example of a prompt
[1477] "Generate an advertisement based on an interview with talent A during their heyday. Analyze user sentiment data in real time and personalize the ad content."
[1478] (Application Example 2)
[1479] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1480] While using celebrities and characters is common in current advertising strategies, their effectiveness has certain limitations, and providing advertising content tailored to individual user emotions in real time is a particularly challenging task. Furthermore, there is a need to generate more effective advertisements by leveraging data from celebrities' peak periods. Against this backdrop, there is a need for methods to increase user engagement by providing more personalized advertising content.
[1481] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1482] In this invention, the server includes means for collecting visual information, audio information, and thought patterns of a talent; means for analyzing the collected visual information, audio information, and thought patterns; means for generating an AI talent using a generation AI model based on the analyzed data; means for incorporating the AI talent into advertising content; means for recognizing emotions from the user's visual information and audio information and adjusting the AI talent's facial expressions and statements in real time; and means for saving the generated AI talent data as an NFT and registering it on the blockchain. This makes it possible to effectively provide personalized advertising content that responds to the user's emotions.
[1483] A "talent" is a person or character used in advertising and entertainment, primarily based on visual and auditory information.
[1484] "Visual information" refers to information that can be recognized visually, such as videos and image data of celebrities.
[1485] "Audio information" refers to information that can be recognized as sound, such as the voice and statements of a celebrity.
[1486] "Thinking patterns" refer to the tendencies in thinking and statements that a celebrity has shown in past interviews and statements.
[1487] A "generative AI model" is an artificial intelligence model that uses deep learning to analyze data and generate new data. Specifically, this includes GANs (Generative Adversarial Networks) and Transformer models.
[1488] An "AI talent" is a virtual talent generated using a generative AI model, based on collected visual information, audio information, and thought patterns of real talents.
[1489] "Advertising content" refers to media content that conveys advertising messages through visual information, audio information, text information, etc.
[1490] "Emotion recognition" is a technology that analyzes a user's visual and auditory information to determine their emotional state.
[1491] "Real-time adjustment" refers to instantly reflecting the user's current emotional state and changing the content accordingly.
[1492] "NFT" stands for "Non-Fungible Token," which is a non-fungible token that uses blockchain technology to prove ownership of digital assets.
[1493] Blockchain is a technology that uses distributed ledger technology to prevent data tampering and ensure transparency.
[1494] This invention provides personalized advertising content that reflects the user's emotions based on the following procedure. Specific hardware and software configurations are shown for each step.
[1495] Data collection
[1496] To collect visual, auditory, and thought patterns of talent, the server retrieves past video files, audio files, and interview records from a database. The databases and storage used for this purpose include cloud-based storage services such as Google Cloud Storage and Amazon S3.
[1497] Data Analysis
[1498] To analyze the collected visual, auditory, and thought patterns, the terminal uses video analysis algorithms (such as OpenCV), audio analysis algorithms (such as Librosa), and natural language processing algorithms (such as the transformers library). This allows for the extraction of the talent's visual, auditory, and thought data.
[1499] Training of generative AI models
[1500] The server prepares a generative AI model using the analyzed visual, auditory, and thought data. It utilizes GANs (Generative Adversarial Networks) and Transformer models (e.g., GPT-3) as generative AI models. These models are based on deep learning and generate AI talents by supplying collected data as training data.
[1501] AI Talent Generate
[1502] The server uses a trained generative AI model to generate AI talents that replicate the visual information, auditory information, and thought patterns of the talent during their prime.
[1503] Creating advertising content
[1504] The device retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into this template. The advertising template includes storyline and design specifications, and the AI talent generates advertising content based on these.
[1505] User emotion recognition
[1506] When a user views advertising content, the device collects the user's visual and auditory information in real time through the smartphone's camera and microphone. Using this data, the server identifies the user's emotions using sentiment analysis algorithms (such as dlib or OpenCV).
[1507] Real-time adjustment of advertising content
[1508] The server adjusts the AI talent's facial expressions and speech in real time based on the results of sentiment analysis. This process uses GPT-2 and GPT-3 models to generate appropriate prompt sentences, which are then reflected in the advertising content.
[1509] Storage and management of AI talent data
[1510] The generated AI talent data is stored as an NFT (Non-Fungible Token) and registered using blockchain technology (such as Ethereum or Hyperledger). This generates a unique digital ID for the AI talent, and ownership and transaction history are recorded on a distributed ledger.
[1511] Specific example
[1512] If a user is watching a video ad on their smartphone, and the camera detects their facial expression and recognizes that they are "surprised," the ad message will automatically adjust to a tone such as "Amazing product!". An example of a prompt message might be, "When the user looks happy, our AI talent will generate the most appropriate ad message."
[1513] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1514] Step 1:
[1515] Collect visual information, auditory information, and thought patterns of the talent.
[1516] The server retrieves past video files, audio files, and interview records of the talent from a database. This data is used as foundational information for analyzing the talent's visual, auditory, and thought patterns during their peak. Specifically, it downloads data from cloud-based storage services (e.g., Google Cloud Storage, Amazon S3).
[1517] Step 2:
[1518] The collected visual information, auditory information, and thought patterns are analyzed.
[1519] The device analyzes data using video analysis algorithms (OpenCV), audio analysis algorithms (Librosa), and natural language processing algorithms (transformers library). Visual data, audio data, and thought data are extracted. Specifically, it detects face regions frame by frame from video files, extracts audio features from audio files, and performs semantic analysis on interview recordings.
[1520] Step 3:
[1521] Train a generative AI model.
[1522] The server prepares a generative AI model using the analyzed visual, audio, and thought data. This generative AI model (GAN or Transformer model) is trained based on deep learning. Specifically, it uses a Python deep learning framework (e.g., TensorFlow, PyTorch) to feed the dataset into the training model and optimize the model parameters.
[1523] Step 4:
[1524] Generate AI talent.
[1525] The server uses a trained generative AI model to generate AI talents that recreate the visual, auditory, and thought patterns of the talent during their prime. Specifically, when a user provides input data, the model generates realistic video and audio of the talent based on that data. This generated data is then used in the next step.
[1526] Step 5:
[1527] Create advertising content.
[1528] The device retrieves an advertising template provided by the company, and the server incorporates the generated AI talent into this template. Specifically, the template includes storyline and design specifications, and the AI talent generates advertising content based on these.
[1529] Step 6:
[1530] Collect and analyze user sentiment data.
[1531] The device collects the user's visual and audio information in real time through the smartphone's camera and microphone. The server processes the collected data using emotion analysis algorithms (such as dlib or OpenCV) to detect emotions. Specifically, it identifies the user's emotions by analyzing facial expressions from the camera image and voice tone from the audio.
[1532] Step 7:
[1533] Adjust advertising content in real time.
[1534] Based on the results of sentiment analysis, the server adjusts the facial expressions and dialogue of the AI talent in the generated advertising content in real time. Specifically, it takes specific prompt text as input to the generating AI model and reflects the resulting text and video in the advertisement, thereby customizing it according to the user's emotions.
[1535] Step 8:
[1536] Store and manage AI talent data.
[1537] The server stores the generated AI talent data as an NFT and registers it using blockchain technology (e.g., Ethereum, Hyperledger). This generates a unique digital ID for the AI talent, and ownership and transaction history are recorded on a distributed ledger. Specifically, it sends transactions to the blockchain network, ensuring that the data is permanently recorded on the blockchain.
[1538] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1539] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1540] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1541] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1542] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1543] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1544] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1545] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1546] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1547] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1548] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1549] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1550] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1551] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1552] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1553] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1554] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1555] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1556] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1557] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1558] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1559] The following is further disclosed regarding the embodiments described above.
[1560] (Claim 1)
[1561] A means of collecting visual information, auditory information, and thought patterns of talent,
[1562] A means of analyzing collected visual information, auditory information, and thought patterns,
[1563] A means of generating AI talent using an AI model based on analyzed data,
[1564] Methods for incorporating AI talent into advertising content such as commercials,
[1565] A method for saving the generated AI talent data as an NFT and registering it on the blockchain,
[1566] A system that includes this.
[1567] (Claim 2)
[1568] The system according to claim 1, wherein the means for analyzing collected visual information, audio information, and thought patterns include a video analysis algorithm, an audio analysis algorithm, and a natural language processing algorithm.
[1569] (Claim 3)
[1570] The system according to claim 1, wherein the means for incorporating the generated AI talent into advertising content is to use an advertising template.
[1571] "Example 1"
[1572] (Claim 1)
[1573] A means of collecting visual information, auditory information, and thought patterns of talent,
[1574] A means of analyzing collected visual information, auditory information, and thought patterns,
[1575] A means of generating AI talent using an AI model based on analyzed data,
[1576] Methods for incorporating AI talent into advertising content,
[1577] A method for saving the generated AI talent data as an NFT and registering it on the blockchain,
[1578] A means of obtaining an ad template and generating an ad with visual and audio consistency,
[1579] A system that includes this.
[1580] (Claim 2)
[1581] The system according to claim 1, wherein the means for analyzing collected visual information, audio information, and thought patterns include a video analysis algorithm, an audio analysis algorithm, and a natural language processing algorithm.
[1582] (Claim 3)
[1583] The system according to claim 1, which incorporates AI talent generated using an advertising template into advertising content.
[1584] "Application Example 1"
[1585] (Claim 1)
[1586] A means of collecting visual information, auditory information, and thought patterns of talent,
[1587] A means of analyzing collected visual information, auditory information, and thought patterns,
[1588] A means of generating AI talent using an AI model based on analyzed data,
[1589] A means of incorporating AI talent into advertising content that is automatically generated based on user input,
[1590] A method for saving the generated AI talent data as an NFT and registering it on the blockchain,
[1591] A system that includes this.
[1592] (Claim 2)
[1593] The system according to claim 1, wherein the means for analyzing collected visual information, audio information, and thought patterns include a video analysis algorithm, an audio analysis algorithm, and a natural language processing algorithm.
[1594] (Claim 3)
[1595] The system according to claim 1, wherein the means for incorporating the generated AI talent into an advertising template provided by the user is an advertising template that uses an advertising template.
[1596] "Example 2 of combining an emotion engine"
[1597] (Claim 1)
[1598] A means of collecting visual information, auditory information, and thought patterns of talent,
[1599] A means of analyzing collected visual information, auditory information, and thought patterns,
[1600] A means of generating AI talent using an AI model based on analyzed data,
[1601] Methods for incorporating AI talent into advertising content,
[1602] A means of storing the generated AI talent data as a non-fungible token and registering it in a distributed ledger,
[1603] A means of collecting and analyzing user sentiment data,
[1604] A means of adjusting advertising content in real time based on sentiment analysis results,
[1605] A system that includes this.
[1606] (Claim 2)
[1607] The system according to claim 1, wherein the means for analyzing collected visual information, auditory information, and thought patterns include an image analysis algorithm, an auditory analysis algorithm, and a natural language processing algorithm.
[1608] (Claim 3)
[1609] The system according to claim 1, wherein the means for incorporating the generated AI talent into advertising content is to use an advertising template.
[1610] "Application example 2 of combining emotional engines"
[1611] (Claim 1)
[1612] A means of collecting visual information, auditory information, and thought patterns of talent,
[1613] A means of analyzing collected visual information, auditory information, and thought patterns,
[1614] A means of generating AI talent using an AI model based on analyzed data,
[1615] Methods for incorporating AI talent into advertising content,
[1616] A means of recognizing emotions from the user's visual and auditory information and adjusting the AI talent's facial expressions and speech in real time,
[1617] A method for saving the generated AI talent data as an NFT and registering it on the blockchain,
[1618] A system that includes this.
[1619] (Claim 2)
[1620] The system according to claim 1, wherein the means for analyzing collected visual information, audio information, and thought patterns include a video analysis algorithm, an audio analysis algorithm, and a natural language processing algorithm.
[1621] (Claim 3)
[1622] The system according to claim 1, wherein the means for incorporating the generated AI talent into advertising content is to use an advertising template. [Explanation of Symbols]
[1623] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of collecting visual information, auditory information, and thought patterns of talent, A means of analyzing collected visual information, auditory information, and thought patterns, A means of generating AI talent using an AI model based on analyzed data, Methods for incorporating AI talent into advertising content such as commercials, A method for saving the generated AI talent data as an NFT and registering it on the blockchain, A system that includes this.
2. The system according to claim 1, wherein the means for analyzing collected visual information, audio information, and thought patterns includes a video analysis algorithm, an audio analysis algorithm, and a natural language processing algorithm.
3. The system according to claim 1, wherein the means for incorporating the generated AI talent into advertising content is to use an advertising template.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A