System

The system uses generative AI to automate ad creation, incorporating virtual talent and real-time user feedback, addressing time and cost inefficiencies and talent risks in traditional advertising processes.

JP2026018032APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119093
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

The traditional advertising creation process is time-consuming and costly, and there is a risk of ad cancellation due to celebrity scandals, with each revision requiring significant resources.

Method used

A system that uses generative AI to automatically generate advertising scripts, videos, and audio, incorporating virtual talent and animation, while allowing for real-time user feedback and revision, and efficiently exports the final advertisement to distribution platforms.

Benefits of technology

This system streamlines the ad creation process, reduces costs, and minimizes the risk of talent-related issues by generating high-quality advertisements quickly and efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018032000001_ABST
    Figure 2026018032000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving a user-input advertisement theme; means for obtaining default settings based on historical advertisement information and market trends; means for generating an advertisement script using a generation AI; means for presenting the generated advertisement script to users and receiving modification instructions; means for generating an advertisement video based on the modification instructions; means for generating and adding appropriate audio and music to the advertisement video; and means for exporting a final version of the advertisement to a delivery platform.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The traditional advertising creation process requires a significant amount of time and cost, from setting the advertising theme to producing the video and selecting the music. Furthermore, advertising production is a high-risk business, as there is a risk of the ad being canceled due to a scandal involving the celebrity appearing in the ad. Furthermore, each time an ad needs to be revised or re-created, it requires significant resources. To solve these issues, there is a need for technology that can quickly generate high-quality ads efficiently and at low cost, while also reducing the risks associated with using celebrities. [Means for solving the problem]

[0005] The present invention solves the aforementioned problems with a system that includes a means for receiving an advertising theme input by a user, a means for obtaining initial settings based on past advertising data and market trends, a means for generating an advertising script using a generation AI, a means for presenting the generated advertising script to a user and receiving revision instructions, a means for generating an advertising video based on the revision instructions, a means for generating appropriate audio and music and adding it to the advertising video, and a means for exporting the final advertisement and sending it to a distribution platform. This system not only speeds up the advertising creation process at low cost, but also reduces the risks associated with the use of talent. Furthermore, it allows for efficient revision and regeneration of advertisements, resulting in the provision of high-quality advertisements. Furthermore, the use of virtual talent and animation in advertising videos reduces the risk of advertisements being canceled due to talent scandals.

[0006] An "advertising theme" is a subject that indicates the overall direction or focus of the advertising content.

[0007] "Initial Settings" are the goals and parameters established at the beginning of the ad creation process, including, for example, target audience, ad length, key message, etc.

[0008] "Generative AI" is a system or software that uses artificial intelligence technology to automatically generate scripts and videos.

[0009] An "advertising script" is a text document that describes the content of an advertising video or audio.

[0010] "Modification instructions" are instructions for changes or improvements that a user makes to the generated advertising script or video.

[0011] "Advertising video" refers to video used as advertising, and may include images, videos, text, effects, etc.

[0012] "Audio and music" refers to audio elements such as narration and background music that are added to advertising videos.

[0013] "Final" means the final version of the advertisement after all revisions have been completed and approved by the user.

[0014] "Distribution platform" refers to a medium, such as a television station or online service, for publishing and distributing advertisements.

[0015] "Virtual talent" is not a real person, but rather a video or animation of a person created by generative AI. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention is a system for streamlining the advertisement creation process and reducing costs, specifically, it uses generation AI to automatically generate advertisement scripts, video, and audio. This system mainly operates in cooperation with a server, terminals, and users.

[0038] 1. Theme input step

[0039] The user inputs an advertising theme on the terminal, for example, specifying the theme "advertising for a new smoothie."

[0040] The terminal transmits the entered advertising theme to the server.

[0041] 2. Initial setting acquisition step

[0042] The server analyzes the past successful advertising data and market trends from the received advertising theme to obtain initial settings, such as setting the target audience to "young people," the length of the advertisement to "30 seconds," and the main message to "refreshment and health."

[0043] 3. Script generation step

[0044] The server uses AI to automatically generate ad scripts based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0045] The server sends the generated script to the terminal, where the user checks it and sends corrections or additional instructions to the server as necessary.

[0046] 4. Image generation step

[0047] The server uses the modified script to generate the ad video using AI. For example, a scene of energetic young men and women enjoying smoothies on the beach can be automatically generated.

[0048] The server sends the generated video to the terminal, where the user can view it.

[0049] 5. Speech and Music Generation Steps

[0050] The server generates appropriate audio and background music and adds them to the ad video. For example, a narration saying, "Refresh yourself this summer with our new fresh smoothie!" can be added along with refreshing background music.

[0051] The server combines the generated audio and video and transmits them to the terminal.

[0052] The user checks the audio and video and sends feedback and correction instructions to the server.

[0053] 6. Final check and correction steps

[0054] The server generates new advertisements based on the user's feedback and sends them to the terminal.

[0055] This process is repeated until the user is satisfied with the final version.

[0056] 7. Ad Output Steps

[0057] The server exports the final ad in the appropriate format and sends it to the device.

[0058] The device stores the ad and prepares it for transmission to the distribution platform.

[0059] The user uses the terminal to carry out the transmission procedure to the distribution platform.

[0060] This system makes the ad creation process fast and efficient. The use of virtual talent and animation also reduces the risk of talent misconduct. For example, a high-quality ad for a new smoothie can be generated quickly and ready for distribution to the user's satisfaction.

[0061] The processing flow will be explained below.

[0062] Step 1:

[0063] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0064] Step 2:

[0065] Based on the advertising theme received by the server, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[0066] Step 3:

[0067] Using the initial settings and ad theme acquired by the server, the AI ​​generator automatically generates an ad script. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0068] Step 4:

[0069] The server sends the generated script to the terminal, where the user confirms it. The user inputs corrections or additional instructions for the script, and the terminal sends them to the server.

[0070] Step 5:

[0071] The server uses the modified script to generate an advertising video using generative AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated video is then sent to the device for the user to view.

[0072] Step 6:

[0073] The server generates or selects appropriate audio and background music and adds them to the advertising video. For example, a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0074] Step 7:

[0075] The user checks the audio and video and inputs feedback and correction instructions. The device sends the input feedback to the server.

[0076] Step 8:

[0077] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[0078] Step 9:

[0079] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0080] Step 10:

[0081] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0082] This process streamlines the ad creation process, reducing costs, and using virtual talent also reduces the risk of talent misconduct.

[0083] Example 1

[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0085] The traditional advertising creation process requires a significant amount of time and money, and involves the risk of talent scheduling and scandals. It is also difficult to guarantee the ad's suitability for the target audience and its quality. A system that can solve these issues and generate high-quality ads quickly and efficiently is needed.

[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0087] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generative AI model, means for presenting the generated advertising script to a user and receiving instructions for modification, means for generating an advertising video based on the modified advertising script, means for generating appropriate audio and background music and adding it to the advertising video, means for reflecting user feedback and regenerating the advertising video, and means for exporting the final advertisement in an appropriate data format and transmitting it to a distribution platform. This makes it possible to create high-quality advertisements in a short period of time and efficiently provide content optimized for the target audience.

[0088] "User" means a person who uses the system to input a theme for creating an advertisement and checks and modifies the generated advertisement script and video.

[0089] "Advertising theme" refers to the subject matter of the content or message that you want to convey in your advertisement, and is entered by the user.

[0090] "Past advertising data" refers to data relating to previously created advertisements, including information on factors contributing to their success and their impact on the market.

[0091] "Market trends" are information that indicates current and projected market needs and consumer preferences.

[0092] "Initial settings" are settings that are set as basic conditions and goals when creating an advertisement, including the target audience, length of the advertisement, and main message.

[0093] A "generative AI model" is an algorithm or model that uses artificial intelligence to automatically generate new content from data.

[0094] An "advertising script" is a written expression of the content of an advertising video, and is automatically generated by a generative AI model.

[0095] "Modification instructions" refers to corrections or additional instructions given by the user to the generated advertising script or video.

[0096] "Advertising video" refers to visual content that is automatically generated using a generative AI model and is created based on an advertising script.

[0097] "Audio and background music" refers to the narration and background music (BGM) added to the advertising video, which are automatically generated by a generative AI model.

[0098] "Feedback" refers to opinions and suggestions for improvement regarding advertising scripts and videos that users provide to the server.

[0099] "Data format" refers to a standardized format for storing and communicating digital data in a particular format, an example of which is the MP4 format.

[0100] A "distribution platform" refers to an online service or website that publishes generated advertisements and delivers them to a large number of users.

[0101] This invention is a system for streamlining the advertising creation process and reducing costs. The system uses a generative AI model to automatically generate advertising scripts, video, and audio. The system mainly operates in cooperation with a server, a terminal, and a user.

[0102] First, the user accesses the system using a terminal and logs in on the user authentication screen. Then, they proceed to a screen for entering the advertising theme, where they enter the advertising theme. For example, they might set an advertising theme such as "Promotion of a newly released smoothie." The terminal then sends this entered theme to the server as an HTTP request.

[0103] The server analyzes the received HTTP request and extracts the advertising theme. Then, the server connects to a database and searches and retrieves past advertising data and market trend data. Based on this data, the server sets initial settings such as the target audience (e.g., young people), ad length (e.g., 30 seconds), and main message (e.g., freshness and health). The initial settings are sent to the terminal in a data format (e.g., JSON) and presented to the user.

[0104] The server then uses the initial settings and ad theme to create a prompt for the generative AI model (e.g., OpenAI's GPT-4). An example prompt is as follows:

[0105] "Generate a 30-second advertising script promoting a new smoothie. The target audience is young people, and the key message is 'refreshing and healthy.'"

[0106] By inputting this prompt into the generative AI model, an advertising script is automatically generated. For example, the generated script might be, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The server then sends the generated script to the device and has the user review it. The user can then enter feedback to modify the script as needed, and the script is sent from the device to the server.

[0107] The server receives feedback from the user and uses the generative AI model again to generate advertising footage based on the revised script. For example, it might generate a scene of energetic young people enjoying themselves on the beach while drinking smoothies. The generated footage is then sent to the device for the user to review.

[0108] The server then generates appropriate audio and background music. Using a generative AI model, it generates refreshing background music and narration, adding a narration such as, "Refresh yourself this summer with our new fresh smoothie!" The generated audio and video are then integrated and the final ad video is sent to the device. The user can review it and provide feedback if necessary.

[0109] Based on the user's feedback, the server regenerates the final ad video and exports it in the appropriate data format (e.g., MP4 format). It is then sent to the device for storage and preparation for uploading to the distribution platform. The user then uses their device to log in to the distribution platform and upload the ad, which is then officially distributed.

[0110] This system makes the ad creation process fast and efficient. The use of virtual characters and animations also reduces the risk of celebrity scandals. For example, an ad promoting a new smoothie can be generated quickly and with high quality, and ready for distribution to the user's satisfaction.

[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0112] Step 1:

[0113] A user accesses the system using a terminal and authenticates on the login screen. After authentication, the user proceeds to the ad theme input screen and inputs an ad theme such as "Advertising a new smoothie." The terminal sends the input ad theme to the server via an HTTP request. The input is the ad theme in text format, and the output is the HTTP request passed to the server.

[0114] Step 2:

[0115] The server analyzes the received HTTP request and extracts the advertising theme. The server connects to the database and searches for and obtains past advertising data and market trend data. During this process, the server searches the database for data on successful advertising and information on current market trends, and generates initial settings such as the target audience, advertising length, and main message. For example, the target audience is set to "young people," the advertising length to "30 seconds," and the main message to "refreshment and health." The input is the advertising theme, and the output is the initial data format (JSON).

[0116] Step 3:

[0117] The server creates a prompt for the generative AI model (e.g., GPT-4) based on the initial settings and advertising theme. For example, a prompt such as "Please generate a 30-second advertising script promoting a newly released smoothie. The target audience is young people, and the main message is 'refreshing and healthy.'" is generated. This prompt is input into the generative AI model to generate an advertising script. The input is the initial settings data and advertising theme, and the output is the generated advertising script.

[0118] Step 4:

[0119] The server sends the generated ad script to the terminal and presents it to the user. The user checks the script and inputs corrections or additional instructions as necessary. The terminal sends this feedback to the server as an HTTP request. The input is the generated ad script, and the output is the user's correction instructions (feedback).

[0120] Step 5:

[0121] The server receives feedback from users and generates ad videos using a generative AI model based on the revised ad script. For example, a scene of energetic young people enjoying themselves on the beach while drinking smoothies is automatically generated. The input is the revised ad script, and the output is the generated ad video.

[0122] Step 6:

[0123] The server generates appropriate audio and background music using a generative AI model. For example, refreshing background music and a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" are added. The server then integrates the generated audio and video and sends the completed advertising video to the device. The user reviews the generated audio and video and enters any additional feedback they may have. The input is the generated advertising video and audio, and the output is the integrated final advertising video.

[0124] Step 7:

[0125] The server receives the user's feedback, regenerates the ad video as needed, and exports the final ad in an appropriate data format (e.g., MP4). The generated video is then sent to the device for storage and preparation for uploading to the distribution platform. The process is completed when the user logs into the distribution platform using their device and uploads the ad. The input is the final feedback, and the output is the final exported ad video.

[0126] The above processing steps ensure that the ad creation process is fast and efficient, enabling high-quality ads to be generated and delivered in a short period of time.

[0127] (Application example 1)

[0128] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0129] The traditional ad creation process was often manual, resulting in time-consuming and costly issues. There was also the risk that continued use of the ad would become difficult due to celebrity scandals or contract issues. Furthermore, there were few ways for users to check and edit ad content in real time, making it difficult to efficiently create high-quality ads.

[0130] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0131] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generation AI, means for presenting the generated script and audio on a visual display device in real time, means for receiving confirmation and correction instructions from the user using the visual display device, and means for exporting the final advertisement and sending it to a distribution platform. This makes it possible to streamline the advertising creation process, reducing time and costs, and enabling users to confirm and correct advertising content in real time, enabling high-quality advertisements to be created in a short period of time.

[0132] "User" means any person or entity that uses the System to create and modify Advertisements.

[0133] The "advertising theme" is the central setting for the content and message of the advertisement to be generated, and is information input by the user.

[0134] "Past advertising data" refers to a collection of data on advertisements that have been created and distributed in the past, and is used as reference information when creating advertisements.

[0135] "Market trends" refers to information about current market trends and consumer behavior patterns.

[0136] "Initial Settings" refers to basic setting information for starting the ad creation process, and is obtained based on past ad data and market trends.

[0137] "Generative AI" is an artificial intelligence technology that uses machine learning and deep learning to automatically generate advertising scripts, video, and audio.

[0138] An "advertising script" is a document that describes the specific content and message of an advertisement.

[0139] A "visual display device" is an electronic device that displays the generated advertising script and video to a user in real time. Examples include smart glasses and head-mounted displays.

[0140] A "distribution platform" is an online service or system for distributing generated advertisements.

[0141] "Modification instructions" are instructions for changes or additions made by the user to the generated advertisement script or video.

[0142] A "generated script" is an advertising document automatically generated by the generation AI.

[0143] "Final Ad" means the final ad content that has been modified and adjusted based on user feedback and correction instructions.

[0144] The system for realizing this invention operates in cooperation with three entities: a server, a terminal, and a user.

[0145] First, the user uses the smart glasses to input the advertising theme by voice. This voice input is picked up by the microphone built into the smart glasses and converted into text data using speech recognition software (speech_recognition library). For example, the user may input "Advertisement for new smoothie release."

[0146] Based on the received advertising theme, the server analyzes past advertising data and market trends to obtain initial settings, such as target audience, advertising length, and key message, by referring to an internal database and external market data.

[0147] The server then uses a generative AI (such as OpenAI's GPT-3) to automatically generate an ad script. The generated script is displayed in real time on the smart glasses' display, and the user can make corrections by voice. For example, the following prompt sentence can be input to the generative AI:

[0148] "You are an ad creator. Please create an ad script for the theme 'Promotion of a new smoothie'."

[0149] Based on user feedback, the server again uses the AI ​​to generate a revised ad script, and then uses text-to-speech software (Google Cloud Text-to-Speech) to generate an appropriate narration voice, which can also be heard in real time on the smart glasses.

[0150] Once the final ad script and audio are completed, the server merges them with the video to automatically generate the ad video, which is based on the latest content reviewed and edited by the user through the smart glasses. Finally, the server exports the generated ad in the appropriate format and sends it to the distribution platform.

[0151] In this way, the server, terminal, and user work together, and prompt sentences using generative AI models are used to streamline the ad creation process and quickly generate high-quality ads.

[0152] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0153] Step 1:

[0154] The user wears the smart glasses and inputs the advertising theme by voice. The voice input is acquired through a microphone built into the smart glasses. This voice data is sent to the device and converted into text data using speech recognition software (speech_recognition library) on the device side. This allows the advertising theme to be acquired in text format.

[0155] Step 2:

[0156] The device transmits the converted text data of the advertising theme to the server. Based on the received advertising theme, the server retrieves initial settings from its internal database and market trend data. For example, it sets the target audience, the length of the advertisement, and the main message. This provides basic setting information for creating the advertisement.

[0157] Step 3:

[0158] The server generates an ad script using a generative AI (OpenAI's GPT-3) based on the initial settings and ad theme. Specifically, it inputs the following prompt to the generative AI:

[0159] "You are an ad creator. Please create an ad script for the theme 'Promotion of a new smoothie'."

[0160] The generated ad script is obtained in text format.

[0161] Step 4:

[0162] The generated script is sent from the server to the device, which then displays it in real time on the smart glasses display. The user can check the content and make corrections using voice or gestures. These corrections are then converted to text on the device and sent back to the server.

[0163] Step 5:

[0164] The server uses the AI ​​to modify and regenerate the ad script based on the modification instructions sent by the user. This process is repeated until the user is satisfied. The modified script is obtained in text format and sent back to the device.

[0165] Step 6:

[0166] The server generates narration using speech synthesis software (Google Cloud Text-to-Speech) based on the finalized ad script, and this audio data is sent to the smart glasses via the device, allowing the user to listen to it in real time.

[0167] Step 7:

[0168] Once the final ad script and audio are completed, the server automatically generates the ad video based on them. The generated video data is sent to the device and can be viewed by the user through smart glasses. This process is repeated until the user is satisfied.

[0169] Step 8:

[0170] Finally, the server exports the final ad in the appropriate format and sends it to the distribution platform, where the user-created ad is ready to be distributed online.

[0171] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0172] This invention is a system for streamlining the advertisement creation process and reducing costs. Specifically, it uses generative AI to automatically generate advertisement scripts, video, and audio, and further combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[0173] 1. Theme input step

[0174] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0175] 2. Initial setting acquisition step

[0176] Based on the advertising theme received by the server, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[0177] 3. Script generation step

[0178] Using the initial settings and ad theme acquired by the server, the AI ​​generator automatically generates an ad script. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0179] The server sends the generated script to the terminal, where the user confirms it. The user inputs corrections or additional instructions for the script, and the terminal sends them to the server.

[0180] 4. Image generation step

[0181] The server uses the modified script to generate an advertising video using generative AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated video is then sent to the device for the user to view.

[0182] 5. Speech and Music Generation Steps

[0183] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0184] The user checks the audio and video and inputs feedback and correction instructions. The device sends the input feedback to the server.

[0185] 6. Feedback step by emotion engine

[0186] The server uses an emotion engine to recognize the user's emotions, for example, by analyzing the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust).

[0187] The server then uses data from the emotion engine to suggest modifications to the ad script, for example, to include a more positive message or image if the user expresses dissatisfaction.

[0188] The server uses the user's emotion data to select video elements that correspond to the user's preferences and emotions when generating advertising videos, and also selects audio and music that correspond to the user's emotions and preferences.

[0189] 7. Final Check and Correction Steps

[0190] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[0191] 8. Ad Output Steps

[0192] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0193] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0194] This system not only speeds up and streamlines the ad creation process, but also utilizes an emotion engine to provide high-quality ads based on user emotions. As a specific example, an ad for a new smoothie product can be generated in a short time in a way that satisfies users, and is ready for distribution.

[0195] The processing flow will be explained below.

[0196] Step 1:

[0197] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0198] Step 2:

[0199] Based on the received advertising theme, the server refers to the database for past successful advertising data and market trends, and obtains initial settings such as target audience, advertising length, and main message. For example, the target audience may be set to "young people," the advertising length to "30 seconds," and the main message to "refreshment and health."

[0200] Step 3:

[0201] The server uses the initial settings and ad theme to automatically generate an ad script using AI generation. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The server then sends the generated script to the device.

[0202] Step 4:

[0203] The user checks the ad script generated on the device. The user inputs corrections or additional instructions for the script, which the device then sends to the server. For example, the user can send feedback such as, "I'd like it to be a bit more lively and casual."

[0204] Step 5:

[0205] The server uses the generative AI to generate advertising footage based on the revised script. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The server then sends the generated footage to the device, where the user can view it.

[0206] Step 6:

[0207] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0208] Step 7:

[0209] The user checks the audio and video on the device and inputs feedback and correction instructions. For example, the user may give feedback such as "Please speak the narration a little more slowly." The device then sends the input feedback to the server.

[0210] Step 8:

[0211] The server uses an emotion engine to recognize the user's emotions, for example, by analyzing the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust).

[0212] Step 9:

[0213] The server then suggests modifications to the ad script based on the emotion engine data, for example, suggesting including more positive messaging or images if the user expresses dissatisfaction.

[0214] Step 10:

[0215] The server uses the user's emotional data to select video elements that reflect the user's preferences and emotions when generating advertising videos. For example, it increases the number of scenes with bright colors and smiling faces. The server also selects audio and music based on the user's emotions and preferences.

[0216] Step 11:

[0217] The server regenerates the revised ad and sends it to the device, repeating this process until the user is satisfied with the final version.

[0218] Step 12:

[0219] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0220] Step 13:

[0221] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0222] Example 2

[0223] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0224] The ad creation process is typically time-consuming and resource-intensive. Furthermore, the quality of the ad may not perfectly match user emotions and preferences. Traditional methods have difficulty incorporating user feedback in real time, especially in generating ad scripts and audio and video. Furthermore, creating ads that reflect user emotions and preferences is difficult, making it challenging to quickly create high-quality, effective ads.

[0225] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0226] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generation AI, means for presenting the generated advertising script to a user and receiving correction instructions, means for generating an advertising video based on the correction instructions, means for generating appropriate voice and music and adding them to the advertising video, means for recognizing the user's emotional state and analyzing the feedback, means for regenerating the advertising script and advertising video based on the emotional state, and means for exporting the final advertisement and sending it to a distribution platform. This streamlines the advertising creation process and enables the rapid generation of high-quality advertisements based on users' emotions and preferences.

[0227] "User" refers to a user who uses the advertisement creation system.

[0228] "Advertising theme" refers to keywords or phrases that outline the content or purpose of an advertisement.

[0229] "Past advertising data" refers to information about advertisements that have been produced and distributed to date, including data such as effectiveness measurement results and creative elements used.

[0230] "Market trends" refers to information that shows current trends in consumer preferences, trends, economic conditions, etc.

[0231] "Initial Settings" refers to the basic parameter settings required for creating an ad, including the target audience, ad length, key message, etc.

[0232] "Generative AI" refers to technologies and models that use artificial intelligence to generate specific content or data.

[0233] An "advertising script" refers to a text scenario or script that serves as the basis for advertising video and audio.

[0234] "Modification instructions" refer to instructions for changes or additions to advertising scripts, videos, etc. submitted by users.

[0235] "Advertising footage" refers to visual content used as advertising, including video content.

[0236] "Audio and music" refers to audio elements such as narration, sound effects, and background music that are added to advertising footage.

[0237] "Emotional state" refers to the psychological state judged from the user's facial expression, voice, etc., and includes excitement, satisfaction, disgust, etc.

[0238] "Regeneration" refers to a new generation process that modifies or improves upon the content that was initially generated.

[0239] "Distribution Platform" refers to an online platform or medium for distributing generated advertisements.

[0240] This invention is a system for streamlining the advertisement creation process and reducing costs, and in particular, it uses generative AI to automatically generate advertisement scripts, video, and audio, and further combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[0241] Hardware and Software Use

[0242] This system utilizes multiple hardware and software components. The main hardware components include servers and terminals. The software components include generative AI models (e.g., OpenAI's GPT-4), emotion engines (e.g., Microsoft Azure's Face API), and speech synthesis technologies (e.g., Amazon Polly).

[0243] The server receives the ad theme and obtains initial settings based on past ad data and market trends. It also generates ad scripts using a generative AI model and generates ad videos based on user modifications. It also generates suitable voice and music, recognizes user emotions using an emotion engine, and reflects user feedback.

[0244] The terminal is a device that is directly operated by the user, and is used to input advertising themes, check the generated advertising script and video, and input correction instructions. It also saves the final advertising data sent from the server and prepares it for transmission to the distribution platform.

[0245] Specific actions and examples

[0246] Below is a concrete example of the ad creation process.

[0247] 1. Entering a theme: The user enters "Advertisement for a new smoothie" into the input form on the device.

[0248] 2. Obtaining initial settings: The server obtains the target audience, optimal ad length, key messages, etc. from past advertising data and market trends.

[0249] 3. Script generation: The server inputs the following prompt into the generative AI model:

[0250] Ad theme: Promotion of a new smoothie

[0251] Target audience: Young people

[0252] Ad length: 30 seconds

[0253] Main message: Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy health and deliciousness all at once.

[0254] 4. Checking and modifying the script: The server sends the generated script to the terminal, where the user can check it. For example, the user may input instructions for modification, such as "I want a more moving phrase."

[0255] 5. Video generation: The server uses the generative AI to generate advertising videos based on the revised script. For example, it generates a scene of energetic young people enjoying themselves on the beach while drinking smoothies.

[0256] 6. Adding audio and music: The server adds a narration saying "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music to the advertising video.

[0257] 7. Feedback using emotion engine: The server uses an emotion engine to recognize the user's emotional state from their facial expressions and tone of voice, and suggests modifications to the advertisement.

[0258] 8. Final review and output: The server exports the final ad in the appropriate format (e.g., MP4 file) and sends it to the device. The user then performs a final review and sends it to the distribution platform.

[0259] This system not only makes the ad creation process fast and efficient, but also uses an emotion engine to provide high-quality ads based on user emotions. As a specific example, an ad for a new smoothie product can be generated in a short time in a way that satisfies users and is ready to be distributed.

[0260] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0261] Step 1:

[0262] Entering the theme

[0263] The user operates the screen of the device to input the advertising theme. Specifically, the user enters "Advertisement for the newly released smoothie" into the input form.

[0264] The device sends the entered advertising theme to the server in the form of an HTTP POST request. The input data is "Advertising theme: Promotion of newly released smoothie."

[0265] Step 2:

[0266] Get initial configuration

[0267] The server analyzes the received advertisement theme and extracts keywords related to the theme (e.g., "new release," "smoothie," "advertisement").

[0268] The server uses these keywords to retrieve data on past successful ads and market trends from a database. Specific initial settings include target audience, ad length, key message, etc.

[0269] The server structures the initial setup data and prepares it as input data for the next step. The output data is information about "target audience: young people, ad length: 30 seconds, main message: health and delicious."

[0270] Step 3:

[0271] Generate scripts

[0272] The server sends the initial settings and advertising theme obtained to the generative AI model as a prompt.

[0273] Specific prompt:

[0274] Ad theme: Promotion of a new smoothie

[0275] Target audience: Young people

[0276] Ad length: 30 seconds

[0277] Main message: Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy health and deliciousness all at once.

[0278] The generative AI model (on the server) generates an ad script based on this prompt, for example, "Refresh yourself this summer with our new fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0279] The server sends the generated script to the terminal, and the output data is the generated ad script.

[0280] Step 4:

[0281] Check and correct the script

[0282] The user checks the advertisement script displayed on the terminal.

[0283] The user inputs corrections or additional instructions for the script, such as "I want a more moving phrase."

[0284] The device retransmits this feedback to the server. The input data is a correction instruction to "make the expression more moving."

[0285] Step 5:

[0286] Video generation

[0287] The server uses a generative AI to generate advertising videos based on the revised script.

[0288] As a specific example, we generate a scene of energetic young people enjoying themselves on the beach while drinking smoothies.

[0289] The server transmits the generated video data to the terminal, where the user confirms it. The output data is the generated advertising video.

[0290] Step 6:

[0291] Adding voice and music

[0292] The server generates and adds appropriate audio (e.g., narration) and music to the advertising video.

[0293] As a specific example, a narration saying "Refresh yourself this summer with our newly released fresh smoothie!" is combined with refreshing background music.

[0294] The server transmits the integrated audio and video as a single advertisement data to the terminal, which the user confirms. The output data is the completed advertisement content.

[0295] Step 7:

[0296] Emotional Engine Feedback

[0297] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust, etc.).

[0298] The server then uses the data from the emotion engine to suggest improvements to the ad script or video. For example, if the user expresses dissatisfaction, it will select a more positive message or image.

[0299] The server regenerates and optimizes the advertising script and video based on the emotional state, and the output data is the advertising script and video corresponding to the emotional state.

[0300] Step 8:

[0301] Final check and output

[0302] The server exports the final ad in the appropriate format (e.g. MP4 file) and sends it to the device.

[0303] The device stores the advertising data and prepares it for transmission to the distribution platform.

[0304] The user uses their device to submit the ad to the distribution platform, which officially publishes the ad. The output data is the final ad file.

[0305] (Application example 2)

[0306] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0307] The traditional ad creation process required a lot of time and money, and it was difficult to effectively reflect user emotions and feedback. Repeated revisions to ad content were particularly time-consuming, potentially reducing the quality of the final ad. Furthermore, there was a lack of a way to quickly and effectively generate ads that responded to the emotions and expectations of today's highly diverse consumers.

[0308] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving an advertising theme input by a user; means for acquiring initial settings based on past advertising data and market trends; means for generating an advertising script using a generation AI; means for presenting the generated advertising script to a user and receiving correction instructions; means for generating an advertising video based on the correction instructions; means for generating appropriate voice and music and adding it to the advertising video; means for analyzing the user's facial expressions and voice and acquiring emotional data; means for re-correcting the advertisement based on the emotional data; and means for exporting the final advertisement and sending it to a distribution platform. This makes it possible to quickly and efficiently advance the advertisement creation process and generate high-quality advertisements that reflect consumer emotions and expectations.

[0309] The "advertising theme" refers to the basic idea or purpose of the content of the advertisement, and is input by the user.

[0310] "Historical Advertising Data" means historical information about advertisements that have been created and distributed, including details of successful and unsuccessful advertisements.

[0311] "Market trends" refers to information about overall market trends, such as consumer preferences and trends, and the actions of competitors.

[0312] "Initial settings" refers to the basic setting information required to create an ad, such as the target audience, key message, and length of the ad.

[0313] "Generative AI" is a technology that uses artificial intelligence to generate text, video, audio, etc., and in this invention is used to generate advertising scripts, video, and audio.

[0314] An "ad script" is the text that makes up the content of an advertisement, which is generated by the generation AI and can be modified by the user.

[0315] "Modification instructions" are instructions given by the user to the advertising script or video, including changes, additions, or deletions of content.

[0316] "Advertising video" refers to video advertising content created by generative AI and is generated based on a script.

[0317] "Audio and music" refers to narration and background music generated to complement the advertising video.

[0318] "Emotion data" is information about the user's emotional state obtained by analyzing changes in the user's facial expressions and voice.

[0319] "Export" refers to the process of converting the final generated advertisement into an appropriate format so that it can be sent to an external distribution platform.

[0320] This invention is a system for streamlining the advertising creation process and reducing costs. Specifically, it uses generative AI to automatically generate advertising scripts, video, and audio, and then combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[0321] 1. Theme input step

[0322] The user inputs an advertising theme into an input form on the terminal. This advertising theme is sent to the server as a specific prompt sentence. For example, the theme is "Advertisement for a new smoothie."

[0323] 2. Initial setting acquisition step

[0324] Based on the received advertising theme, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[0325] 3. Script generation step

[0326] The server uses AI to automatically generate an ad script based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The generated script is sent to the device, where the user can review and edit it.

[0327] 4. Image generation step

[0328] The server uses the modified script to generate advertising footage using AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated footage is then sent to the device for the user to view.

[0329] 5. Speech and Music Generation Steps

[0330] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device. The user can check the audio and video and input feedback and corrections.

[0331] 6. Feedback step by emotion engine

[0332] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust). Based on the emotion engine data, it suggests modifications to the advertising script and video. For example, if the user expresses dissatisfaction, it may include more positive messages and images. The server uses the user's emotion data to select video elements that correspond to the user's preferences and emotions when generating advertising videos. It also selects audio and music based on the user's emotions and preferences.

[0333] 7. Final Check and Correction Steps

[0334] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[0335] 8. Ad Output Steps

[0336] The server exports the final ad in the appropriate format and sends it to the device, which stores the ad and prepares it for transmission to the TV station or other distribution platform, and the user uses the device to submit it to the distribution platform, where it is published.

[0337] Hardware and software used

[0338] Hardware: The device used by the user (smartphone, tablet, HMD, smart glasses, etc.)

[0339] Software: Generative AI models, emotion engines, database management systems, interface software

[0340] Prompt Sentence Examples

[0341] "Please generate an advertising script for our newly released fresh smoothie, targeting health-conscious people in their 20s and 30s. The main message is 'delicious and healthy.'"

[0342] This makes it possible to quickly and efficiently generate high-quality advertisements that reflect user emotions and feedback.

[0343] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0344] Step 1:

[0345] The user inputs an advertising theme into the input form on the terminal. The input advertising theme might be, for example, "Advertising a new smoothie." This theme is sent to the server as a prompt. The input at this stage is the advertising theme, and the output is the theme sent to the server.

[0346] Step 2:

[0347] Based on the received advertising theme, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, key message, etc. The input of this stage is the advertising theme, and the output is the initial setting information.

[0348] Step 3:

[0349] The server uses the generation AI to automatically generate an ad script based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy health and deliciousness all at once." The input at this stage is the initial settings information and ad theme, and the output is the generated ad script.

[0350] Step 4:

[0351] The server sends the generated ad script to the device, where the user can review and modify it. User feedback and modification instructions are sent to the server via the device. The input at this stage is the generated ad script, and the output is modification instructions from the user.

[0352] Step 5:

[0353] The server generates the ad video based on the modification instructions. Using generative AI, it can generate a scene of energetic young men and women enjoying smoothies on the beach, for example. The input at this stage is the modified ad script, and the output is the generated ad video.

[0354] Step 6:

[0355] The server generates appropriate audio and background music to add to the ad video. For example, a narration voice saying "Refresh yourself this summer with our new fresh smoothie!" and refreshing background music are added. The generated audio and video are sent to the terminal, where the user can review them and input feedback and corrections. The input at this stage is the generated ad video, and the output is the ad video with the added audio and music.

[0356] Step 7:

[0357] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust). Based on the emotion data, it suggests modifications to the ad script or video. The input of this stage is the user's emotion data, and the output is the proposed modification of the ad content.

[0358] Step 8:

[0359] The server regenerates the ad based on the user's feedback and sends it to the device. This process is repeated until the user is satisfied with the final version. The input at this stage is suggested revisions based on emotional data, and the output is the final revised version of the ad.

[0360] Step 9:

[0361] The server exports the final ad in the appropriate format and sends it to the device, which stores the ad and prepares it for transmission to the TV station or other distribution platform. The input to this stage is the final ad data, and the output is the ad data ready for distribution.

[0362] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0363] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0364] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0365] [Second embodiment]

[0366] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0367] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0368] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0369] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0370] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0371] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0372] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0373] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0374] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0375] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0376] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0377] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0378] This invention is a system for streamlining the advertisement creation process and reducing costs, specifically, it uses generation AI to automatically generate advertisement scripts, video, and audio. This system mainly operates in cooperation with a server, terminals, and users.

[0379] 1. Theme input step

[0380] The user inputs an advertising theme on the terminal, for example, specifying the theme "advertising for a new smoothie."

[0381] The terminal transmits the entered advertising theme to the server.

[0382] 2. Initial setting acquisition step

[0383] The server analyzes the past successful advertising data and market trends from the received advertising theme to obtain initial settings, such as setting the target audience to "young people," the length of the advertisement to "30 seconds," and the main message to "refreshment and health."

[0384] 3. Script generation step

[0385] The server uses AI to automatically generate ad scripts based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0386] The server sends the generated script to the terminal, where the user checks it and sends corrections or additional instructions to the server as necessary.

[0387] 4. Image generation step

[0388] The server uses the modified script to generate the ad video using AI. For example, a scene of energetic young men and women enjoying smoothies on the beach can be automatically generated.

[0389] The server sends the generated video to the terminal, where the user can view it.

[0390] 5. Speech and Music Generation Steps

[0391] The server generates appropriate audio and background music and adds them to the ad video. For example, a narration saying, "Refresh yourself this summer with our new fresh smoothie!" can be added along with refreshing background music.

[0392] The server combines the generated audio and video and transmits them to the terminal.

[0393] The user checks the audio and video and sends feedback and correction instructions to the server.

[0394] 6. Final check and correction steps

[0395] The server generates new advertisements based on the user's feedback and sends them to the terminal.

[0396] This process is repeated until the user is satisfied with the final version.

[0397] 7. Ad Output Steps

[0398] The server exports the final ad in the appropriate format and sends it to the device.

[0399] The device stores the ad and prepares it for transmission to the distribution platform.

[0400] The user uses the terminal to carry out the transmission procedure to the distribution platform.

[0401] This system makes the ad creation process fast and efficient. The use of virtual talent and animation also reduces the risk of talent misconduct. For example, a high-quality ad for a new smoothie can be generated quickly and ready for distribution to the user's satisfaction.

[0402] The processing flow will be explained below.

[0403] Step 1:

[0404] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0405] Step 2:

[0406] Based on the advertising theme received by the server, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[0407] Step 3:

[0408] Using the initial settings and ad theme acquired by the server, the AI ​​generator automatically generates an ad script. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0409] Step 4:

[0410] The server sends the generated script to the terminal, where the user confirms it. The user inputs corrections or additional instructions for the script, and the terminal sends them to the server.

[0411] Step 5:

[0412] The server uses the modified script to generate an advertising video using generative AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated video is then sent to the device for the user to view.

[0413] Step 6:

[0414] The server generates or selects appropriate audio and background music and adds them to the advertising video. For example, a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0415] Step 7:

[0416] The user checks the audio and video and inputs feedback and correction instructions. The device sends the input feedback to the server.

[0417] Step 8:

[0418] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[0419] Step 9:

[0420] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0421] Step 10:

[0422] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0423] This process streamlines the ad creation process, reducing costs, and using virtual talent also reduces the risk of talent misconduct.

[0424] Example 1

[0425] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0426] The traditional advertising creation process requires a significant amount of time and money, and involves the risk of talent scheduling and scandals. It is also difficult to guarantee the ad's suitability for the target audience and its quality. A system that can solve these issues and generate high-quality ads quickly and efficiently is needed.

[0427] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0428] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generative AI model, means for presenting the generated advertising script to a user and receiving instructions for modification, means for generating an advertising video based on the modified advertising script, means for generating appropriate audio and background music and adding it to the advertising video, means for reflecting user feedback and regenerating the advertising video, and means for exporting the final advertisement in an appropriate data format and transmitting it to a distribution platform. This makes it possible to create high-quality advertisements in a short period of time and efficiently provide content optimized for the target audience.

[0429] "User" means a person who uses the system to input a theme for creating an advertisement and checks and modifies the generated advertisement script and video.

[0430] "Advertising theme" refers to the subject matter of the content or message that you want to convey in your advertisement, and is entered by the user.

[0431] "Past advertising data" refers to data relating to previously created advertisements, including information on factors contributing to their success and their impact on the market.

[0432] "Market trends" are information that indicates current and projected market needs and consumer preferences.

[0433] "Initial settings" are settings that are set as basic conditions and goals when creating an advertisement, including the target audience, length of the advertisement, and main message.

[0434] A "generative AI model" is an algorithm or model that uses artificial intelligence to automatically generate new content from data.

[0435] An "advertising script" is a written expression of the content of an advertising video, and is automatically generated by a generative AI model.

[0436] "Modification instructions" refers to corrections or additional instructions given by the user to the generated advertising script or video.

[0437] "Advertising video" refers to visual content that is automatically generated using a generative AI model and is created based on an advertising script.

[0438] "Audio and background music" refers to the narration and background music (BGM) added to the advertising video, which are automatically generated by a generative AI model.

[0439] "Feedback" refers to opinions and suggestions for improvement regarding advertising scripts and videos that users provide to the server.

[0440] "Data format" refers to a standardized format for storing and communicating digital data in a particular format, an example of which is the MP4 format.

[0441] A "distribution platform" refers to an online service or website that publishes generated advertisements and delivers them to a large number of users.

[0442] This invention is a system for streamlining the advertising creation process and reducing costs. The system uses a generative AI model to automatically generate advertising scripts, video, and audio. The system mainly operates in cooperation with a server, a terminal, and a user.

[0443] First, the user accesses the system using a terminal and logs in on the user authentication screen. Then, they proceed to a screen for entering the advertising theme, where they enter the advertising theme. For example, they might set an advertising theme such as "Promotion of a newly released smoothie." The terminal then sends this entered theme to the server as an HTTP request.

[0444] The server analyzes the received HTTP request and extracts the advertising theme. Then, the server connects to a database and searches and retrieves past advertising data and market trend data. Based on this data, the server sets initial settings such as the target audience (e.g., young people), ad length (e.g., 30 seconds), and main message (e.g., freshness and health). The initial settings are sent to the terminal in a data format (e.g., JSON) and presented to the user.

[0445] The server then uses the initial settings and ad theme to create a prompt for the generative AI model (e.g., OpenAI's GPT-4). An example prompt is as follows:

[0446] "Generate a 30-second advertising script promoting a new smoothie. The target audience is young people, and the key message is 'refreshing and healthy.'"

[0447] By inputting this prompt into the generative AI model, an advertising script is automatically generated. For example, the generated script might be, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The server then sends the generated script to the device and has the user review it. The user can then enter feedback to modify the script as needed, and the script is sent from the device to the server.

[0448] The server receives feedback from the user and uses the generative AI model again to generate advertising footage based on the revised script. For example, it might generate a scene of energetic young people enjoying themselves on the beach while drinking smoothies. The generated footage is then sent to the device for the user to review.

[0449] The server then generates appropriate audio and background music. Using a generative AI model, it generates refreshing background music and narration, adding a narration such as, "Refresh yourself this summer with our new fresh smoothie!" The generated audio and video are then integrated and the final ad video is sent to the device. The user can review it and provide feedback if necessary.

[0450] Based on the user's feedback, the server regenerates the final ad video and exports it in the appropriate data format (e.g., MP4 format). It is then sent to the device for storage and preparation for uploading to the distribution platform. The user then uses their device to log in to the distribution platform and upload the ad, which is then officially distributed.

[0451] This system makes the ad creation process fast and efficient. The use of virtual characters and animations also reduces the risk of celebrity scandals. For example, an ad promoting a new smoothie can be generated quickly and with high quality, and ready for distribution to the user's satisfaction.

[0452] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0453] Step 1:

[0454] A user accesses the system using a terminal and authenticates on the login screen. After authentication, the user proceeds to the ad theme input screen and inputs an ad theme such as "Advertising a new smoothie." The terminal sends the input ad theme to the server via an HTTP request. The input is the ad theme in text format, and the output is the HTTP request passed to the server.

[0455] Step 2:

[0456] The server analyzes the received HTTP request and extracts the advertising theme. The server connects to the database and searches for and obtains past advertising data and market trend data. During this process, the server searches the database for data on successful advertising and information on current market trends, and generates initial settings such as the target audience, advertising length, and main message. For example, the target audience is set to "young people," the advertising length to "30 seconds," and the main message to "refreshment and health." The input is the advertising theme, and the output is the initial data format (JSON).

[0457] Step 3:

[0458] The server creates a prompt for the generative AI model (e.g., GPT-4) based on the initial settings and advertising theme. For example, a prompt such as "Please generate a 30-second advertising script promoting a newly released smoothie. The target audience is young people, and the main message is 'refreshing and healthy.'" is generated. This prompt is input into the generative AI model to generate an advertising script. The input is the initial settings data and advertising theme, and the output is the generated advertising script.

[0459] Step 4:

[0460] The server sends the generated ad script to the terminal and presents it to the user. The user checks the script and inputs corrections or additional instructions as necessary. The terminal sends this feedback to the server as an HTTP request. The input is the generated ad script, and the output is the user's correction instructions (feedback).

[0461] Step 5:

[0462] The server receives feedback from users and generates ad videos using a generative AI model based on the revised ad script. For example, a scene of energetic young people enjoying themselves on the beach while drinking smoothies is automatically generated. The input is the revised ad script, and the output is the generated ad video.

[0463] Step 6:

[0464] The server generates appropriate audio and background music using a generative AI model. For example, refreshing background music and a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" are added. The server then integrates the generated audio and video and sends the completed advertising video to the device. The user reviews the generated audio and video and enters any additional feedback they may have. The input is the generated advertising video and audio, and the output is the integrated final advertising video.

[0465] Step 7:

[0466] The server receives the user's feedback, regenerates the ad video as needed, and exports the final ad in an appropriate data format (e.g., MP4). The generated video is then sent to the device for storage and preparation for uploading to the distribution platform. The process is completed when the user logs into the distribution platform using their device and uploads the ad. The input is the final feedback, and the output is the final exported ad video.

[0467] The above processing steps ensure that the ad creation process is fast and efficient, enabling high-quality ads to be generated and delivered in a short period of time.

[0468] (Application example 1)

[0469] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0470] The traditional ad creation process was often manual, resulting in time-consuming and costly issues. There was also the risk that continued use of the ad would become difficult due to celebrity scandals or contract issues. Furthermore, there were few ways for users to check and edit ad content in real time, making it difficult to efficiently create high-quality ads.

[0471] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0472] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generation AI, means for presenting the generated script and audio on a visual display device in real time, means for receiving confirmation and correction instructions from the user using the visual display device, and means for exporting the final advertisement and sending it to a distribution platform. This makes it possible to streamline the advertising creation process, reducing time and costs, and enabling users to confirm and correct advertising content in real time, enabling high-quality advertisements to be created in a short period of time.

[0473] "User" means any person or entity that uses the System to create and modify Advertisements.

[0474] The "advertising theme" is the central setting for the content and message of the advertisement to be generated, and is information input by the user.

[0475] "Past advertising data" refers to a collection of data on advertisements that have been created and distributed in the past, and is used as reference information when creating advertisements.

[0476] "Market trends" refers to information about current market trends and consumer behavior patterns.

[0477] "Initial Settings" refers to basic setting information for starting the ad creation process, and is obtained based on past ad data and market trends.

[0478] "Generative AI" is an artificial intelligence technology that uses machine learning and deep learning to automatically generate advertising scripts, video, and audio.

[0479] An "advertising script" is a document that describes the specific content and message of an advertisement.

[0480] A "visual display device" is an electronic device that displays the generated advertising script and video to a user in real time. Examples include smart glasses and head-mounted displays.

[0481] A "distribution platform" is an online service or system for distributing generated advertisements.

[0482] "Modification instructions" are instructions for changes or additions made by the user to the generated advertisement script or video.

[0483] A "generated script" is an advertising document automatically generated by the generation AI.

[0484] "Final Ad" means the final ad content that has been modified and adjusted based on user feedback and correction instructions.

[0485] The system for realizing this invention operates in cooperation with three entities: a server, a terminal, and a user.

[0486] First, the user uses the smart glasses to input the advertising theme by voice. This voice input is picked up by the microphone built into the smart glasses and converted into text data using speech recognition software (speech_recognition library). For example, the user may input "Advertisement for new smoothie release."

[0487] Based on the received advertising theme, the server analyzes past advertising data and market trends to obtain initial settings, such as target audience, advertising length, and key message, by referring to an internal database and external market data.

[0488] The server then uses a generative AI (such as OpenAI's GPT-3) to automatically generate an ad script. The generated script is displayed in real time on the smart glasses' display, and the user can make corrections by voice. For example, the following prompt sentence can be input to the generative AI:

[0489] "You are an ad creator. Please create an ad script for the theme 'Promotion of a new smoothie'."

[0490] Based on user feedback, the server again uses the AI ​​to generate a revised ad script, and then uses text-to-speech software (Google Cloud Text-to-Speech) to generate an appropriate narration voice, which can also be heard in real time on the smart glasses.

[0491] Once the final ad script and audio are completed, the server merges them with the video to automatically generate the ad video, which is based on the latest content reviewed and edited by the user through the smart glasses. Finally, the server exports the generated ad in the appropriate format and sends it to the distribution platform.

[0492] In this way, the server, terminal, and user work together, and prompt sentences using generative AI models are used to streamline the ad creation process and quickly generate high-quality ads.

[0493] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0494] Step 1:

[0495] The user wears the smart glasses and inputs the advertising theme by voice. The voice input is acquired through a microphone built into the smart glasses. This voice data is sent to the device and converted into text data using speech recognition software (speech_recognition library) on the device side. This allows the advertising theme to be acquired in text format.

[0496] Step 2:

[0497] The device transmits the converted text data of the advertising theme to the server. Based on the received advertising theme, the server retrieves initial settings from its internal database and market trend data. For example, it sets the target audience, the length of the advertisement, and the main message. This provides basic setting information for creating the advertisement.

[0498] Step 3:

[0499] The server generates an ad script using a generative AI (OpenAI's GPT-3) based on the initial settings and ad theme. Specifically, it inputs the following prompt to the generative AI:

[0500] "You are an ad creator. Please create an ad script for the theme 'Promotion of a new smoothie'."

[0501] The generated ad script is obtained in text format.

[0502] Step 4:

[0503] The generated script is sent from the server to the device, which then displays it in real time on the smart glasses display. The user can check the content and make corrections using voice or gestures. These corrections are then converted to text on the device and sent back to the server.

[0504] Step 5:

[0505] The server uses the AI ​​to modify and regenerate the ad script based on the modification instructions sent by the user. This process is repeated until the user is satisfied. The modified script is obtained in text format and sent back to the device.

[0506] Step 6:

[0507] The server generates narration using speech synthesis software (Google Cloud Text-to-Speech) based on the finalized ad script, and this audio data is sent to the smart glasses via the device, allowing the user to listen to it in real time.

[0508] Step 7:

[0509] Once the final ad script and audio are completed, the server automatically generates the ad video based on them. The generated video data is sent to the device and can be viewed by the user through smart glasses. This process is repeated until the user is satisfied.

[0510] Step 8:

[0511] Finally, the server exports the final ad in the appropriate format and sends it to the distribution platform, where the user-created ad is ready to be distributed online.

[0512] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0513] This invention is a system for streamlining the advertisement creation process and reducing costs. Specifically, it uses generative AI to automatically generate advertisement scripts, video, and audio, and further combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[0514] 1. Theme input step

[0515] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0516] 2. Initial setting acquisition step

[0517] Based on the advertising theme received by the server, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[0518] 3. Script generation step

[0519] Using the initial settings and ad theme acquired by the server, the AI ​​generator automatically generates an ad script. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0520] The server sends the generated script to the terminal, where the user confirms it. The user inputs corrections or additional instructions for the script, and the terminal sends them to the server.

[0521] 4. Image generation step

[0522] The server uses the modified script to generate an advertising video using generative AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated video is then sent to the device for the user to view.

[0523] 5. Speech and Music Generation Steps

[0524] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0525] The user checks the audio and video and inputs feedback and correction instructions. The device sends the input feedback to the server.

[0526] 6. Feedback step by emotion engine

[0527] The server uses an emotion engine to recognize the user's emotions, for example, by analyzing the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust).

[0528] The server then uses data from the emotion engine to suggest modifications to the ad script, for example, to include a more positive message or image if the user expresses dissatisfaction.

[0529] The server uses the user's emotion data to select video elements that correspond to the user's preferences and emotions when generating advertising videos, and also selects audio and music that correspond to the user's emotions and preferences.

[0530] 7. Final Check and Correction Steps

[0531] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[0532] 8. Ad Output Steps

[0533] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0534] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0535] This system not only speeds up and streamlines the ad creation process, but also utilizes an emotion engine to provide high-quality ads based on user emotions. As a specific example, an ad for a new smoothie product can be generated in a short time in a way that satisfies users, and is ready for distribution.

[0536] The processing flow will be explained below.

[0537] Step 1:

[0538] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0539] Step 2:

[0540] Based on the received advertising theme, the server refers to the database for past successful advertising data and market trends, and obtains initial settings such as target audience, advertising length, and main message. For example, the target audience may be set to "young people," the advertising length to "30 seconds," and the main message to "refreshment and health."

[0541] Step 3:

[0542] The server uses the initial settings and ad theme to automatically generate an ad script using AI generation. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The server then sends the generated script to the device.

[0543] Step 4:

[0544] The user checks the ad script generated on the device. The user inputs corrections or additional instructions for the script, which the device then sends to the server. For example, the user can send feedback such as, "I'd like it to be a bit more lively and casual."

[0545] Step 5:

[0546] The server uses the generative AI to generate advertising footage based on the revised script. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The server then sends the generated footage to the device, where the user can view it.

[0547] Step 6:

[0548] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0549] Step 7:

[0550] The user checks the audio and video on the device and inputs feedback and correction instructions. For example, the user may give feedback such as "Please speak the narration a little more slowly." The device then sends the input feedback to the server.

[0551] Step 8:

[0552] The server uses an emotion engine to recognize the user's emotions, for example, by analyzing the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust).

[0553] Step 9:

[0554] The server then suggests modifications to the ad script based on the emotion engine data, for example, suggesting including more positive messaging or images if the user expresses dissatisfaction.

[0555] Step 10:

[0556] The server uses the user's emotional data to select video elements that reflect the user's preferences and emotions when generating advertising videos. For example, it increases the number of scenes with bright colors and smiling faces. The server also selects audio and music based on the user's emotions and preferences.

[0557] Step 11:

[0558] The server regenerates the revised ad and sends it to the device, repeating this process until the user is satisfied with the final version.

[0559] Step 12:

[0560] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0561] Step 13:

[0562] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0563] Example 2

[0564] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0565] The ad creation process is typically time-consuming and resource-intensive. Furthermore, the quality of the ad may not perfectly match user emotions and preferences. Traditional methods have difficulty incorporating user feedback in real time, especially in generating ad scripts and audio and video. Furthermore, creating ads that reflect user emotions and preferences is difficult, making it challenging to quickly create high-quality, effective ads.

[0566] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0567] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generation AI, means for presenting the generated advertising script to a user and receiving correction instructions, means for generating an advertising video based on the correction instructions, means for generating appropriate voice and music and adding them to the advertising video, means for recognizing the user's emotional state and analyzing the feedback, means for regenerating the advertising script and advertising video based on the emotional state, and means for exporting the final advertisement and sending it to a distribution platform. This streamlines the advertising creation process and enables the rapid generation of high-quality advertisements based on users' emotions and preferences.

[0568] "User" refers to a user who uses the advertisement creation system.

[0569] "Advertising theme" refers to keywords or phrases that outline the content or purpose of an advertisement.

[0570] "Past advertising data" refers to information about advertisements that have been produced and distributed to date, including data such as effectiveness measurement results and creative elements used.

[0571] "Market trends" refers to information that shows current trends in consumer preferences, trends, economic conditions, etc.

[0572] "Initial Settings" refers to the basic parameter settings required for creating an ad, including the target audience, ad length, key message, etc.

[0573] "Generative AI" refers to technologies and models that use artificial intelligence to generate specific content or data.

[0574] An "advertising script" refers to a text scenario or script that serves as the basis for advertising video and audio.

[0575] "Modification instructions" refer to instructions for changes or additions to advertising scripts, videos, etc. submitted by users.

[0576] "Advertising footage" refers to visual content used as advertising, including video content.

[0577] "Audio and music" refers to audio elements such as narration, sound effects, and background music that are added to advertising footage.

[0578] "Emotional state" refers to the psychological state judged from the user's facial expression, voice, etc., and includes excitement, satisfaction, disgust, etc.

[0579] "Regeneration" refers to a new generation process that modifies or improves upon the content that was initially generated.

[0580] "Distribution Platform" refers to an online platform or medium for distributing generated advertisements.

[0581] This invention is a system for streamlining the advertisement creation process and reducing costs, and in particular, it uses generative AI to automatically generate advertisement scripts, video, and audio, and further combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[0582] Hardware and Software Use

[0583] This system utilizes multiple hardware and software components. The main hardware components include servers and terminals. The software components include generative AI models (e.g., OpenAI's GPT-4), emotion engines (e.g., Microsoft Azure's Face API), and speech synthesis technologies (e.g., Amazon Polly).

[0584] The server receives the ad theme and obtains initial settings based on past ad data and market trends. It also generates ad scripts using a generative AI model and generates ad videos based on user modifications. It also generates suitable voice and music, recognizes user emotions using an emotion engine, and reflects user feedback.

[0585] The terminal is a device that is directly operated by the user, and is used to input advertising themes, check the generated advertising script and video, and input correction instructions. It also saves the final advertising data sent from the server and prepares it for transmission to the distribution platform.

[0586] Specific actions and examples

[0587] Below is a concrete example of the ad creation process.

[0588] 1. Entering a theme: The user enters "Advertisement for a new smoothie" into the input form on the device.

[0589] 2. Obtaining initial settings: The server obtains the target audience, optimal ad length, key messages, etc. from past advertising data and market trends.

[0590] 3. Script generation: The server inputs the following prompt into the generative AI model:

[0591] Ad theme: Promotion of a new smoothie

[0592] Target audience: Young people

[0593] Ad length: 30 seconds

[0594] Main message: Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy health and deliciousness all at once.

[0595] 4. Checking and modifying the script: The server sends the generated script to the terminal, where the user can check it. For example, the user may input instructions for modification, such as "I want a more moving phrase."

[0596] 5. Video generation: The server uses the generative AI to generate advertising videos based on the revised script. For example, it generates a scene of energetic young people enjoying themselves on the beach while drinking smoothies.

[0597] 6. Adding audio and music: The server adds a narration saying "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music to the advertising video.

[0598] 7. Feedback using emotion engine: The server uses an emotion engine to recognize the user's emotional state from their facial expressions and tone of voice, and suggests modifications to the advertisement.

[0599] 8. Final review and output: The server exports the final ad in the appropriate format (e.g., MP4 file) and sends it to the device. The user then performs a final review and sends it to the distribution platform.

[0600] This system not only makes the ad creation process fast and efficient, but also uses an emotion engine to provide high-quality ads based on user emotions. As a specific example, an ad for a new smoothie product can be generated in a short time in a way that satisfies users and is ready to be distributed.

[0601] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0602] Step 1:

[0603] Entering the theme

[0604] The user operates the screen of the device to input the advertising theme. Specifically, the user enters "Advertisement for the newly released smoothie" into the input form.

[0605] The device sends the entered advertising theme to the server in the form of an HTTP POST request. The input data is "Advertising theme: Promotion of newly released smoothie."

[0606] Step 2:

[0607] Get initial configuration

[0608] The server analyzes the received advertisement theme and extracts keywords related to the theme (e.g., "new release," "smoothie," "advertisement").

[0609] The server uses these keywords to retrieve data on past successful ads and market trends from a database. Specific initial settings include target audience, ad length, key message, etc.

[0610] The server structures the initial setup data and prepares it as input data for the next step. The output data is information about "target audience: young people, ad length: 30 seconds, main message: health and delicious."

[0611] Step 3:

[0612] Generate scripts

[0613] The server sends the initial settings and advertising theme obtained to the generative AI model as a prompt.

[0614] Specific prompt:

[0615] Ad theme: Promotion of a new smoothie

[0616] Target audience: Young people

[0617] Ad length: 30 seconds

[0618] Main message: Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy health and deliciousness all at once.

[0619] The generative AI model (on the server) generates an ad script based on this prompt, for example, "Refresh yourself this summer with our new fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0620] The server sends the generated script to the terminal, and the output data is the generated ad script.

[0621] Step 4:

[0622] Check and correct the script

[0623] The user checks the advertisement script displayed on the terminal.

[0624] The user inputs corrections or additional instructions for the script, such as "I want a more moving phrase."

[0625] The device retransmits this feedback to the server. The input data is a correction instruction to "make the expression more moving."

[0626] Step 5:

[0627] Video generation

[0628] The server uses a generative AI to generate advertising videos based on the revised script.

[0629] As a specific example, we generate a scene of energetic young people enjoying themselves on the beach while drinking smoothies.

[0630] The server transmits the generated video data to the terminal, where the user confirms it. The output data is the generated advertising video.

[0631] Step 6:

[0632] Adding voice and music

[0633] The server generates and adds appropriate audio (e.g., narration) and music to the advertising video.

[0634] As a specific example, a narration saying "Refresh yourself this summer with our newly released fresh smoothie!" is combined with refreshing background music.

[0635] The server transmits the integrated audio and video as a single advertisement data to the terminal, which the user confirms. The output data is the completed advertisement content.

[0636] Step 7:

[0637] Emotional Engine Feedback

[0638] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust, etc.).

[0639] The server then uses the data from the emotion engine to suggest improvements to the ad script or video. For example, if the user expresses dissatisfaction, it will select a more positive message or image.

[0640] The server regenerates and optimizes the advertising script and video based on the emotional state, and the output data is the advertising script and video corresponding to the emotional state.

[0641] Step 8:

[0642] Final check and output

[0643] The server exports the final ad in the appropriate format (e.g. MP4 file) and sends it to the device.

[0644] The device stores the advertising data and prepares it for transmission to the distribution platform.

[0645] The user uses their device to submit the ad to the distribution platform, which officially publishes the ad. The output data is the final ad file.

[0646] (Application example 2)

[0647] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0648] The traditional ad creation process required a lot of time and money, and it was difficult to effectively reflect user emotions and feedback. Repeated revisions to ad content were particularly time-consuming, potentially reducing the quality of the final ad. Furthermore, there was a lack of a way to quickly and effectively generate ads that responded to the emotions and expectations of today's highly diverse consumers.

[0649] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving an advertising theme input by a user; means for acquiring initial settings based on past advertising data and market trends; means for generating an advertising script using a generation AI; means for presenting the generated advertising script to a user and receiving correction instructions; means for generating an advertising video based on the correction instructions; means for generating appropriate voice and music and adding it to the advertising video; means for analyzing the user's facial expressions and voice and acquiring emotional data; means for re-correcting the advertisement based on the emotional data; and means for exporting the final advertisement and sending it to a distribution platform. This makes it possible to quickly and efficiently advance the advertisement creation process and generate high-quality advertisements that reflect consumer emotions and expectations.

[0650] The "advertising theme" refers to the basic idea or purpose of the content of the advertisement, and is input by the user.

[0651] "Historical Advertising Data" means historical information about advertisements that have been created and distributed, including details of successful and unsuccessful advertisements.

[0652] "Market trends" refers to information about overall market trends, such as consumer preferences and trends, and the actions of competitors.

[0653] "Initial settings" refers to the basic setting information required to create an ad, such as the target audience, key message, and length of the ad.

[0654] "Generative AI" is a technology that uses artificial intelligence to generate text, video, audio, etc., and in this invention is used to generate advertising scripts, video, and audio.

[0655] An "ad script" is the text that makes up the content of an advertisement, which is generated by the generation AI and can be modified by the user.

[0656] "Modification instructions" are instructions given by the user to the advertising script or video, including changes, additions, or deletions of content.

[0657] "Advertising video" refers to video advertising content created by generative AI and is generated based on a script.

[0658] "Audio and music" refers to narration and background music generated to complement the advertising video.

[0659] "Emotion data" is information about the user's emotional state obtained by analyzing changes in the user's facial expressions and voice.

[0660] "Export" refers to the process of converting the final generated advertisement into an appropriate format so that it can be sent to an external distribution platform.

[0661] This invention is a system for streamlining the advertising creation process and reducing costs. Specifically, it uses generative AI to automatically generate advertising scripts, video, and audio, and then combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[0662] 1. Theme input step

[0663] The user inputs an advertising theme into an input form on the terminal. This advertising theme is sent to the server as a specific prompt sentence. For example, the theme is "Advertisement for a new smoothie."

[0664] 2. Initial setting acquisition step

[0665] Based on the received advertising theme, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[0666] 3. Script generation step

[0667] The server uses AI to automatically generate an ad script based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The generated script is sent to the device, where the user can review and edit it.

[0668] 4. Image generation step

[0669] The server uses the modified script to generate advertising footage using AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated footage is then sent to the device for the user to view.

[0670] 5. Speech and Music Generation Steps

[0671] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device. The user can check the audio and video and input feedback and corrections.

[0672] 6. Feedback step by emotion engine

[0673] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust). Based on the emotion engine data, it suggests modifications to the advertising script and video. For example, if the user expresses dissatisfaction, it may include more positive messages and images. The server uses the user's emotion data to select video elements that correspond to the user's preferences and emotions when generating advertising videos. It also selects audio and music based on the user's emotions and preferences.

[0674] 7. Final Check and Correction Steps

[0675] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[0676] 8. Ad Output Steps

[0677] The server exports the final ad in the appropriate format and sends it to the device, which stores the ad and prepares it for transmission to the TV station or other distribution platform, and the user uses the device to submit it to the distribution platform, where it is published.

[0678] Hardware and software used

[0679] Hardware: The device used by the user (smartphone, tablet, HMD, smart glasses, etc.)

[0680] Software: Generative AI models, emotion engines, database management systems, interface software

[0681] Prompt Sentence Examples

[0682] "Please generate an advertising script for our newly released fresh smoothie, targeting health-conscious people in their 20s and 30s. The main message is 'delicious and healthy.'"

[0683] This makes it possible to quickly and efficiently generate high-quality advertisements that reflect user emotions and feedback.

[0684] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0685] Step 1:

[0686] The user inputs an advertising theme into the input form on the terminal. The input advertising theme might be, for example, "Advertising a new smoothie." This theme is sent to the server as a prompt. The input at this stage is the advertising theme, and the output is the theme sent to the server.

[0687] Step 2:

[0688] Based on the received advertising theme, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, key message, etc. The input of this stage is the advertising theme, and the output is the initial setting information.

[0689] Step 3:

[0690] The server uses the generation AI to automatically generate an ad script based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy health and deliciousness all at once." The input at this stage is the initial settings information and ad theme, and the output is the generated ad script.

[0691] Step 4:

[0692] The server sends the generated ad script to the device, where the user can review and modify it. User feedback and modification instructions are sent to the server via the device. The input at this stage is the generated ad script, and the output is modification instructions from the user.

[0693] Step 5:

[0694] The server generates the ad video based on the modification instructions. Using generative AI, it can generate a scene of energetic young men and women enjoying smoothies on the beach, for example. The input at this stage is the modified ad script, and the output is the generated ad video.

[0695] Step 6:

[0696] The server generates appropriate audio and background music to add to the ad video. For example, a narration voice saying "Refresh yourself this summer with our new fresh smoothie!" and refreshing background music are added. The generated audio and video are sent to the terminal, where the user can review them and input feedback and corrections. The input at this stage is the generated ad video, and the output is the ad video with the added audio and music.

[0697] Step 7:

[0698] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust). Based on the emotion data, it suggests modifications to the ad script or video. The input of this stage is the user's emotion data, and the output is the proposed modification of the ad content.

[0699] Step 8:

[0700] The server regenerates the ad based on the user's feedback and sends it to the device. This process is repeated until the user is satisfied with the final version. The input at this stage is suggested revisions based on emotional data, and the output is the final revised version of the ad.

[0701] Step 9:

[0702] The server exports the final ad in the appropriate format and sends it to the device, which stores the ad and prepares it for transmission to the TV station or other distribution platform. The input to this stage is the final ad data, and the output is the ad data ready for distribution.

[0703] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0704] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0705] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0706] [Third embodiment]

[0707] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0708] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0709] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0710] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0711] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0712] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0713] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0714] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0715] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0716] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0717] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0718] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0719] This invention is a system for streamlining the advertisement creation process and reducing costs, specifically, it uses generation AI to automatically generate advertisement scripts, video, and audio. This system mainly operates in cooperation with a server, terminals, and users.

[0720] 1. Theme input step

[0721] The user inputs an advertising theme on the terminal, for example, specifying the theme "advertising for a new smoothie."

[0722] The terminal transmits the entered advertising theme to the server.

[0723] 2. Initial setting acquisition step

[0724] The server analyzes the past successful advertising data and market trends from the received advertising theme to obtain initial settings, such as setting the target audience to "young people," the length of the advertisement to "30 seconds," and the main message to "refreshment and health."

[0725] 3. Script generation step

[0726] The server uses AI to automatically generate ad scripts based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0727] The server sends the generated script to the terminal, where the user checks it and sends corrections or additional instructions to the server as necessary.

[0728] 4. Image generation step

[0729] The server uses the modified script to generate the ad video using AI. For example, a scene of energetic young men and women enjoying smoothies on the beach can be automatically generated.

[0730] The server sends the generated video to the terminal, where the user can view it.

[0731] 5. Speech and Music Generation Steps

[0732] The server generates appropriate audio and background music and adds them to the ad video. For example, a narration saying, "Refresh yourself this summer with our new fresh smoothie!" can be added along with refreshing background music.

[0733] The server combines the generated audio and video and transmits them to the terminal.

[0734] The user checks the audio and video and sends feedback and correction instructions to the server.

[0735] 6. Final check and correction steps

[0736] The server generates new advertisements based on the user's feedback and sends them to the terminal.

[0737] This process is repeated until the user is satisfied with the final version.

[0738] 7. Ad Output Steps

[0739] The server exports the final ad in the appropriate format and sends it to the device.

[0740] The device stores the ad and prepares it for transmission to the distribution platform.

[0741] The user uses the terminal to carry out the transmission procedure to the distribution platform.

[0742] This system makes the ad creation process fast and efficient. The use of virtual talent and animation also reduces the risk of talent misconduct. For example, a high-quality ad for a new smoothie can be generated quickly and ready for distribution to the user's satisfaction.

[0743] The processing flow will be explained below.

[0744] Step 1:

[0745] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0746] Step 2:

[0747] Based on the advertising theme received by the server, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[0748] Step 3:

[0749] Using the initial settings and ad theme acquired by the server, the AI ​​generator automatically generates an ad script. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0750] Step 4:

[0751] The server sends the generated script to the terminal, where the user confirms it. The user inputs corrections or additional instructions for the script, and the terminal sends them to the server.

[0752] Step 5:

[0753] The server uses the modified script to generate an advertising video using generative AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated video is then sent to the device for the user to view.

[0754] Step 6:

[0755] The server generates or selects appropriate audio and background music and adds them to the advertising video. For example, a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0756] Step 7:

[0757] The user checks the audio and video and inputs feedback and correction instructions. The device sends the input feedback to the server.

[0758] Step 8:

[0759] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[0760] Step 9:

[0761] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0762] Step 10:

[0763] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0764] This process streamlines the ad creation process, reducing costs, and using virtual talent also reduces the risk of talent misconduct.

[0765] Example 1

[0766] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0767] The traditional advertising creation process requires a significant amount of time and money, and involves the risk of talent scheduling and scandals. It is also difficult to guarantee the ad's suitability for the target audience and its quality. A system that can solve these issues and generate high-quality ads quickly and efficiently is needed.

[0768] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0769] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generative AI model, means for presenting the generated advertising script to a user and receiving instructions for modification, means for generating an advertising video based on the modified advertising script, means for generating appropriate audio and background music and adding it to the advertising video, means for reflecting user feedback and regenerating the advertising video, and means for exporting the final advertisement in an appropriate data format and transmitting it to a distribution platform. This makes it possible to create high-quality advertisements in a short period of time and efficiently provide content optimized for the target audience.

[0770] "User" means a person who uses the system to input a theme for creating an advertisement and checks and modifies the generated advertisement script and video.

[0771] "Advertising theme" refers to the subject matter of the content or message that you want to convey in your advertisement, and is entered by the user.

[0772] "Past advertising data" refers to data relating to previously created advertisements, including information on factors contributing to their success and their impact on the market.

[0773] "Market trends" are information that indicates current and projected market needs and consumer preferences.

[0774] "Initial settings" are settings that are set as basic conditions and goals when creating an advertisement, including the target audience, length of the advertisement, and main message.

[0775] A "generative AI model" is an algorithm or model that uses artificial intelligence to automatically generate new content from data.

[0776] An "advertising script" is a written expression of the content of an advertising video, and is automatically generated by a generative AI model.

[0777] "Modification instructions" refers to corrections or additional instructions given by the user to the generated advertising script or video.

[0778] "Advertising video" refers to visual content that is automatically generated using a generative AI model and is created based on an advertising script.

[0779] "Audio and background music" refers to the narration and background music (BGM) added to the advertising video, which are automatically generated by a generative AI model.

[0780] "Feedback" refers to opinions and suggestions for improvement regarding advertising scripts and videos that users provide to the server.

[0781] "Data format" refers to a standardized format for storing and communicating digital data in a particular format, an example of which is the MP4 format.

[0782] A "distribution platform" refers to an online service or website that publishes generated advertisements and delivers them to a large number of users.

[0783] This invention is a system for streamlining the advertising creation process and reducing costs. The system uses a generative AI model to automatically generate advertising scripts, video, and audio. The system mainly operates in cooperation with a server, a terminal, and a user.

[0784] First, the user accesses the system using a terminal and logs in on the user authentication screen. Then, they proceed to a screen for entering the advertising theme, where they enter the advertising theme. For example, they might set an advertising theme such as "Promotion of a newly released smoothie." The terminal then sends this entered theme to the server as an HTTP request.

[0785] The server analyzes the received HTTP request and extracts the advertising theme. Then, the server connects to a database and searches and retrieves past advertising data and market trend data. Based on this data, the server sets initial settings such as the target audience (e.g., young people), ad length (e.g., 30 seconds), and main message (e.g., freshness and health). The initial settings are sent to the terminal in a data format (e.g., JSON) and presented to the user.

[0786] The server then uses the initial settings and ad theme to create a prompt for the generative AI model (e.g., OpenAI's GPT-4). An example prompt is as follows:

[0787] "Generate a 30-second advertising script promoting a new smoothie. The target audience is young people, and the key message is 'refreshing and healthy.'"

[0788] By inputting this prompt into the generative AI model, an advertising script is automatically generated. For example, the generated script might be, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The server then sends the generated script to the device and has the user review it. The user can then enter feedback to modify the script as needed, and the script is sent from the device to the server.

[0789] The server receives feedback from the user and uses the generative AI model again to generate advertising footage based on the revised script. For example, it might generate a scene of energetic young people enjoying themselves on the beach while drinking smoothies. The generated footage is then sent to the device for the user to review.

[0790] The server then generates appropriate audio and background music. Using a generative AI model, it generates refreshing background music and narration, adding a narration such as, "Refresh yourself this summer with our new fresh smoothie!" The generated audio and video are then integrated and the final ad video is sent to the device. The user can review it and provide feedback if necessary.

[0791] Based on the user's feedback, the server regenerates the final ad video and exports it in the appropriate data format (e.g., MP4 format). It is then sent to the device for storage and preparation for uploading to the distribution platform. The user then uses their device to log in to the distribution platform and upload the ad, which is then officially distributed.

[0792] This system makes the ad creation process fast and efficient. The use of virtual characters and animations also reduces the risk of celebrity scandals. For example, an ad promoting a new smoothie can be generated quickly and with high quality, and ready for distribution to the user's satisfaction.

[0793] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0794] Step 1:

[0795] A user accesses the system using a terminal and authenticates on the login screen. After authentication, the user proceeds to the ad theme input screen and inputs an ad theme such as "Advertising a new smoothie." The terminal sends the input ad theme to the server via an HTTP request. The input is the ad theme in text format, and the output is the HTTP request passed to the server.

[0796] Step 2:

[0797] The server analyzes the received HTTP request and extracts the advertising theme. The server connects to the database and searches for and obtains past advertising data and market trend data. During this process, the server searches the database for data on successful advertising and information on current market trends, and generates initial settings such as the target audience, advertising length, and main message. For example, the target audience is set to "young people," the advertising length to "30 seconds," and the main message to "refreshment and health." The input is the advertising theme, and the output is the initial data format (JSON).

[0798] Step 3:

[0799] The server creates a prompt for the generative AI model (e.g., GPT-4) based on the initial settings and advertising theme. For example, a prompt such as "Please generate a 30-second advertising script promoting a newly released smoothie. The target audience is young people, and the main message is 'refreshing and healthy.'" is generated. This prompt is input into the generative AI model to generate an advertising script. The input is the initial settings data and advertising theme, and the output is the generated advertising script.

[0800] Step 4:

[0801] The server sends the generated ad script to the terminal and presents it to the user. The user checks the script and inputs corrections or additional instructions as necessary. The terminal sends this feedback to the server as an HTTP request. The input is the generated ad script, and the output is the user's correction instructions (feedback).

[0802] Step 5:

[0803] The server receives feedback from users and generates ad videos using a generative AI model based on the revised ad script. For example, a scene of energetic young people enjoying themselves on the beach while drinking smoothies is automatically generated. The input is the revised ad script, and the output is the generated ad video.

[0804] Step 6:

[0805] The server generates appropriate audio and background music using a generative AI model. For example, refreshing background music and a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" are added. The server then integrates the generated audio and video and sends the completed advertising video to the device. The user reviews the generated audio and video and enters any additional feedback they may have. The input is the generated advertising video and audio, and the output is the integrated final advertising video.

[0806] Step 7:

[0807] The server receives the user's feedback, regenerates the ad video as needed, and exports the final ad in an appropriate data format (e.g., MP4). The generated video is then sent to the device for storage and preparation for uploading to the distribution platform. The process is completed when the user logs into the distribution platform using their device and uploads the ad. The input is the final feedback, and the output is the final exported ad video.

[0808] The above processing steps ensure that the ad creation process is fast and efficient, enabling high-quality ads to be generated and delivered in a short period of time.

[0809] (Application example 1)

[0810] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0811] The traditional ad creation process was often manual, resulting in time-consuming and costly issues. There was also the risk that continued use of the ad would become difficult due to celebrity scandals or contract issues. Furthermore, there were few ways for users to check and edit ad content in real time, making it difficult to efficiently create high-quality ads.

[0812] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0813] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generation AI, means for presenting the generated script and audio on a visual display device in real time, means for receiving confirmation and correction instructions from the user using the visual display device, and means for exporting the final advertisement and sending it to a distribution platform. This makes it possible to streamline the advertising creation process, reducing time and costs, and enabling users to confirm and correct advertising content in real time, enabling high-quality advertisements to be created in a short period of time.

[0814] "User" means any person or entity that uses the System to create and modify Advertisements.

[0815] The "advertising theme" is the central setting for the content and message of the advertisement to be generated, and is information input by the user.

[0816] "Past advertising data" refers to a collection of data on advertisements that have been created and distributed in the past, and is used as reference information when creating advertisements.

[0817] "Market trends" refers to information about current market trends and consumer behavior patterns.

[0818] "Initial Settings" refers to basic setting information for starting the ad creation process, and is obtained based on past ad data and market trends.

[0819] "Generative AI" is an artificial intelligence technology that uses machine learning and deep learning to automatically generate advertising scripts, video, and audio.

[0820] An "advertising script" is a document that describes the specific content and message of an advertisement.

[0821] A "visual display device" is an electronic device that displays the generated advertising script and video to a user in real time. Examples include smart glasses and head-mounted displays.

[0822] A "distribution platform" is an online service or system for distributing generated advertisements.

[0823] "Modification instructions" are instructions for changes or additions made by the user to the generated advertisement script or video.

[0824] A "generated script" is an advertising document automatically generated by the generation AI.

[0825] "Final Ad" means the final ad content that has been modified and adjusted based on user feedback and correction instructions.

[0826] The system for realizing this invention operates in cooperation with three entities: a server, a terminal, and a user.

[0827] First, the user uses the smart glasses to input the advertising theme by voice. This voice input is picked up by the microphone built into the smart glasses and converted into text data using speech recognition software (speech_recognition library). For example, the user may input "Advertisement for new smoothie release."

[0828] Based on the received advertising theme, the server analyzes past advertising data and market trends to obtain initial settings, such as target audience, advertising length, and key message, by referring to an internal database and external market data.

[0829] The server then uses a generative AI (such as OpenAI's GPT-3) to automatically generate an ad script. The generated script is displayed in real time on the smart glasses' display, and the user can make corrections by voice. For example, the following prompt sentence can be input to the generative AI:

[0830] "You are an ad creator. Please create an ad script for the theme 'Promotion of a new smoothie'."

[0831] Based on user feedback, the server again uses the AI ​​to generate a revised ad script, and then uses text-to-speech software (Google Cloud Text-to-Speech) to generate an appropriate narration voice, which can also be heard in real time on the smart glasses.

[0832] Once the final ad script and audio are completed, the server merges them with the video to automatically generate the ad video, which is based on the latest content reviewed and edited by the user through the smart glasses. Finally, the server exports the generated ad in the appropriate format and sends it to the distribution platform.

[0833] In this way, the server, terminal, and user work together, and prompt sentences using generative AI models are used to streamline the ad creation process and quickly generate high-quality ads.

[0834] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0835] Step 1:

[0836] The user wears the smart glasses and inputs the advertising theme by voice. The voice input is acquired through a microphone built into the smart glasses. This voice data is sent to the device and converted into text data using speech recognition software (speech_recognition library) on the device side. This allows the advertising theme to be acquired in text format.

[0837] Step 2:

[0838] The device transmits the converted text data of the advertising theme to the server. Based on the received advertising theme, the server retrieves initial settings from its internal database and market trend data. For example, it sets the target audience, the length of the advertisement, and the main message. This provides basic setting information for creating the advertisement.

[0839] Step 3:

[0840] The server generates an ad script using a generative AI (OpenAI's GPT-3) based on the initial settings and ad theme. Specifically, it inputs the following prompt to the generative AI:

[0841] "You are an ad creator. Please create an ad script for the theme 'Promotion of a new smoothie'."

[0842] The generated ad script is obtained in text format.

[0843] Step 4:

[0844] The generated script is sent from the server to the device, which then displays it in real time on the smart glasses display. The user can check the content and make corrections using voice or gestures. These corrections are then converted to text on the device and sent back to the server.

[0845] Step 5:

[0846] The server uses the AI ​​to modify and regenerate the ad script based on the modification instructions sent by the user. This process is repeated until the user is satisfied. The modified script is obtained in text format and sent back to the device.

[0847] Step 6:

[0848] The server generates narration using speech synthesis software (Google Cloud Text-to-Speech) based on the finalized ad script, and this audio data is sent to the smart glasses via the device, allowing the user to listen to it in real time.

[0849] Step 7:

[0850] Once the final ad script and audio are completed, the server automatically generates the ad video based on them. The generated video data is sent to the device and can be viewed by the user through smart glasses. This process is repeated until the user is satisfied.

[0851] Step 8:

[0852] Finally, the server exports the final ad in the appropriate format and sends it to the distribution platform, where the user-created ad is ready to be distributed online.

[0853] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0854] This invention is a system for streamlining the advertisement creation process and reducing costs. Specifically, it uses generative AI to automatically generate advertisement scripts, video, and audio, and further combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[0855] 1. Theme input step

[0856] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0857] 2. Initial setting acquisition step

[0858] Based on the advertising theme received by the server, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[0859] 3. Script generation step

[0860] Using the initial settings and ad theme acquired by the server, the AI ​​generator automatically generates an ad script. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0861] The server sends the generated script to the terminal, where the user confirms it. The user inputs corrections or additional instructions for the script, and the terminal sends them to the server.

[0862] 4. Image generation step

[0863] The server uses the modified script to generate an advertising video using generative AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated video is then sent to the device for the user to view.

[0864] 5. Speech and Music Generation Steps

[0865] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0866] The user checks the audio and video and inputs feedback and correction instructions. The device sends the input feedback to the server.

[0867] 6. Feedback step by emotion engine

[0868] The server uses an emotion engine to recognize the user's emotions, for example, by analyzing the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust).

[0869] The server then uses data from the emotion engine to suggest modifications to the ad script, for example, to include a more positive message or image if the user expresses dissatisfaction.

[0870] The server uses the user's emotion data to select video elements that correspond to the user's preferences and emotions when generating advertising videos, and also selects audio and music that correspond to the user's emotions and preferences.

[0871] 7. Final Check and Correction Steps

[0872] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[0873] 8. Ad Output Steps

[0874] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0875] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0876] This system not only speeds up and streamlines the ad creation process, but also utilizes an emotion engine to provide high-quality ads based on user emotions. As a specific example, an ad for a new smoothie product can be generated in a short time in a way that satisfies users, and is ready for distribution.

[0877] The processing flow will be explained below.

[0878] Step 1:

[0879] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[0880] Step 2:

[0881] Based on the received advertising theme, the server refers to the database for past successful advertising data and market trends, and obtains initial settings such as target audience, advertising length, and main message. For example, the target audience may be set to "young people," the advertising length to "30 seconds," and the main message to "refreshment and health."

[0882] Step 3:

[0883] The server uses the initial settings and ad theme to automatically generate an ad script using AI generation. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The server then sends the generated script to the device.

[0884] Step 4:

[0885] The user checks the ad script generated on the device. The user inputs corrections or additional instructions for the script, which the device then sends to the server. For example, the user can send feedback such as, "I'd like it to be a bit more lively and casual."

[0886] Step 5:

[0887] The server uses the generative AI to generate advertising footage based on the revised script. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The server then sends the generated footage to the device, where the user can view it.

[0888] Step 6:

[0889] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[0890] Step 7:

[0891] The user checks the audio and video on the device and inputs feedback and correction instructions. For example, the user may give feedback such as "Please speak the narration a little more slowly." The device then sends the input feedback to the server.

[0892] Step 8:

[0893] The server uses an emotion engine to recognize the user's emotions, for example, by analyzing the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust).

[0894] Step 9:

[0895] The server then suggests modifications to the ad script based on the emotion engine data, for example, suggesting including more positive messaging or images if the user expresses dissatisfaction.

[0896] Step 10:

[0897] The server uses the user's emotional data to select video elements that reflect the user's preferences and emotions when generating advertising videos. For example, it increases the number of scenes with bright colors and smiling faces. The server also selects audio and music based on the user's emotions and preferences.

[0898] Step 11:

[0899] The server regenerates the revised ad and sends it to the device, repeating this process until the user is satisfied with the final version.

[0900] Step 12:

[0901] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[0902] Step 13:

[0903] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[0904] Example 2

[0905] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0906] The ad creation process is typically time-consuming and resource-intensive. Furthermore, the quality of the ad may not perfectly match user emotions and preferences. Traditional methods have difficulty incorporating user feedback in real time, especially in generating ad scripts and audio and video. Furthermore, creating ads that reflect user emotions and preferences is difficult, making it challenging to quickly create high-quality, effective ads.

[0907] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0908] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generation AI, means for presenting the generated advertising script to a user and receiving correction instructions, means for generating an advertising video based on the correction instructions, means for generating appropriate voice and music and adding them to the advertising video, means for recognizing the user's emotional state and analyzing the feedback, means for regenerating the advertising script and advertising video based on the emotional state, and means for exporting the final advertisement and sending it to a distribution platform. This streamlines the advertising creation process and enables the rapid generation of high-quality advertisements based on users' emotions and preferences.

[0909] "User" refers to a user who uses the advertisement creation system.

[0910] "Advertising theme" refers to keywords or phrases that outline the content or purpose of an advertisement.

[0911] "Past advertising data" refers to information about advertisements that have been produced and distributed to date, including data such as effectiveness measurement results and creative elements used.

[0912] "Market trends" refers to information that shows current trends in consumer preferences, trends, economic conditions, etc.

[0913] "Initial Settings" refers to the basic parameter settings required for creating an ad, including the target audience, ad length, key message, etc.

[0914] "Generative AI" refers to technologies and models that use artificial intelligence to generate specific content or data.

[0915] An "advertising script" refers to a text scenario or script that serves as the basis for advertising video and audio.

[0916] "Modification instructions" refer to instructions for changes or additions to advertising scripts, videos, etc. submitted by users.

[0917] "Advertising footage" refers to visual content used as advertising, including video content.

[0918] "Audio and music" refers to audio elements such as narration, sound effects, and background music that are added to advertising footage.

[0919] "Emotional state" refers to the psychological state judged from the user's facial expression, voice, etc., and includes excitement, satisfaction, disgust, etc.

[0920] "Regeneration" refers to a new generation process that modifies or improves upon the content that was initially generated.

[0921] "Distribution Platform" refers to an online platform or medium for distributing generated advertisements.

[0922] This invention is a system for streamlining the advertisement creation process and reducing costs, and in particular, it uses generative AI to automatically generate advertisement scripts, video, and audio, and further combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[0923] Hardware and Software Use

[0924] This system utilizes multiple hardware and software components. The main hardware components include servers and terminals. The software components include generative AI models (e.g., OpenAI's GPT-4), emotion engines (e.g., Microsoft Azure's Face API), and speech synthesis technologies (e.g., Amazon Polly).

[0925] The server receives the ad theme and obtains initial settings based on past ad data and market trends. It also generates ad scripts using a generative AI model and generates ad videos based on user modifications. It also generates suitable voice and music, recognizes user emotions using an emotion engine, and reflects user feedback.

[0926] The terminal is a device that is directly operated by the user, and is used to input advertising themes, check the generated advertising script and video, and input correction instructions. It also saves the final advertising data sent from the server and prepares it for transmission to the distribution platform.

[0927] Specific actions and examples

[0928] Below is a concrete example of the ad creation process.

[0929] 1. Entering a theme: The user enters "Advertisement for a new smoothie" into the input form on the device.

[0930] 2. Obtaining initial settings: The server obtains the target audience, optimal ad length, key messages, etc. from past advertising data and market trends.

[0931] 3. Script generation: The server inputs the following prompt into the generative AI model:

[0932] Ad theme: Promotion of a new smoothie

[0933] Target audience: Young people

[0934] Ad length: 30 seconds

[0935] Main message: Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy health and deliciousness all at once.

[0936] 4. Checking and modifying the script: The server sends the generated script to the terminal, where the user can check it. For example, the user may input instructions for modification, such as "I want a more moving phrase."

[0937] 5. Video generation: The server uses the generative AI to generate advertising videos based on the revised script. For example, it generates a scene of energetic young people enjoying themselves on the beach while drinking smoothies.

[0938] 6. Adding audio and music: The server adds a narration saying "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music to the advertising video.

[0939] 7. Feedback using emotion engine: The server uses an emotion engine to recognize the user's emotional state from their facial expressions and tone of voice, and suggests modifications to the advertisement.

[0940] 8. Final review and output: The server exports the final ad in the appropriate format (e.g., MP4 file) and sends it to the device. The user then performs a final review and sends it to the distribution platform.

[0941] This system not only makes the ad creation process fast and efficient, but also uses an emotion engine to provide high-quality ads based on user emotions. As a specific example, an ad for a new smoothie product can be generated in a short time in a way that satisfies users and is ready to be distributed.

[0942] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0943] Step 1:

[0944] Entering the theme

[0945] The user operates the screen of the device to input the advertising theme. Specifically, the user enters "Advertisement for the newly released smoothie" into the input form.

[0946] The device sends the entered advertising theme to the server in the form of an HTTP POST request. The input data is "Advertising theme: Promotion of newly released smoothie."

[0947] Step 2:

[0948] Get initial configuration

[0949] The server analyzes the received advertisement theme and extracts keywords related to the theme (e.g., "new release," "smoothie," "advertisement").

[0950] The server uses these keywords to retrieve data on past successful ads and market trends from a database. Specific initial settings include target audience, ad length, key message, etc.

[0951] The server structures the initial setup data and prepares it as input data for the next step. The output data is information about "target audience: young people, ad length: 30 seconds, main message: health and delicious."

[0952] Step 3:

[0953] Generate scripts

[0954] The server sends the initial settings and advertising theme obtained to the generative AI model as a prompt.

[0955] Specific prompt:

[0956] Ad theme: Promotion of a new smoothie

[0957] Target audience: Young people

[0958] Ad length: 30 seconds

[0959] Main message: Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy health and deliciousness all at once.

[0960] The generative AI model (on the server) generates an ad script based on this prompt, for example, "Refresh yourself this summer with our new fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[0961] The server sends the generated script to the terminal, and the output data is the generated ad script.

[0962] Step 4:

[0963] Check and correct the script

[0964] The user checks the advertisement script displayed on the terminal.

[0965] The user inputs corrections or additional instructions for the script, such as "I want a more moving phrase."

[0966] The device retransmits this feedback to the server. The input data is a correction instruction to "make the expression more moving."

[0967] Step 5:

[0968] Video generation

[0969] The server uses a generative AI to generate advertising videos based on the revised script.

[0970] As a specific example, we generate a scene of energetic young people enjoying themselves on the beach while drinking smoothies.

[0971] The server transmits the generated video data to the terminal, where the user confirms it. The output data is the generated advertising video.

[0972] Step 6:

[0973] Adding voice and music

[0974] The server generates and adds appropriate audio (e.g., narration) and music to the advertising video.

[0975] As a specific example, a narration saying "Refresh yourself this summer with our newly released fresh smoothie!" is combined with refreshing background music.

[0976] The server transmits the integrated audio and video as a single advertisement data to the terminal, which the user confirms. The output data is the completed advertisement content.

[0977] Step 7:

[0978] Emotional Engine Feedback

[0979] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust, etc.).

[0980] The server then uses the data from the emotion engine to suggest improvements to the ad script or video. For example, if the user expresses dissatisfaction, it will select a more positive message or image.

[0981] The server regenerates and optimizes the advertising script and video based on the emotional state, and the output data is the advertising script and video corresponding to the emotional state.

[0982] Step 8:

[0983] Final check and output

[0984] The server exports the final ad in the appropriate format (e.g. MP4 file) and sends it to the device.

[0985] The device stores the advertising data and prepares it for transmission to the distribution platform.

[0986] The user uses their device to submit the ad to the distribution platform, which officially publishes the ad. The output data is the final ad file.

[0987] (Application example 2)

[0988] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0989] The traditional ad creation process required a lot of time and money, and it was difficult to effectively reflect user emotions and feedback. Repeated revisions to ad content were particularly time-consuming, potentially reducing the quality of the final ad. Furthermore, there was a lack of a way to quickly and effectively generate ads that responded to the emotions and expectations of today's highly diverse consumers.

[0990] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving an advertising theme input by a user; means for acquiring initial settings based on past advertising data and market trends; means for generating an advertising script using a generation AI; means for presenting the generated advertising script to a user and receiving correction instructions; means for generating an advertising video based on the correction instructions; means for generating appropriate voice and music and adding it to the advertising video; means for analyzing the user's facial expressions and voice and acquiring emotional data; means for re-correcting the advertisement based on the emotional data; and means for exporting the final advertisement and sending it to a distribution platform. This makes it possible to quickly and efficiently advance the advertisement creation process and generate high-quality advertisements that reflect consumer emotions and expectations.

[0991] The "advertising theme" refers to the basic idea or purpose of the content of the advertisement, and is input by the user.

[0992] "Historical Advertising Data" means historical information about advertisements that have been created and distributed, including details of successful and unsuccessful advertisements.

[0993] "Market trends" refers to information about overall market trends, such as consumer preferences and trends, and the actions of competitors.

[0994] "Initial settings" refers to the basic setting information required to create an ad, such as the target audience, key message, and length of the ad.

[0995] "Generative AI" is a technology that uses artificial intelligence to generate text, video, audio, etc., and in this invention is used to generate advertising scripts, video, and audio.

[0996] An "ad script" is the text that makes up the content of an advertisement, which is generated by the generation AI and can be modified by the user.

[0997] "Modification instructions" are instructions given by the user to the advertising script or video, including changes, additions, or deletions of content.

[0998] "Advertising video" refers to video advertising content created by generative AI and is generated based on a script.

[0999] "Audio and music" refers to narration and background music generated to complement the advertising video.

[1000] "Emotion data" is information about the user's emotional state obtained by analyzing changes in the user's facial expressions and voice.

[1001] "Export" refers to the process of converting the final generated advertisement into an appropriate format so that it can be sent to an external distribution platform.

[1002] This invention is a system for streamlining the advertising creation process and reducing costs. Specifically, it uses generative AI to automatically generate advertising scripts, video, and audio, and then combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[1003] 1. Theme input step

[1004] The user inputs an advertising theme into an input form on the terminal. This advertising theme is sent to the server as a specific prompt sentence. For example, the theme is "Advertisement for a new smoothie."

[1005] 2. Initial setting acquisition step

[1006] Based on the received advertising theme, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[1007] 3. Script generation step

[1008] The server uses AI to automatically generate an ad script based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The generated script is sent to the device, where the user can review and edit it.

[1009] 4. Image generation step

[1010] The server uses the modified script to generate advertising footage using AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated footage is then sent to the device for the user to view.

[1011] 5. Speech and Music Generation Steps

[1012] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device. The user can check the audio and video and input feedback and corrections.

[1013] 6. Feedback step by emotion engine

[1014] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust). Based on the emotion engine data, it suggests modifications to the advertising script and video. For example, if the user expresses dissatisfaction, it may include more positive messages and images. The server uses the user's emotion data to select video elements that correspond to the user's preferences and emotions when generating advertising videos. It also selects audio and music based on the user's emotions and preferences.

[1015] 7. Final Check and Correction Steps

[1016] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[1017] 8. Ad Output Steps

[1018] The server exports the final ad in the appropriate format and sends it to the device, which stores the ad and prepares it for transmission to the TV station or other distribution platform, and the user uses the device to submit it to the distribution platform, where it is published.

[1019] Hardware and software used

[1020] Hardware: The device used by the user (smartphone, tablet, HMD, smart glasses, etc.)

[1021] Software: Generative AI models, emotion engines, database management systems, interface software

[1022] Prompt Sentence Examples

[1023] "Please generate an advertising script for our newly released fresh smoothie, targeting health-conscious people in their 20s and 30s. The main message is 'delicious and healthy.'"

[1024] This makes it possible to quickly and efficiently generate high-quality advertisements that reflect user emotions and feedback.

[1025] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1026] Step 1:

[1027] The user inputs an advertising theme into the input form on the terminal. The input advertising theme might be, for example, "Advertising a new smoothie." This theme is sent to the server as a prompt. The input at this stage is the advertising theme, and the output is the theme sent to the server.

[1028] Step 2:

[1029] Based on the received advertising theme, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, key message, etc. The input of this stage is the advertising theme, and the output is the initial setting information.

[1030] Step 3:

[1031] The server uses the generation AI to automatically generate an ad script based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy health and deliciousness all at once." The input at this stage is the initial settings information and ad theme, and the output is the generated ad script.

[1032] Step 4:

[1033] The server sends the generated ad script to the device, where the user can review and modify it. User feedback and modification instructions are sent to the server via the device. The input at this stage is the generated ad script, and the output is modification instructions from the user.

[1034] Step 5:

[1035] The server generates the ad video based on the modification instructions. Using generative AI, it can generate a scene of energetic young men and women enjoying smoothies on the beach, for example. The input at this stage is the modified ad script, and the output is the generated ad video.

[1036] Step 6:

[1037] The server generates appropriate audio and background music to add to the ad video. For example, a narration voice saying "Refresh yourself this summer with our new fresh smoothie!" and refreshing background music are added. The generated audio and video are sent to the terminal, where the user can review them and input feedback and corrections. The input at this stage is the generated ad video, and the output is the ad video with the added audio and music.

[1038] Step 7:

[1039] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust). Based on the emotion data, it suggests modifications to the ad script or video. The input of this stage is the user's emotion data, and the output is the proposed modification of the ad content.

[1040] Step 8:

[1041] The server regenerates the ad based on the user's feedback and sends it to the device. This process is repeated until the user is satisfied with the final version. The input at this stage is suggested revisions based on emotional data, and the output is the final revised version of the ad.

[1042] Step 9:

[1043] The server exports the final ad in the appropriate format and sends it to the device, which stores the ad and prepares it for transmission to the TV station or other distribution platform. The input to this stage is the final ad data, and the output is the ad data ready for distribution.

[1044] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1045] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1046] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1047] [Fourth embodiment]

[1048] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1049] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1050] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1051] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1052] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1053] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1054] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1055] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1056] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1057] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1058] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1059] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1060] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1061] This invention is a system for streamlining the advertisement creation process and reducing costs, specifically, it uses generation AI to automatically generate advertisement scripts, video, and audio. This system mainly operates in cooperation with a server, terminals, and users.

[1062] 1. Theme input step

[1063] The user inputs an advertising theme on the terminal, for example, specifying the theme "advertising for a new smoothie."

[1064] The terminal transmits the entered advertising theme to the server.

[1065] 2. Initial setting acquisition step

[1066] The server analyzes the past successful advertising data and market trends from the received advertising theme to obtain initial settings, such as setting the target audience to "young people," the length of the advertisement to "30 seconds," and the main message to "refreshment and health."

[1067] 3. Script generation step

[1068] The server uses AI to automatically generate ad scripts based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[1069] The server sends the generated script to the terminal, where the user checks it and sends corrections or additional instructions to the server as necessary.

[1070] 4. Image generation step

[1071] The server uses the modified script to generate the ad video using AI. For example, a scene of energetic young men and women enjoying smoothies on the beach can be automatically generated.

[1072] The server sends the generated video to the terminal, where the user can view it.

[1073] 5. Speech and Music Generation Steps

[1074] The server generates appropriate audio and background music and adds them to the ad video. For example, a narration saying, "Refresh yourself this summer with our new fresh smoothie!" can be added along with refreshing background music.

[1075] The server combines the generated audio and video and transmits them to the terminal.

[1076] The user checks the audio and video and sends feedback and correction instructions to the server.

[1077] 6. Final check and correction steps

[1078] The server generates new advertisements based on the user's feedback and sends them to the terminal.

[1079] This process is repeated until the user is satisfied with the final version.

[1080] 7. Ad Output Steps

[1081] The server exports the final ad in the appropriate format and sends it to the device.

[1082] The device stores the ad and prepares it for transmission to the distribution platform.

[1083] The user uses the terminal to carry out the transmission procedure to the distribution platform.

[1084] This system makes the ad creation process fast and efficient. The use of virtual talent and animation also reduces the risk of talent misconduct. For example, a high-quality ad for a new smoothie can be generated quickly and ready for distribution to the user's satisfaction.

[1085] The processing flow will be explained below.

[1086] Step 1:

[1087] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[1088] Step 2:

[1089] Based on the advertising theme received by the server, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[1090] Step 3:

[1091] Using the initial settings and ad theme acquired by the server, the AI ​​generator automatically generates an ad script. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[1092] Step 4:

[1093] The server sends the generated script to the terminal, where the user confirms it. The user inputs corrections or additional instructions for the script, and the terminal sends them to the server.

[1094] Step 5:

[1095] The server uses the modified script to generate an advertising video using generative AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated video is then sent to the device for the user to view.

[1096] Step 6:

[1097] The server generates or selects appropriate audio and background music and adds them to the advertising video. For example, a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[1098] Step 7:

[1099] The user checks the audio and video and inputs feedback and correction instructions. The device sends the input feedback to the server.

[1100] Step 8:

[1101] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[1102] Step 9:

[1103] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[1104] Step 10:

[1105] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[1106] This process streamlines the ad creation process, reducing costs, and using virtual talent also reduces the risk of talent misconduct.

[1107] Example 1

[1108] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1109] The traditional advertising creation process requires a significant amount of time and money, and involves the risk of talent scheduling and scandals. It is also difficult to guarantee the ad's suitability for the target audience and its quality. A system that can solve these issues and generate high-quality ads quickly and efficiently is needed.

[1110] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1111] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generative AI model, means for presenting the generated advertising script to a user and receiving instructions for modification, means for generating an advertising video based on the modified advertising script, means for generating appropriate audio and background music and adding it to the advertising video, means for reflecting user feedback and regenerating the advertising video, and means for exporting the final advertisement in an appropriate data format and transmitting it to a distribution platform. This makes it possible to create high-quality advertisements in a short period of time and efficiently provide content optimized for the target audience.

[1112] "User" means a person who uses the system to input a theme for creating an advertisement and checks and modifies the generated advertisement script and video.

[1113] "Advertising theme" refers to the subject matter of the content or message that you want to convey in your advertisement, and is entered by the user.

[1114] "Past advertising data" refers to data relating to previously created advertisements, including information on factors contributing to their success and their impact on the market.

[1115] "Market trends" are information that indicates current and projected market needs and consumer preferences.

[1116] "Initial settings" are settings that are set as basic conditions and goals when creating an advertisement, including the target audience, length of the advertisement, and main message.

[1117] A "generative AI model" is an algorithm or model that uses artificial intelligence to automatically generate new content from data.

[1118] An "advertising script" is a written expression of the content of an advertising video, and is automatically generated by a generative AI model.

[1119] "Modification instructions" refers to corrections or additional instructions given by the user to the generated advertising script or video.

[1120] "Advertising video" refers to visual content that is automatically generated using a generative AI model and is created based on an advertising script.

[1121] "Audio and background music" refers to the narration and background music (BGM) added to the advertising video, which are automatically generated by a generative AI model.

[1122] "Feedback" refers to opinions and suggestions for improvement regarding advertising scripts and videos that users provide to the server.

[1123] "Data format" refers to a standardized format for storing and communicating digital data in a particular format, an example of which is the MP4 format.

[1124] A "distribution platform" refers to an online service or website that publishes generated advertisements and delivers them to a large number of users.

[1125] This invention is a system for streamlining the advertising creation process and reducing costs. The system uses a generative AI model to automatically generate advertising scripts, video, and audio. The system mainly operates in cooperation with a server, a terminal, and a user.

[1126] First, the user accesses the system using a terminal and logs in on the user authentication screen. Then, they proceed to a screen for entering the advertising theme, where they enter the advertising theme. For example, they might set an advertising theme such as "Promotion of a newly released smoothie." The terminal then sends this entered theme to the server as an HTTP request.

[1127] The server analyzes the received HTTP request and extracts the advertising theme. Then, the server connects to a database and searches and retrieves past advertising data and market trend data. Based on this data, the server sets initial settings such as the target audience (e.g., young people), ad length (e.g., 30 seconds), and main message (e.g., freshness and health). The initial settings are sent to the terminal in a data format (e.g., JSON) and presented to the user.

[1128] The server then uses the initial settings and ad theme to create a prompt for the generative AI model (e.g., OpenAI's GPT-4). An example prompt is as follows:

[1129] "Generate a 30-second advertising script promoting a new smoothie. The target audience is young people, and the key message is 'refreshing and healthy.'"

[1130] By inputting this prompt into the generative AI model, an advertising script is automatically generated. For example, the generated script might be, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The server then sends the generated script to the device and has the user review it. The user can then enter feedback to modify the script as needed, and the script is sent from the device to the server.

[1131] The server receives feedback from the user and uses the generative AI model again to generate advertising footage based on the revised script. For example, it might generate a scene of energetic young people enjoying themselves on the beach while drinking smoothies. The generated footage is then sent to the device for the user to review.

[1132] The server then generates appropriate audio and background music. Using a generative AI model, it generates refreshing background music and narration, adding a narration such as, "Refresh yourself this summer with our new fresh smoothie!" The generated audio and video are then integrated and the final ad video is sent to the device. The user can review it and provide feedback if necessary.

[1133] Based on the user's feedback, the server regenerates the final ad video and exports it in the appropriate data format (e.g., MP4 format). It is then sent to the device for storage and preparation for uploading to the distribution platform. The user then uses their device to log in to the distribution platform and upload the ad, which is then officially distributed.

[1134] This system makes the ad creation process fast and efficient. The use of virtual characters and animations also reduces the risk of celebrity scandals. For example, an ad promoting a new smoothie can be generated quickly and with high quality, and ready for distribution to the user's satisfaction.

[1135] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1136] Step 1:

[1137] A user accesses the system using a terminal and authenticates on the login screen. After authentication, the user proceeds to the ad theme input screen and inputs an ad theme such as "Advertising a new smoothie." The terminal sends the input ad theme to the server via an HTTP request. The input is the ad theme in text format, and the output is the HTTP request passed to the server.

[1138] Step 2:

[1139] The server analyzes the received HTTP request and extracts the advertising theme. The server connects to the database and searches for and obtains past advertising data and market trend data. During this process, the server searches the database for data on successful advertising and information on current market trends, and generates initial settings such as the target audience, advertising length, and main message. For example, the target audience is set to "young people," the advertising length to "30 seconds," and the main message to "refreshment and health." The input is the advertising theme, and the output is the initial data format (JSON).

[1140] Step 3:

[1141] The server creates a prompt for the generative AI model (e.g., GPT-4) based on the initial settings and advertising theme. For example, a prompt such as "Please generate a 30-second advertising script promoting a newly released smoothie. The target audience is young people, and the main message is 'refreshing and healthy.'" is generated. This prompt is input into the generative AI model to generate an advertising script. The input is the initial settings data and advertising theme, and the output is the generated advertising script.

[1142] Step 4:

[1143] The server sends the generated ad script to the terminal and presents it to the user. The user checks the script and inputs corrections or additional instructions as necessary. The terminal sends this feedback to the server as an HTTP request. The input is the generated ad script, and the output is the user's correction instructions (feedback).

[1144] Step 5:

[1145] The server receives feedback from users and generates ad videos using a generative AI model based on the revised ad script. For example, a scene of energetic young people enjoying themselves on the beach while drinking smoothies is automatically generated. The input is the revised ad script, and the output is the generated ad video.

[1146] Step 6:

[1147] The server generates appropriate audio and background music using a generative AI model. For example, refreshing background music and a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" are added. The server then integrates the generated audio and video and sends the completed advertising video to the device. The user reviews the generated audio and video and enters any additional feedback they may have. The input is the generated advertising video and audio, and the output is the integrated final advertising video.

[1148] Step 7:

[1149] The server receives the user's feedback, regenerates the ad video as needed, and exports the final ad in an appropriate data format (e.g., MP4). The generated video is then sent to the device for storage and preparation for uploading to the distribution platform. The process is completed when the user logs into the distribution platform using their device and uploads the ad. The input is the final feedback, and the output is the final exported ad video.

[1150] The above processing steps ensure that the ad creation process is fast and efficient, enabling high-quality ads to be generated and delivered in a short period of time.

[1151] (Application example 1)

[1152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1153] The traditional ad creation process was often manual, resulting in time-consuming and costly issues. There was also the risk that continued use of the ad would become difficult due to celebrity scandals or contract issues. Furthermore, there were few ways for users to check and edit ad content in real time, making it difficult to efficiently create high-quality ads.

[1154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1155] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generation AI, means for presenting the generated script and audio on a visual display device in real time, means for receiving confirmation and correction instructions from the user using the visual display device, and means for exporting the final advertisement and sending it to a distribution platform. This makes it possible to streamline the advertising creation process, reducing time and costs, and enabling users to confirm and correct advertising content in real time, enabling high-quality advertisements to be created in a short period of time.

[1156] "User" means any person or entity that uses the System to create and modify Advertisements.

[1157] The "advertising theme" is the central setting for the content and message of the advertisement to be generated, and is information input by the user.

[1158] "Past advertising data" refers to a collection of data on advertisements that have been created and distributed in the past, and is used as reference information when creating advertisements.

[1159] "Market trends" refers to information about current market trends and consumer behavior patterns.

[1160] "Initial Settings" refers to basic setting information for starting the ad creation process, and is obtained based on past ad data and market trends.

[1161] "Generative AI" is an artificial intelligence technology that uses machine learning and deep learning to automatically generate advertising scripts, video, and audio.

[1162] An "advertising script" is a document that describes the specific content and message of an advertisement.

[1163] A "visual display device" is an electronic device that displays the generated advertising script and video to a user in real time. Examples include smart glasses and head-mounted displays.

[1164] A "distribution platform" is an online service or system for distributing generated advertisements.

[1165] "Modification instructions" are instructions for changes or additions made by the user to the generated advertisement script or video.

[1166] A "generated script" is an advertising document automatically generated by the generation AI.

[1167] "Final Ad" means the final ad content that has been modified and adjusted based on user feedback and correction instructions.

[1168] The system for realizing this invention operates in cooperation with three entities: a server, a terminal, and a user.

[1169] First, the user uses the smart glasses to input the advertising theme by voice. This voice input is picked up by the microphone built into the smart glasses and converted into text data using speech recognition software (speech_recognition library). For example, the user may input "Advertisement for new smoothie release."

[1170] Based on the received advertising theme, the server analyzes past advertising data and market trends to obtain initial settings, such as target audience, advertising length, and key message, by referring to an internal database and external market data.

[1171] The server then uses a generative AI (such as OpenAI's GPT-3) to automatically generate an ad script. The generated script is displayed in real time on the smart glasses' display, and the user can make corrections by voice. For example, the following prompt sentence can be input to the generative AI:

[1172] "You are an ad creator. Please create an ad script for the theme 'Promotion of a new smoothie'."

[1173] Based on user feedback, the server again uses the AI ​​to generate a revised ad script, and then uses text-to-speech software (Google Cloud Text-to-Speech) to generate an appropriate narration voice, which can also be heard in real time on the smart glasses.

[1174] Once the final ad script and audio are completed, the server merges them with the video to automatically generate the ad video, which is based on the latest content reviewed and edited by the user through the smart glasses. Finally, the server exports the generated ad in the appropriate format and sends it to the distribution platform.

[1175] In this way, the server, terminal, and user work together, and prompt sentences using generative AI models are used to streamline the ad creation process and quickly generate high-quality ads.

[1176] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1177] Step 1:

[1178] The user wears the smart glasses and inputs the advertising theme by voice. The voice input is acquired through a microphone built into the smart glasses. This voice data is sent to the device and converted into text data using speech recognition software (speech_recognition library) on the device side. This allows the advertising theme to be acquired in text format.

[1179] Step 2:

[1180] The device transmits the converted text data of the advertising theme to the server. Based on the received advertising theme, the server retrieves initial settings from its internal database and market trend data. For example, it sets the target audience, the length of the advertisement, and the main message. This provides basic setting information for creating the advertisement.

[1181] Step 3:

[1182] The server generates an ad script using a generative AI (OpenAI's GPT-3) based on the initial settings and ad theme. Specifically, it inputs the following prompt to the generative AI:

[1183] "You are an ad creator. Please create an ad script for the theme 'Promotion of a new smoothie'."

[1184] The generated ad script is obtained in text format.

[1185] Step 4:

[1186] The generated script is sent from the server to the device, which then displays it in real time on the smart glasses display. The user can check the content and make corrections using voice or gestures. These corrections are then converted to text on the device and sent back to the server.

[1187] Step 5:

[1188] The server uses the AI ​​to modify and regenerate the ad script based on the modification instructions sent by the user. This process is repeated until the user is satisfied. The modified script is obtained in text format and sent back to the device.

[1189] Step 6:

[1190] The server generates narration using speech synthesis software (Google Cloud Text-to-Speech) based on the finalized ad script, and this audio data is sent to the smart glasses via the device, allowing the user to listen to it in real time.

[1191] Step 7:

[1192] Once the final ad script and audio are completed, the server automatically generates the ad video based on them. The generated video data is sent to the device and can be viewed by the user through smart glasses. This process is repeated until the user is satisfied.

[1193] Step 8:

[1194] Finally, the server exports the final ad in the appropriate format and sends it to the distribution platform, where the user-created ad is ready to be distributed online.

[1195] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1196] This invention is a system for streamlining the advertisement creation process and reducing costs. Specifically, it uses generative AI to automatically generate advertisement scripts, video, and audio, and further combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[1197] 1. Theme input step

[1198] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[1199] 2. Initial setting acquisition step

[1200] Based on the advertising theme received by the server, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[1201] 3. Script generation step

[1202] Using the initial settings and ad theme acquired by the server, the AI ​​generator automatically generates an ad script. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[1203] The server sends the generated script to the terminal, where the user confirms it. The user inputs corrections or additional instructions for the script, and the terminal sends them to the server.

[1204] 4. Image generation step

[1205] The server uses the modified script to generate an advertising video using generative AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated video is then sent to the device for the user to view.

[1206] 5. Speech and Music Generation Steps

[1207] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[1208] The user checks the audio and video and inputs feedback and correction instructions. The device sends the input feedback to the server.

[1209] 6. Feedback step by emotion engine

[1210] The server uses an emotion engine to recognize the user's emotions, for example, by analyzing the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust).

[1211] The server then uses data from the emotion engine to suggest modifications to the ad script, for example, to include a more positive message or image if the user expresses dissatisfaction.

[1212] The server uses the user's emotion data to select video elements that correspond to the user's preferences and emotions when generating advertising videos, and also selects audio and music that correspond to the user's emotions and preferences.

[1213] 7. Final Check and Correction Steps

[1214] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[1215] 8. Ad Output Steps

[1216] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[1217] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[1218] This system not only speeds up and streamlines the ad creation process, but also utilizes an emotion engine to provide high-quality ads based on user emotions. As a specific example, an ad for a new smoothie product can be generated in a short time in a way that satisfies users, and is ready for distribution.

[1219] The processing flow will be explained below.

[1220] Step 1:

[1221] The user inputs an advertising theme into the input form on the terminal. For example, the user selects the theme "Advertisement for Newly Released Smoothie." The terminal transmits the input advertising theme to the server.

[1222] Step 2:

[1223] Based on the received advertising theme, the server refers to the database for past successful advertising data and market trends, and obtains initial settings such as target audience, advertising length, and main message. For example, the target audience may be set to "young people," the advertising length to "30 seconds," and the main message to "refreshment and health."

[1224] Step 3:

[1225] The server uses the initial settings and ad theme to automatically generate an ad script using AI generation. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The server then sends the generated script to the device.

[1226] Step 4:

[1227] The user checks the ad script generated on the device. The user inputs corrections or additional instructions for the script, which the device then sends to the server. For example, the user can send feedback such as, "I'd like it to be a bit more lively and casual."

[1228] Step 5:

[1229] The server uses the generative AI to generate advertising footage based on the revised script. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The server then sends the generated footage to the device, where the user can view it.

[1230] Step 6:

[1231] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device.

[1232] Step 7:

[1233] The user checks the audio and video on the device and inputs feedback and correction instructions. For example, the user may give feedback such as "Please speak the narration a little more slowly." The device then sends the input feedback to the server.

[1234] Step 8:

[1235] The server uses an emotion engine to recognize the user's emotions, for example, by analyzing the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust).

[1236] Step 9:

[1237] The server then suggests modifications to the ad script based on the emotion engine data, for example, suggesting including more positive messaging or images if the user expresses dissatisfaction.

[1238] Step 10:

[1239] The server uses the user's emotional data to select video elements that reflect the user's preferences and emotions when generating advertising videos. For example, it increases the number of scenes with bright colors and smiling faces. The server also selects audio and music based on the user's emotions and preferences.

[1240] Step 11:

[1241] The server regenerates the revised ad and sends it to the device, repeating this process until the user is satisfied with the final version.

[1242] Step 12:

[1243] The server exports the final ad in the appropriate format and sends it to the device, which stores it and prepares it for transmission to the TV station or other distribution platform.

[1244] Step 13:

[1245] The user uses their device to carry out the transmission procedure to the distribution platform, which then makes the advertisement public.

[1246] Example 2

[1247] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1248] The ad creation process is typically time-consuming and resource-intensive. Furthermore, the quality of the ad may not perfectly match user emotions and preferences. Traditional methods have difficulty incorporating user feedback in real time, especially in generating ad scripts and audio and video. Furthermore, creating ads that reflect user emotions and preferences is difficult, making it challenging to quickly create high-quality, effective ads.

[1249] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1250] In this invention, the server includes means for receiving an advertising theme input by a user, means for acquiring initial settings based on past advertising data and market trends, means for generating an advertising script using a generation AI, means for presenting the generated advertising script to a user and receiving correction instructions, means for generating an advertising video based on the correction instructions, means for generating appropriate voice and music and adding them to the advertising video, means for recognizing the user's emotional state and analyzing the feedback, means for regenerating the advertising script and advertising video based on the emotional state, and means for exporting the final advertisement and sending it to a distribution platform. This streamlines the advertising creation process and enables the rapid generation of high-quality advertisements based on users' emotions and preferences.

[1251] "User" refers to a user who uses the advertisement creation system.

[1252] "Advertising theme" refers to keywords or phrases that outline the content or purpose of an advertisement.

[1253] "Past advertising data" refers to information about advertisements that have been produced and distributed to date, including data such as effectiveness measurement results and creative elements used.

[1254] "Market trends" refers to information that shows current trends in consumer preferences, trends, economic conditions, etc.

[1255] "Initial Settings" refers to the basic parameter settings required for creating an ad, including the target audience, ad length, key message, etc.

[1256] "Generative AI" refers to technologies and models that use artificial intelligence to generate specific content or data.

[1257] An "advertising script" refers to a text scenario or script that serves as the basis for advertising video and audio.

[1258] "Modification instructions" refer to instructions for changes or additions to advertising scripts, videos, etc. submitted by users.

[1259] "Advertising footage" refers to visual content used as advertising, including video content.

[1260] "Audio and music" refers to audio elements such as narration, sound effects, and background music that are added to advertising footage.

[1261] "Emotional state" refers to the psychological state judged from the user's facial expression, voice, etc., and includes excitement, satisfaction, disgust, etc.

[1262] "Regeneration" refers to a new generation process that modifies or improves upon the content that was initially generated.

[1263] "Distribution Platform" refers to an online platform or medium for distributing generated advertisements.

[1264] This invention is a system for streamlining the advertisement creation process and reducing costs, and in particular, it uses generative AI to automatically generate advertisement scripts, video, and audio, and further combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[1265] Hardware and Software Use

[1266] This system utilizes multiple hardware and software components. The main hardware components include servers and terminals. The software components include generative AI models (e.g., OpenAI's GPT-4), emotion engines (e.g., Microsoft Azure's Face API), and speech synthesis technologies (e.g., Amazon Polly).

[1267] The server receives the ad theme and obtains initial settings based on past ad data and market trends. It also generates ad scripts using a generative AI model and generates ad videos based on user modifications. It also generates suitable voice and music, recognizes user emotions using an emotion engine, and reflects user feedback.

[1268] The terminal is a device that is directly operated by the user, and is used to input advertising themes, check the generated advertising script and video, and input correction instructions. It also saves the final advertising data sent from the server and prepares it for transmission to the distribution platform.

[1269] Specific actions and examples

[1270] Below is a concrete example of the ad creation process.

[1271] 1. Entering a theme: The user enters "Advertisement for a new smoothie" into the input form on the device.

[1272] 2. Obtaining initial settings: The server obtains the target audience, optimal ad length, key messages, etc. from past advertising data and market trends.

[1273] 3. Script generation: The server inputs the following prompt into the generative AI model:

[1274] Ad theme: Promotion of a new smoothie

[1275] Target audience: Young people

[1276] Ad length: 30 seconds

[1277] Main message: Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy health and deliciousness all at once.

[1278] 4. Checking and modifying the script: The server sends the generated script to the terminal, where the user can check it. For example, the user may input instructions for modification, such as "I want a more moving phrase."

[1279] 5. Video generation: The server uses the generative AI to generate advertising videos based on the revised script. For example, it generates a scene of energetic young people enjoying themselves on the beach while drinking smoothies.

[1280] 6. Adding audio and music: The server adds a narration saying "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music to the advertising video.

[1281] 7. Feedback using emotion engine: The server uses an emotion engine to recognize the user's emotional state from their facial expressions and tone of voice, and suggests modifications to the advertisement.

[1282] 8. Final review and output: The server exports the final ad in the appropriate format (e.g., MP4 file) and sends it to the device. The user then performs a final review and sends it to the distribution platform.

[1283] This system not only makes the ad creation process fast and efficient, but also uses an emotion engine to provide high-quality ads based on user emotions. As a specific example, an ad for a new smoothie product can be generated in a short time in a way that satisfies users and is ready to be distributed.

[1284] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1285] Step 1:

[1286] Entering the theme

[1287] The user operates the screen of the device to input the advertising theme. Specifically, the user enters "Advertisement for the newly released smoothie" into the input form.

[1288] The device sends the entered advertising theme to the server in the form of an HTTP POST request. The input data is "Advertising theme: Promotion of newly released smoothie."

[1289] Step 2:

[1290] Get initial configuration

[1291] The server analyzes the received advertisement theme and extracts keywords related to the theme (e.g., "new release," "smoothie," "advertisement").

[1292] The server uses these keywords to retrieve data on past successful ads and market trends from a database. Specific initial settings include target audience, ad length, key message, etc.

[1293] The server structures the initial setup data and prepares it as input data for the next step. The output data is information about "target audience: young people, ad length: 30 seconds, main message: health and delicious."

[1294] Step 3:

[1295] Generate scripts

[1296] The server sends the initial settings and advertising theme obtained to the generative AI model as a prompt.

[1297] Specific prompt:

[1298] Ad theme: Promotion of a new smoothie

[1299] Target audience: Young people

[1300] Ad length: 30 seconds

[1301] Main message: Refresh yourself this summer with our newly released fresh smoothie! We offer a new experience that lets you enjoy health and deliciousness all at once.

[1302] The generative AI model (on the server) generates an ad script based on this prompt, for example, "Refresh yourself this summer with our new fresh smoothie! We offer a new experience that lets you enjoy both health and deliciousness at the same time."

[1303] The server sends the generated script to the terminal, and the output data is the generated ad script.

[1304] Step 4:

[1305] Check and correct the script

[1306] The user checks the advertisement script displayed on the terminal.

[1307] The user inputs corrections or additional instructions for the script, such as "I want a more moving phrase."

[1308] The device retransmits this feedback to the server. The input data is a correction instruction to "make the expression more moving."

[1309] Step 5:

[1310] Video generation

[1311] The server uses a generative AI to generate advertising videos based on the revised script.

[1312] As a specific example, we generate a scene of energetic young people enjoying themselves on the beach while drinking smoothies.

[1313] The server transmits the generated video data to the terminal, where the user confirms it. The output data is the generated advertising video.

[1314] Step 6:

[1315] Adding voice and music

[1316] The server generates and adds appropriate audio (e.g., narration) and music to the advertising video.

[1317] As a specific example, a narration saying "Refresh yourself this summer with our newly released fresh smoothie!" is combined with refreshing background music.

[1318] The server transmits the integrated audio and video as a single advertisement data to the terminal, which the user confirms. The output data is the completed advertisement content.

[1319] Step 7:

[1320] Emotional Engine Feedback

[1321] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust, etc.).

[1322] The server then uses the data from the emotion engine to suggest improvements to the ad script or video. For example, if the user expresses dissatisfaction, it will select a more positive message or image.

[1323] The server regenerates and optimizes the advertising script and video based on the emotional state, and the output data is the advertising script and video corresponding to the emotional state.

[1324] Step 8:

[1325] Final check and output

[1326] The server exports the final ad in the appropriate format (e.g. MP4 file) and sends it to the device.

[1327] The device stores the advertising data and prepares it for transmission to the distribution platform.

[1328] The user uses their device to submit the ad to the distribution platform, which officially publishes the ad. The output data is the final ad file.

[1329] (Application example 2)

[1330] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1331] The traditional ad creation process required a lot of time and money, and it was difficult to effectively reflect user emotions and feedback. Repeated revisions to ad content were particularly time-consuming, potentially reducing the quality of the final ad. Furthermore, there was a lack of a way to quickly and effectively generate ads that responded to the emotions and expectations of today's highly diverse consumers.

[1332] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving an advertising theme input by a user; means for acquiring initial settings based on past advertising data and market trends; means for generating an advertising script using a generation AI; means for presenting the generated advertising script to a user and receiving correction instructions; means for generating an advertising video based on the correction instructions; means for generating appropriate voice and music and adding it to the advertising video; means for analyzing the user's facial expressions and voice and acquiring emotional data; means for re-correcting the advertisement based on the emotional data; and means for exporting the final advertisement and sending it to a distribution platform. This makes it possible to quickly and efficiently advance the advertisement creation process and generate high-quality advertisements that reflect consumer emotions and expectations.

[1333] The "advertising theme" refers to the basic idea or purpose of the content of the advertisement, and is input by the user.

[1334] "Historical Advertising Data" means historical information about advertisements that have been created and distributed, including details of successful and unsuccessful advertisements.

[1335] "Market trends" refers to information about overall market trends, such as consumer preferences and trends, and the actions of competitors.

[1336] "Initial settings" refers to the basic setting information required to create an ad, such as the target audience, key message, and length of the ad.

[1337] "Generative AI" is a technology that uses artificial intelligence to generate text, video, audio, etc., and in this invention is used to generate advertising scripts, video, and audio.

[1338] An "ad script" is the text that makes up the content of an advertisement, which is generated by the generation AI and can be modified by the user.

[1339] "Modification instructions" are instructions given by the user to the advertising script or video, including changes, additions, or deletions of content.

[1340] "Advertising video" refers to video advertising content created by generative AI and is generated based on a script.

[1341] "Audio and music" refers to narration and background music generated to complement the advertising video.

[1342] "Emotion data" is information about the user's emotional state obtained by analyzing changes in the user's facial expressions and voice.

[1343] "Export" refers to the process of converting the final generated advertisement into an appropriate format so that it can be sent to an external distribution platform.

[1344] This invention is a system for streamlining the advertising creation process and reducing costs. Specifically, it uses generative AI to automatically generate advertising scripts, video, and audio, and then combines it with an emotion engine to generate advertisements based on user emotions. This system mainly operates in cooperation with a server, terminals, and users.

[1345] 1. Theme input step

[1346] The user inputs an advertising theme into an input form on the terminal. This advertising theme is sent to the server as a specific prompt sentence. For example, the theme is "Advertisement for a new smoothie."

[1347] 2. Initial setting acquisition step

[1348] Based on the received advertising theme, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, and key messages.

[1349] 3. Script generation step

[1350] The server uses AI to automatically generate an ad script based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy both health and deliciousness at the same time." The generated script is sent to the device, where the user can review and edit it.

[1351] 4. Image generation step

[1352] The server uses the modified script to generate advertising footage using AI. For example, it could generate a scene of energetic young men and women enjoying smoothies on the beach. The generated footage is then sent to the device for the user to view.

[1353] 5. Speech and Music Generation Steps

[1354] The server generates appropriate audio and background music and adds them to the advertising video. For example, a narration voice saying, "Refresh yourself this summer with our newly released fresh smoothie!" and refreshing background music are added. The generated audio and video are integrated and sent to the device. The user can check the audio and video and input feedback and corrections.

[1355] 6. Feedback step by emotion engine

[1356] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust). Based on the emotion engine data, it suggests modifications to the advertising script and video. For example, if the user expresses dissatisfaction, it may include more positive messages and images. The server uses the user's emotion data to select video elements that correspond to the user's preferences and emotions when generating advertising videos. It also selects audio and music based on the user's emotions and preferences.

[1357] 7. Final Check and Correction Steps

[1358] The server regenerates the ad based on the user's feedback and sends it to the device, repeating this process until the user is satisfied with the final version.

[1359] 8. Ad Output Steps

[1360] The server exports the final ad in the appropriate format and sends it to the device, which stores the ad and prepares it for transmission to the TV station or other distribution platform, and the user uses the device to submit it to the distribution platform, where it is published.

[1361] Hardware and software used

[1362] Hardware: The device used by the user (smartphone, tablet, HMD, smart glasses, etc.)

[1363] Software: Generative AI models, emotion engines, database management systems, interface software

[1364] Prompt Sentence Examples

[1365] "Please generate an advertising script for our newly released fresh smoothie, targeting health-conscious people in their 20s and 30s. The main message is 'delicious and healthy.'"

[1366] This makes it possible to quickly and efficiently generate high-quality advertisements that reflect user emotions and feedback.

[1367] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1368] Step 1:

[1369] The user inputs an advertising theme into the input form on the terminal. The input advertising theme might be, for example, "Advertising a new smoothie." This theme is sent to the server as a prompt. The input at this stage is the advertising theme, and the output is the theme sent to the server.

[1370] Step 2:

[1371] Based on the received advertising theme, the server references past successful advertising data and market trends from a database to obtain initial settings such as target audience, advertising length, key message, etc. The input of this stage is the advertising theme, and the output is the initial setting information.

[1372] Step 3:

[1373] The server uses the generation AI to automatically generate an ad script based on the initial settings and ad theme. For example, a script might be generated that reads, "Refresh yourself this summer with our newly released fresh smoothie! We'll bring you a new experience where you can enjoy health and deliciousness all at once." The input at this stage is the initial settings information and ad theme, and the output is the generated ad script.

[1374] Step 4:

[1375] The server sends the generated ad script to the device, where the user can review and modify it. User feedback and modification instructions are sent to the server via the device. The input at this stage is the generated ad script, and the output is modification instructions from the user.

[1376] Step 5:

[1377] The server generates the ad video based on the modification instructions. Using generative AI, it can generate a scene of energetic young men and women enjoying smoothies on the beach, for example. The input at this stage is the modified ad script, and the output is the generated ad video.

[1378] Step 6:

[1379] The server generates appropriate audio and background music to add to the ad video. For example, a narration voice saying "Refresh yourself this summer with our new fresh smoothie!" and refreshing background music are added. The generated audio and video are sent to the terminal, where the user can review them and input feedback and corrections. The input at this stage is the generated ad video, and the output is the ad video with the added audio and music.

[1380] Step 7:

[1381] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to identify their emotional state (excitement, satisfaction, disgust). Based on the emotion data, it suggests modifications to the ad script or video. The input of this stage is the user's emotion data, and the output is the proposed modification of the ad content.

[1382] Step 8:

[1383] The server regenerates the ad based on the user's feedback and sends it to the device. This process is repeated until the user is satisfied with the final version. The input at this stage is suggested revisions based on emotional data, and the output is the final revised version of the ad.

[1384] Step 9:

[1385] The server exports the final ad in the appropriate format and sends it to the device, which stores the ad and prepares it for transmission to the TV station or other distribution platform. The input to this stage is the final ad data, and the output is the ad data ready for distribution.

[1386] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1387] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1388] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1389] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1390] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1391] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1392] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1393] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1394] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1395] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1396] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1397] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1398] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1399] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1400] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1401] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1402] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1403] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1404] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1405] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1406] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1407] The following is further disclosed regarding the above embodiment.

[1408] (Claim 1)

[1409] A means for receiving an advertising theme input by a user;

[1410] a means for obtaining an initial configuration based on historical advertising data and market trends;

[1411] A means for generating an ad script using a generative AI;

[1412] a means for presenting the generated advertisement script to a user and receiving instructions for correction;

[1413] means for generating an advertising video based on the correction instruction;

[1414] means for generating and adding appropriate sounds and music to the advertising video;

[1415] A means to export the final ad and send it to a distribution platform;

[1416] A system including:

[1417] (Claim 2)

[1418] 10. The system of claim 1, further comprising means for using virtual talent and animation in the advertising video.

[1419] (Claim 3)

[1420] 10. The system of claim 1, further comprising means for reflecting feedback from users and re-presenting the regenerated advertisement.

[1421] "Example 1"

[1422] (Claim 1)

[1423] A means for receiving an advertising theme input by a user;

[1424] a means for obtaining an initial configuration based on historical advertising data and market trends;

[1425] a means for generating an ad script using a generative AI model;

[1426] a means for presenting the generated advertisement script to a user and receiving instructions for correction;

[1427] a means for generating an advertisement video based on the modified advertisement script;

[1428] means for generating and adding appropriate audio and background music to the advertising video;

[1429] A means of regenerating advertising videos by reflecting user feedback;

[1430] A means to export the final ad in the appropriate data format and send it to a distribution platform;

[1431] A system including:

[1432] (Claim 2)

[1433] 10. The system of claim 1, further comprising means for using virtual characters and animations in the advertising video.

[1434] (Claim 3)

[1435] 10. The system of claim 1, further comprising means for regenerating the advertising video based on user feedback.

[1436] "Application Example 1"

[1437] (Claim 1)

[1438] A means for receiving an advertising theme input by a user;

[1439] a means for obtaining an initial configuration based on historical advertising data and market trends;

[1440] A means for generating an ad script using a generative AI;

[1441] a means for presenting the generated advertisement script to a user and receiving instructions for correction;

[1442] means for generating an advertising video based on the correction instruction;

[1443] means for generating and adding appropriate sounds and music to the advertising video;

[1444] means for presenting the generated script and audio in real time on a visual display device;

[1445] means for receiving user confirmation and / or correction instructions using a visual display device;

[1446] A means to export the final ad and send it to a distribution platform;

[1447] A system including:

[1448] (Claim 2)

[1449] 10. The system of claim 1, further comprising means for using virtual talent and animation in the advertising video.

[1450] (Claim 3)

[1451] 10. The system of claim 1, further comprising means for reflecting feedback from users and re-presenting the regenerated advertisement.

[1452] "Example 2: Combining Emotion Engines"

[1453] (Claim 1)

[1454] A means for receiving an advertising theme input by a user;

[1455] a means for obtaining an initial configuration based on historical advertising data and market trends;

[1456] A means for generating an ad script using a generative AI;

[1457] a means for presenting the generated advertisement script to a user and receiving instructions for correction;

[1458] means for generating an advertising video based on the correction instruction;

[1459] means for generating and adding appropriate sounds and music to the advertising video;

[1460] a means for recognizing the user's emotional state and analyzing the feedback;

[1461] means for regenerating advertising scripts and advertising videos based on the emotional state;

[1462] A means to export the final ad and send it to a distribution platform;

[1463] A system including:

[1464] (Claim 2)

[1465] 10. The system of claim 1, further comprising means for using virtual talent and animation in the advertising video.

[1466] (Claim 3)

[1467] 10. The system of claim 1, further comprising means for reflecting feedback from users and re-presenting the regenerated advertisement.

[1468] "Application example 2 when combining emotion engines"

[1469] (Claim 1)

[1470] A means for receiving an advertising theme input by a user;

[1471] a means for obtaining an initial configuration based on historical advertising data and market trends;

[1472] A means for generating an ad script using a generative AI;

[1473] a means for presenting the generated advertisement script to a user and receiving instructions for correction;

[1474] means for generating an advertising video based on the correction instruction;

[1475] means for generating and adding appropriate sounds and music to the advertising video;

[1476] A means of analyzing the user's facial expressions and voice to obtain emotional data;

[1477] a means for re-modifying the advertisement based on the sentiment data;

[1478] A means to export the final ad and send it to a distribution platform;

[1479] A system including:

[1480] (Claim 2)

[1481] 10. The system of claim 1, further comprising means for using virtual talent and animation in the advertising video.

[1482] (Claim 3)

[1483] 10. The system of claim 1, further comprising means for reflecting feedback from users and re-presenting the regenerated advertisement. [Explanation of symbols]

[1484] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving an advertising theme input by a user; a means for obtaining an initial configuration based on historical advertising data and market trends; A means for generating an ad script using a generative AI; a means for presenting the generated advertisement script to a user and receiving instructions for correction; means for generating an advertising video based on the correction instruction; means for generating and adding appropriate sounds and music to the advertising video; A means to export the final ad and send it to a distribution platform; A system including:

2. 10. The system of claim 1, further comprising means for using virtual talent and animation in advertising videos.

3. 10. The system of claim 1, further comprising means for reflecting feedback from users and re-presenting the regenerated advertisement.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A