System
A system that analyzes user input text and storyboards to automatically generate videos, addressing the inefficiencies and costs of traditional video production while ensuring proper compensation for material usage.
Patent Information
- Application Number
- JP2024126393
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2026-02-13
Smart Images

Figure 2026024072000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Video explanations are an extremely effective way to intuitively convey information that is difficult to understand through text or illustrations. However, video production requires a lot of time and money, and many companies and individuals find it a high hurdle. This has led to a demand for an easy way to produce high-quality videos. Furthermore, there is also a need for a method that overcomes copyright issues and ensures appropriate compensation for providers of materials such as illustrations and audio. [Means for solving the problem]
[0005] The present invention is a system that has a means for receiving text and storyboard data entered by a user and analyzing them. The system also has a means for searching for and acquiring the necessary materials based on the analysis results, and a means for generating payment information for the use of the materials. Once the user confirms that payment has been completed, a means for automatically generating a video using the materials and the analysis results is activated, and the generated video is provided to the user. This series of processes significantly reduces the effort and cost involved in video production, enabling high-quality videos to be used quickly and efficiently.
[0006] A "user" is an entity that uses the video generation system to input text and storyboards and receives the generated video.
[0007] "Text" is text data entered by the user to explain the content of the video.
[0008] A "storyboard" is image data uploaded by a user to visually show the structure and scenes of a video.
[0009] "Means for receiving data" refers to a method or device by which the system accepts text and storyboard data sent by the user.
[0010] The "analyzing means" refers to a method or device that analyzes the received text and storyboard to identify the necessary materials and audio.
[0011] "Materials" refers to data such as illustrations and audio required to generate a video.
[0012] "Search and acquisition means" refers to a method or device for searching and acquiring the necessary materials from a database based on the analysis results.
[0013] The "means for generating payment information" refers to a method or device that calculates the fee to be paid by the user for the material used and generates billing information.
[0014] A "payment verification means" is a method or device for verifying that a user has completed a payment.
[0015] "Means for automatically generating moving images" refers to a method or device for automatically generating moving images based on acquired materials and analysis results.
[0016] "Means for providing videos" refers to a method or device for delivering the generated videos to users. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention relates to a system that automatically generates high-quality animation based on text and storyboards input by a user. This system has the following main functions:
[0039] 1. Data reception function
[0040] The user uses a terminal to access a dedicated web page or application, input text and storyboards, and this data is sent to the server via an HTTP request.
[0041] 2. Data analysis function
[0042] The server analyzes the received text and storyboard data. Specifically, it uses natural language processing technology to analyze the text and image analysis technology to analyze the storyboard. This allows it to identify the materials needed to generate the video.
[0043] 3. Material search and acquisition function
[0044] The server searches and retrieves the necessary materials, such as illustrations and audio, from a materials database based on the analysis results. This process involves searching the materials database using SQL queries.
[0045] 4. Payment information generation function
[0046] The server generates payment information for the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates total billing information.
[0047] 5. Payment confirmation function
[0048] The generated billing information is presented to the user for payment confirmation, and the payment is processed based on the payment information provided by the user. The server then confirms that the payment has been completed.
[0049] 6. Video generation function
[0050] Once payment is complete, the server uses the analysis results and acquired materials to run the AI video generation service and generate a video. The AI video generation service automatically generates a video according to the provided materials and instructions.
[0051] 7. Video provision function
[0052] Finally, the server provides the generated video to the user by sending a download link for the generated video file to the user's device, which the user can use to download or stream the video.
[0053] Specific examples
[0054] Example 1: Manual video for manufacturing industry
[0055] A user in the manufacturing industry enters the steps for assembling part A in text and uploads a simple sketch as a storyboard. The server receives this and performs text and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The server then searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[0056] Example 2: Educational lesson videos
[0057] Educators enter a text description of the solar system and upload a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio of the solar system. It then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[0058] This system will enable users to easily generate and use high-quality videos, significantly reducing the time and cost required for video production and is expected to be used in a wide range of applications.
[0059] The processing flow will be explained below.
[0060] Step 1:
[0061] The user uses the device to access a dedicated web page or application, enter text, and upload storyboards, thereby providing information about the content of the video.
[0062] Specific behavior:
[0063] The user logs in to a dedicated web page or application.
[0064] The user enters a description of the video in a text box.
[0065] The user opens a file selection dialog and selects an image file to upload a storyboard.
[0066] Once you have completed the entry and upload, click the "Submit" button.
[0067] Step 2:
[0068] The device sends the input text and storyboard data to the server as an HTTP request.
[0069] Specific behavior:
[0070] The device encodes the input text and storyboard as JSON format or multipart form data.
[0071] The encoded data is sent as an HTTP POST request to the server's API endpoint.
[0072] Step 3:
[0073] The server analyzes the received text and storyboard data, specifically using natural language processing technology for the text and image analysis technology for the storyboard.
[0074] Specific behavior:
[0075] The server receives the HTTP request and passes the data to the analysis module.
[0076] A natural language processing engine analyzes the text data and identifies the necessary keywords and context.
[0077] An image analysis engine processes the storyboard and extracts specific elements and scenes.
[0078] Step 4:
[0079] The server searches for and acquires the necessary materials based on the analysis results, including searching for illustrations and audio files from a materials database.
[0080] Specific behavior:
[0081] The server executes an SQL query against the material database and searches for materials using search criteria based on the analysis results.
[0082] The path of the corresponding material file is obtained and temporarily saved.
[0083] Step 5:
[0084] The server generates payment information for the used material, including information about the material provider and a totaling process for the usage fee.
[0085] Specific behavior:
[0086] The server obtains information about the material provider and the usage fee from the material database.
[0087] Billing information is generated based on the acquired information and presented to the user.
[0088] Step 6:
[0089] The user confirms the payment information provided by the server and makes the payment, after which the material becomes officially available to the user.
[0090] Specific behavior:
[0091] The server generates billing information and sends it to the user in HTML format.
[0092] The user visits the payment page and enters the required payment information.
[0093] The terminal sends the payment information to the server and waits for the payment to be completed.
[0094] Step 7:
[0095] The server verifies that payment has been made and prepares the video for generation, thereby confirming permission to legally use the material.
[0096] Specific behavior:
[0097] Your server uses the payment gateway API to confirm the payment is successful.
[0098] Once payment is confirmed, the permission to use the materials will be updated in our internal system.
[0099] Step 8:
[0100] The server uses the analysis results and the acquired materials to send the data to an AI video generation service, which then automatically generates the video.
[0101] Specific behavior:
[0102] The server sends the necessary material files and analysis results to the API of the AI video generation service.
[0103] The AI video generation service generates a video based on the provided data and returns the generated video file to the server.
[0104] Step 9:
[0105] The server sends a download link for the generated video to the user's device, allowing the user to access the video.
[0106] Specific behavior:
[0107] The server creates a URL for the generated video file and generates a link in an HTML email or on the dashboard.
[0108] When users click on the link, they can download or stream the video on their device.
[0109] Example 1
[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0111] Demand for high-quality video content is increasing in many fields today. However, video production requires specialized skills and is time-consuming and costly. Furthermore, there are not enough methods available for users to easily create high-quality videos based on their own ideas and information. Therefore, there is a need for a system that can efficiently create and provide high-quality videos without specialized knowledge.
[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0113] In this invention, the server includes means for receiving information and visual material data input by a user, means for analyzing the information and visual material, and means for searching for and acquiring necessary data based on the analysis results, thereby enabling users without specialized knowledge to easily turn their ideas into high-quality videos.
[0114] "User" refers to an individual or organization that utilizes this system to provide information and visual materials and request the creation of a video.
[0115] "Visual materials" refers to materials that convey information visually, such as storyboards and image files that users upload to the system.
[0116] "Data receiving means" refers to a mechanism by which the server receives information and visual materials entered by the user via the network.
[0117] "Analysis means" refers to the process of breaking down and interpreting received information and visual materials, and extracting and identifying the necessary data.
[0118] "Search and acquisition means" refers to the function of searching the database for the necessary data based on the analysis results and acquiring the appropriate materials.
[0119] The "fee information generating means" is a mechanism for calculating the fee for the user based on the data and materials used and creating billing information.
[0120] "Payment confirmation means" refers to a process for confirming that the user has completed payment.
[0121] "Video generation means" refers to a function that automatically generates high-quality video using AI technology, etc., based on the analysis results and acquired materials.
[0122] The "video providing means" refers to a mechanism for providing the generated video in a form that allows users to download or stream the video.
[0123] The present invention relates to a system for automatically generating and providing high-quality video based on user-provided information and visual materials. The system includes the following main components:
[0124] Data reception
[0125] The user uses a device to access a dedicated web page or application. The user inputs and uploads text information and visual materials such as storyboards for the video they want to create. The device then sends this data to the server via an HTTP POST request. The server then receives the data provided by the user.
[0126] Data analysis
[0127] The server analyzes the received text and visual materials. For text data, it uses a natural language processing engine (e.g., SpaCy, BERT) to analyze the text and extract keywords and important location information. For image data, it uses image analysis technology (e.g., OpenCV, TensorFlow) to analyze the storyboard and identify objects. This identifies the materials needed to generate the video.
[0128] Material Search and Acquisition
[0129] Based on the analysis results, the server searches for the necessary data (illustrations, audio files, etc.) from the material database and retrieves them using SQL queries. These materials are used as components necessary for video generation.
[0130] Payment information generation
[0131] The server generates fee information for the use of the material. Specifically, it retrieves information about the material provider and the usage fee from the database, adds them up, and generates an invoice to be presented to the user.
[0132] Payment confirmation
[0133] Provide the user with a link to the payment page and confirm that they wish to complete the payment. The server receives notification that the payment has been completed and proceeds to the next processing step.
[0134] Image Generation
[0135] The server generates videos using AI video generation services (e.g., OpenAI's DALL-E, GPT-4) based on the analysis results and acquired materials. The video generation process is fully automated, and high-quality videos are generated according to the instructions provided by the user.
[0136] Video provided by
[0137] Finally, the server provides the user with a download link for the generated video, which the user can use to download or stream the video, allowing the user to obtain high-quality video in a convenient way.
[0138] Specific examples
[0139] Example 1: Manual video for manufacturing industry
[0140] The user inputs the assembly instructions and associated storyboard for a specific product from their device. The server analyzes the instructions, identifies and acquires the necessary illustrations and audio materials, and generates payment information. Once the user completes payment, the server generates a video using an AI video generation service and provides the user with a download link for the final video.
[0141] prompt:
[0142] "Assembly steps for part A:
[0143] 1. Take out part A.
[0144] 2. Connect to part B.
[0145] 3. Tighten the screws.
[0146] "
[0147] Image: [Illustration: Part A, Part B, Screw diagram]
[0148] Example 2: Educational lesson videos
[0149] The user, an educator, enters a text description of the solar system and uploads a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio of the solar system. It then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user.
[0150] prompt:
[0151] "Solar System Description:
[0152] 1. The sun is at the center and the other planets revolve around it.
[0153] 2. Mercury is the planet closest to the sun.
[0154] 3. Venus is the brightest star known.
[0155] "
[0156] Image: [Diagram showing the solar system, with diagrams of each planet]
[0157] This allows users to easily create high-quality videos without specialized knowledge or skills, and provides a system that can be used for a variety of purposes.
[0158] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0159] Step 1: Receiving data
[0160] Users use their devices to access a dedicated web page or application and input and upload written information about the video they want to create and visual materials such as storyboards.
[0161] Input: Text information and storyboard data entered by the user
[0162] The terminal sends this data to the server as an HTTP POST request.
[0163] Output: Text information and storyboard data received by the server
[0164] Step 2: Data analysis
[0165] The server analyzes the received text information using a natural language processing engine (e.g., SpaCy, BERT) to extract keywords and important location information.
[0166] Input: Received text information
[0167] Output: Extracted keywords and location information
[0168] The server analyzes the storyboard using image analysis technology (e.g., OpenCV, TensorFlow) and identifies the objects.
[0169] Input: Received storyboard data
[0170] Output: Identified objects
[0171] Step 3: Search and acquire materials
[0172] Based on the analysis results, the server searches for the necessary data (illustrations, audio files, etc.) from the material database.
[0173] Input: Analysis results (extracted keywords and identified objects)
[0174] Output: Searched material data
[0175] The server uses an SQL query to search the materials database and retrieve the appropriate materials.
[0176] Input: SQL query
[0177] Output: Acquired material data
[0178] Step 4: Generate payment information
[0179] The server generates fee information based on the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates a combined invoice.
[0180] Input: Acquired material data
[0181] Output: Generated invoice
[0182] Step 5: Payment confirmation
[0183] The server presents the generated invoice to the user and provides a link to a payment page.
[0184] Input: Generated Invoice
[0185] Output: Payment page link
[0186] The user enters payment information and completes the payment.
[0187] Input: User's payment information
[0188] Once the server receives notification that payment has been completed, it proceeds to the next processing step.
[0189] Output: Confirmation of successful payment
[0190] Step 6: Image generation
[0191] The server generates video using an AI video generation service (e.g., OpenAI's DALL-E, GPT-4) based on the analysis results and acquired materials.
[0192] Input: Analysis results and acquired materials
[0193] Output: Generated video file
[0194] The server stores the generated video files.
[0195] Input: Generated video file
[0196] Output: Saved video file
[0197] Step 7: Provide footage
[0198] The server provides the user with a download link for the generated video.
[0199] Input: Saved video file
[0200] Output: Video download link
[0201] Users can use this link to download or stream the video.
[0202] Input: Video download link
[0203] Output: Downloaded or streamed video
[0204] Through these processing steps, users can easily create high-quality videos for a variety of uses, even without specialized knowledge or skills.
[0205] (Application example 1)
[0206] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0207] Conventional video production systems require a lot of time and effort for the entire process of collecting, editing, and creating the materials needed to create a video, making it difficult for average users to easily create high-quality videos. Furthermore, the payment process for using the materials is complicated, placing a burden on users. The present invention aims to solve these problems by providing a system that allows users to easily create high-quality videos and complete payments smoothly.
[0208] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0209] In this invention, the server includes means for receiving text and storyboard data entered by a user, means for analyzing the text and storyboard, means for searching for and acquiring necessary materials based on the analysis results, means for generating payment information for the use of the materials, means for confirming that the user has completed payment, means for automatically generating a video using the materials and the analysis results, means for providing the generated video to the user, means for providing a user interface integrated into a smartphone application, and means for using AI to generate a video based on materials acquired from a materials database and confirm payment. This allows users to generate high-quality videos in a short amount of time and complete payment procedures easily and quickly.
[0210] "User-input text" refers to data based on text provided by a user through the system's input interface.
[0211] A "storyboard" is drawing data based on a visual guide entered by the user.
[0212] A "receiving means" is a part of the system that has the function of capturing data sent by a user.
[0213] The "analyzing means" is a part of the system that has the function of analyzing the received data and extracting the necessary information.
[0214] "Necessary materials" refers to data such as images and audio required to generate video.
[0215] A "search and retrieval means" is a part of a system that has the ability to search for information in a database and retrieve requested material.
[0216] A "means for generating payment information" is a part of the system that has the functionality to calculate fees for use of material and generate that information.
[0217] A "means for confirming payment completion" is a part of the system that has the function of confirming that a user's payment has been successfully made.
[0218] The "means for automatically generating a video" is a part of a system that has the function of automatically creating a video using the acquired materials and analysis results.
[0219] A "means for providing videos to users" is a part of the system that has the function of making the generated videos available to users.
[0220] A "smartphone application" is a program that runs on a smartphone and provides an interface that can be operated directly by the user.
[0221] A "user interface" is a feature that serves as an entry point for a user to interact with a system.
[0222] A "material database" is a data storage that stores materials necessary for video generation.
[0223] "AI-based means" refers to a part of a system that has the function of generating videos using artificial intelligence technology.
[0224] The embodiment of the present invention is a system that automatically generates high-quality videos based on text and storyboards entered by the user. This system is composed of a server and a user's terminal (smartphone application).
[0225] System Program
[0226] The system includes the following key features:
[0227] 1. Data reception function: The server receives text and storyboards sent from the user's device. The user inputs and sends this data via a smartphone application.
[0228] 2. Data analysis function: The server analyzes the received text using natural language processing technology (e.g., spaCy or NLTK) and analyzes the storyboard using image analysis technology (e.g., OpenCV or TensorFlow).
[0229] 3. Material search and retrieval function: Based on the analysis results, the server searches and retrieves related images, audio, and other materials from the material database. Here, the database search is performed using SQL queries.
[0230] 4. Payment information generation function: The server calculates the usage fee for the acquired material and presents payment information to the user. At this time, information about the material provider is also obtained.
[0231] 5. Payment confirmation function: The server processes the payment based on the payment information provided by the user and confirms that the payment has been completed. Possible payment services used include Stripe and PayPal.
[0232] 6. Video generation function: The server automatically generates videos using AI based on the acquired materials and analysis results. AI video generation uses services such as DeepArt and Runway ML.
[0233] 7. Video provision function: The generated video is provided to the user from the server. The user can stream or download the video via the URL of the generated video.
[0234] Natural language description of the process
[0235] Data reception and analysis: The text and storyboards entered by the user using the smartphone application are sent to the server via HTTP requests. The server stores the received data, analyzes the text using natural language processing (NLP) technology, and analyzes the storyboards using image analysis technology.
[0236] Material search and retrieval: Based on the analysis results, the server performs a database search to retrieve the required image and audio materials. The material database stores materials by multiple categories, so the required materials can be quickly retrieved using the appropriate SQL query.
[0237] Payment processing: Calculate the usage fee for the acquired materials and present the payment information to the user. The server confirms that the user has completed the payment and proceeds to the next step.
[0238] Video generation and provision: Using AI technology, a video is generated by combining the necessary materials and analysis results. The generated video is provided to the user as a URL link. The user can view or download the generated video via the link.
[0239] Examples of concrete examples and prompts
[0240] Examples:
[0241] In the education field, users can enter a text description of each planet in the solar system and upload a corresponding hand-drawn sketch of the planet as a storyboard. This data is then analyzed to retrieve images of the corresponding planet from, for example, a NASA image database, and audio material from LibriVox. The resulting video can then be used in the user's educational activities.
[0242] Example prompt sentence:
[0243] "Generate a video about each planet in the solar system. Create a high-quality video based on the text and storyboard below.
[0244] Text: The solar system has the sun at its center and the planets that orbit it are Mercury, Venus, Earth, Mars, Jupiter, Saturn, Uranus, and Neptune.
[0245] Storyboard: Hand-drawn planet sketch (image data)
[0246] The above is a specific embodiment of the present invention, which allows users to easily generate high-quality videos and use them immediately.
[0247] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0248] Step 1:
[0249] The user inputs text and storyboards using a smartphone application and sends them.
[0250] Input: Text data, storyboard data
[0251] Output: HTTP request data
[0252] Specific operation: The user enters text into the application's input form and uploads image files that will serve as storyboards. This data is sent from the device to the server as an HTTP request.
[0253] Step 2:
[0254] The server receives the HTTP request and stores the data.
[0255] Input: HTTP request data
[0256] Output: Saved text data, saved storyboard data
[0257] Specific operation: The server stores the received data in temporary storage and prepares it for the next analysis step.
[0258] Step 3:
[0259] The server analyzes the received data.
[0260] Input: Saved text data, saved storyboard data
[0261] Output: Analyzed text data, analyzed storyboard data
[0262] Specific operation: The server analyzes the text using natural language processing technology (e.g., spaCy or NLTK) to extract keywords and structural information, and also analyzes the storyboard using image analysis technology (e.g., OpenCV or TensorFlow) to extract the objects and layout information contained therein.
[0263] Step 4:
[0264] The server searches the material database based on the analysis results and retrieves the necessary images and audio materials.
[0265] Input: Analyzed text data, analyzed storyboard data
[0266] Output: Acquired image material data, acquired audio material data
[0267] Specific operation: The server uses an SQL query to search the material database based on the analysis results and retrieve relevant images and audio materials.
[0268] Step 5:
[0269] The server calculates the usage fee for the acquired material and generates payment information.
[0270] Input: Acquired image material data, acquired audio material data
[0271] Output: Payment information data
[0272] Specific operation: The server retrieves the information of the material provider and the usage fee data, combines them, and generates payment information, which is then presented to the user.
[0273] Step 6:
[0274] The server processes the payment based on the user's payment information and confirms that the payment has been completed.
[0275] Input: Payment information data, user's payment information
[0276] Output: Payment confirmation data
[0277] Specific operation: The server uses a payment service such as Stripe or PayPal to check whether the user's payment has been successfully completed. If the confirmation is successful, it proceeds to the next step.
[0278] Step 7:
[0279] Using the materials acquired by the server and the analysis results, videos are automatically generated using AI.
[0280] Input: Acquired image material data, acquired audio material data, analyzed text data, analyzed storyboard data
[0281] Output: Generated video data
[0282] Specific operation: The server uses an AI video generation service (e.g., DeepArt or Runway ML) to generate a video that combines the acquired materials and analysis results.
[0283] Step 8:
[0284] The server provides the generated video to the user.
[0285] Input: Generated video data
[0286] Output: Video data provided to users (URL link, etc.)
[0287] Specific operation: The server saves the generated video in cloud storage and provides the URL to the user, who can use this URL to stream or download the video.
[0288] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0289] This invention relates to a system that automatically generates high-quality videos based on text and storyboards entered by the user, as well as a system that recognizes the user's emotions and reflects them in the videos. This system has the following main functions:
[0290] 1. Data reception function
[0291] The user uses a terminal to access a dedicated web page or application, input text and storyboards, and this data is sent to the server via an HTTP request.
[0292] 2. Data analysis function
[0293] The server analyzes the received text and storyboard data using natural language processing and image analysis technologies.
[0294] 3. Emotion engine function
[0295] The server is equipped with an emotion engine for recognizing emotions contained in text and storyboard data. The emotion engine performs emotion analysis when analyzing text, and also performs emotion analysis on storyboards.
[0296] 4. Material search and acquisition function
[0297] The server searches and retrieves the necessary materials from a materials database based on the analysis and emotion analysis results, including appropriate illustrations, audio, and other materials.
[0298] 5. Payment information generation function
[0299] The server generates payment information for the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates total billing information.
[0300] 6. Payment confirmation function
[0301] The generated billing information is presented to the user for payment confirmation, and the payment is processed based on the payment information provided by the user. The server then confirms that the payment has been completed.
[0302] 7. Video generation function
[0303] Once payment is complete, the server uses the analysis results and acquired materials to run the AI video generation service and generate a video, including a function to adjust the expressions in the video based on the emotion engine.
[0304] 8. Video provision function
[0305] Finally, the server provides the generated video to the user by sending a download link for the generated video file to the user's device, which the user can use to download or stream the video.
[0306] Specific examples
[0307] Example 1: Manual video for manufacturing industry
[0308] A user working in the manufacturing industry enters the steps for assembling part A in text and uploads a simple sketch as a storyboard. The server receives this and performs natural language processing and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The emotion engine then analyzes the user's intentions and emotions from the input text and generates a narration that reflects kindness or strictness. The server searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[0309] Example 2: Educational lesson videos
[0310] Educators enter a text description of the solar system and upload a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio for the solar system. The emotion engine then performs an emotional analysis of the entered text and selects passionate narration or a calm tone. The system then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[0311] This system enables users to easily generate and use high-quality, emotionally-reflective videos, significantly reducing the time and cost required for video production and is expected to be used in a wide range of applications.
[0312] The processing flow will be explained below.
[0313] Step 1:
[0314] The user uses the device to access a dedicated web page or application, enter text, and upload storyboards, thereby providing information about the content of the video.
[0315] Specific behavior:
[0316] The user logs in to a dedicated web page or application.
[0317] The user enters a description of the video in a text box.
[0318] The user opens a file selection dialog and selects an image file to upload a storyboard.
[0319] Once you have completed the entry and upload, click the "Submit" button.
[0320] Step 2:
[0321] The device sends the input text and storyboard data to the server as an HTTP request.
[0322] Specific behavior:
[0323] The device encodes the input text and storyboard as JSON format or multipart form data.
[0324] The encoded data is sent as an HTTP POST request to the server's API endpoint.
[0325] Step 3:
[0326] The server analyzes the received text and storyboard data, specifically using natural language processing technology for the text and image analysis technology for the storyboard.
[0327] Specific behavior:
[0328] The server receives the HTTP request and passes the data to the analysis module.
[0329] A natural language processing engine analyzes the text data and identifies the necessary keywords and context.
[0330] An image analysis engine processes the storyboard and extracts specific elements and scenes.
[0331] Step 4:
[0332] The server uses an emotion engine to recognize emotions contained in the received text and storyboard data.
[0333] Specific behavior:
[0334] The emotion engine performs sentiment analysis on the text to identify the user's intentions and emotions.
[0335] The emotion engine performs an emotion analysis of the storyboard and identifies the emotional nuances of the depicted scene.
[0336] Step 5:
[0337] The server searches for and acquires the necessary materials based on the analysis and emotion analysis results, including searching for illustrations and audio files from a materials database.
[0338] Specific behavior:
[0339] The server executes an SQL query against the material database and searches for materials using search criteria based on the analysis results and emotion analysis results.
[0340] The path of the corresponding material file is obtained and temporarily saved.
[0341] Step 6:
[0342] The server generates payment information for the used material, including information about the material provider and a totaling process for the usage fee.
[0343] Specific behavior:
[0344] The server obtains information about the material provider and the usage fee from the material database.
[0345] Billing information is generated based on the acquired information and presented to the user.
[0346] Step 7:
[0347] The user confirms the payment information provided by the server and makes the payment, after which the material becomes officially available to the user.
[0348] Specific behavior:
[0349] The server generates billing information and sends it to the user in HTML format.
[0350] The user visits the payment page and enters the required payment information.
[0351] The terminal sends the payment information to the server and waits for the payment to be completed.
[0352] Step 8:
[0353] The server verifies that payment has been made and prepares the video for generation, thereby confirming permission to legally use the material.
[0354] Specific behavior:
[0355] Your server uses the payment gateway API to confirm the payment is successful.
[0356] Once payment is confirmed, the permission to use the materials will be updated in our internal system.
[0357] Step 9:
[0358] The server uses the analysis results and the acquired materials to send the data to an AI video generation service, which then automatically generates the video.
[0359] Specific behavior:
[0360] The server sends the necessary material files and analysis results to the API of the AI video generation service.
[0361] The AI video generation service generates a video based on the provided data and returns the generated video file to the server.
[0362] Based on the emotion engine, the tone of the video's narration and visual expression are adjusted.
[0363] Step 10:
[0364] The server sends a download link for the generated video to the user's device, allowing the user to access the video.
[0365] Specific behavior:
[0366] The server creates a URL for the generated video file and generates a link in an HTML email or on the dashboard.
[0367] When users click on the link, they can download or stream the video on their device.
[0368] Example 2
[0369] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0370] Conventional video generation systems require a lot of manual work when generating videos based on user-entered text and storyboards. Furthermore, they lacked a means to reflect the user's emotions in the video, making it difficult to generate personalized, high-quality videos. This increased the time and cost required for video production, making them unusable for many users.
[0371] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving text and storyboard data input by the user, means for analyzing the text and storyboard, means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results, means for generating payment information for use of the materials, means for confirming that the user has completed payment, means for automatically generating a video using the materials and the analysis results and emotion analysis results, and means for providing the generated video to the user. This enables the user to automatically generate high-quality videos that reflect emotions with little effort and quickly use them.
[0372] "User" refers to the person who inputs text and storyboards to use the system.
[0373] "Data receiving means" refers to the function of receiving text and storyboard data entered by the user.
[0374] "Data analysis means" refers to a function for analyzing received text and storyboards.
[0375] "Emotion analysis means" refers to a function that identifies the user's emotions contained in text or storyboards based on the analysis results.
[0376] "Material search means" refers to a function for searching for and acquiring necessary materials based on the analysis results and emotion analysis results.
[0377] "Material acquisition means" refers to a function for acquiring materials identified by the search means from a database or external service.
[0378] The "payment information generating means" refers to a function for generating payment information for the materials used.
[0379] "Payment Verification Method" refers to the function that verifies that a user has completed a payment.
[0380] "Video generation means" refers to a function that automatically generates videos based on acquired materials and analysis results.
[0381] "Video providing means" refers to the function of providing the generated video to the user.
[0382] The present invention relates to a system that automatically generates high-quality videos based on text and storyboard data entered by a user, and further enables personalized video creation by recognizing the user's emotions and reflecting them in the videos. This system includes the following components.
[0383] First, the user accesses a dedicated web page or application using their device and enters text and storyboards for the video they want to create. This input data is then sent from the device to the server via an HTTP request.
[0384] The server temporarily stores the received data and then begins processing to analyze it. Natural language processing (NLP) technology is used to analyze the text data, specifically using Python's NLTK and Spacy libraries. Image analysis technology is used to analyze the storyboards, specifically using the OpenCV library for contour detection and shape recognition.
[0385] Furthermore, we use sentiment analysis tools to identify the emotions contained in the analyzed text and storyboards. For sentiment analysis of text, we use natural language processing services such as Google Cloud NLP, and for sentiment analysis of storyboards, we use machine learning frameworks such as TensorFlow.
[0386] Based on the analysis results and sentiment analysis results, the necessary materials (e.g., illustrations and audio) are searched and acquired from a stock material database. The material database is connected to the API of an online stock material site, allowing necessary materials to be searched and acquired programmatically.
[0387] Next, the server generates payment information for the materials used. This information is generated by retrieving information about the material provider and the usage fee from a database and adding them up. The database used here is MySQL, a common relational database.
[0388] Once the user's payment is confirmed, the server generates a video using AI video generation techniques, adjusting the expressions in the video based on the results of sentiment analysis. The video is generated using OpenAI's CLIP model and DeepAI's API.
[0389] Finally, the generated video is served to the user: the server sends a download link for the generated video to the user's device, which the user can use to download or stream the video.
[0390] Specific examples
[0391] Example 1: Manual video for manufacturing industry
[0392] The user enters the "steps to assemble part A" in text on their device and uploads a simple sketch as a storyboard. The server receives this and performs natural language processing and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The emotion engine then analyzes the user's intentions and emotions from the entered text and generates a narration that reflects kindness or strictness. The server searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[0393] Example 2: Educational lesson videos
[0394] A user who is an educator enters a text description of the solar system and uploads a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio for the solar system. The emotion engine also performs emotional analysis of the entered text and selects passionate narration or a calm tone. The server searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[0395] Prompt Sentence Examples
[0396] Example 1: Manual video for manufacturing industry
[0397] "Please upload a written description of the steps to assemble part A and a simple sketch as a storyboard. After analysis, we will identify the illustrations of the parts and the audio material of the assembly steps and generate a narration that reflects kindness or strictness."
[0398] Example 2: Educational lesson videos
[0399] "Enter a description of the solar system in text and upload a simple storyboard for each planet. After analysis, we'll generate a video with illustrations of the solar system and a choice of passionate or calming narration."
[0400] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0401] Step 1: Data entry
[0402] The user uses a terminal to access a dedicated web page or application and inputs text and storyboards related to the video they want to generate. Specifically, the user inputs the text "Assembly procedure for part A" and uploads the storyboard for part A. The input format is a text file and an image file. This results in the text and storyboards as input data.
[0403] Step 2: Send data
[0404] Data entered by the user is sent from the terminal to the server via an HTTP request. For example, a POST request is used to send text data and image data to the server's endpoint. The data is received by the server and temporarily stored in a database or system memory.
[0405] Step 3: Text analysis
[0406] The server analyzes the received text data using natural language processing (NLP) techniques, specifically Python's NLTK and Spacy libraries. For example, it tokenizes the text, tags parts of speech, and performs syntactic analysis. This process extracts syntactic and semantic data from the input text.
[0407] Step 4: Storyboard Analysis
[0408] The server analyzes the storyboard it receives using image analysis technology. The OpenCV library is used for image analysis, specifically contour detection and shape recognition. For example, it can detect a specific shape (e.g., the shape of part A) from the uploaded storyboard and extract the corresponding data.
[0409] Step 5: Sentiment Analysis
[0410] The server uses an emotion engine to identify emotions contained in the text and storyboards. It uses machine learning frameworks such as Google Cloud NLP for emotion analysis of text and TensorFlow for emotion analysis of storyboards. Specifically, it uses an emotion model to identify emotions such as "joy," "sadness," and "anger."
[0411] Step 6: Material Search
[0412] Based on the analysis and emotion analysis results, the server searches for the necessary materials from a material database. The material database is connected to an online stock material site using an API. For example, it searches for "illustration of part A" or "gentle-toned voice narration" to identify appropriate materials.
[0413] Step 7: Obtaining Materials
[0414] The server retrieves the searched material. This involves downloading data from the stock material site via API. The retrieved material is then stored in a directory or database managed by the server.
[0415] Step 8: Generate payment information
[0416] The server generates payment information for the materials used. The payment information is generated by retrieving information about the material provider and the usage fee from a database and adding them up. Specifically, a specific SQL query is executed using a MySQL database to calculate the usage fee.
[0417] Step 9: Confirm payment
[0418] The server presents the generated billing information to the user. The user enters the payment information (e.g., credit card information). The server processes the payment using an online payment system (e.g., PayPal or Stripe). If the payment is successful, the server confirms it.
[0419] Step 10: Video Generation
[0420] After the user's payment is confirmed, the server uses the analysis results and the acquired materials to execute the AI video generation service. The emotion analysis results are taken into account and the expressions in the video are adjusted. The video is generated using OpenAI's CLIP model and DeepAI's API. The generated video file is saved in the appropriate format.
[0421] Step 11: Submit your video
[0422] The server provides the generated video to the user by generating a download link for the video and sending it to the user's device. The user can then use the link to download or stream the video.
[0423] (Application example 2)
[0424] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0425] Conventional video generation systems simply generate videos based on user-entered text and storyboards, making it difficult to generate personalized videos that reflect the user's emotions. Furthermore, there is a lack of systems that efficiently execute the entire process of delivering generated videos to target users. For this reason, the advertising industry needs a way to quickly generate and deliver effective video ads that match the interests and emotions of target users.
[0426] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0427] In this invention, the server includes means for receiving text and storyboard data input by the user, means for analyzing the text and storyboard, and means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results. This makes it possible to generate high-quality videos that reflect emotions based on the user's input data and distribute them to target users.
[0428] "Means for receiving text and storyboard data entered by the user" refers to a function that allows a user to use a dedicated application or web page via a terminal to enter text and storyboard data and send it to the server.
[0429] The "means for analyzing text and storyboards" refers to a means for analyzing received text and storyboard data using natural language processing and image analysis techniques to understand the content.
[0430] "Means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results" refers to means for searching for and acquiring necessary video and audio materials from a materials database based on the analyzed data and the results of emotion analysis.
[0431] The "means for generating payment information for use of material" is a function for calculating the fee for the used material and generating corresponding payment information.
[0432] The "means for confirming that the user has completed the payment" refers to a means for confirming that the user has made the payment based on the generated payment information.
[0433] "Means for automatically generating videos using materials, analysis results, and emotion analysis results" refers to means for automatically generating videos using AI video generation technology based on the acquired materials and analysis results.
[0434] "Means for providing the generated video to the user" means means for providing the generated video to the user and enabling the user to download or stream the video.
[0435] "Means for distributing the generated video to target users" refers to means for distributing the generated video to target users designated by the advertiser, thereby increasing the effectiveness of the advertisement.
[0436] The embodiments for carrying out the present invention are described in detail below.
[0437] The system aims to enable advertisers to generate personalized video ads using smartphone applications and efficiently deliver them to target users.
[0438] 1. User Input and Data Receipt
[0439] Advertisers use a dedicated smartphone application to input text and storyboards for their advertisements, which are then sent to the server via HTTP requests.
[0440] 2. Data Analysis
[0441] The server analyzes the received text and storyboard data using a natural language processing library (e.g., SpaCy) and an image analysis library (e.g., OpenCV). The analysis includes syntactic analysis of the text data and feature extraction of the image data.
[0442] 3. Emotion recognition
[0443] The emotion engine analyzes emotions based on the analyzed data. A common emotion analysis API (e.g., IBM Watson Natural Language Understanding, Google Cloud Natural Language API) is used as the emotion engine.
[0444] 4. Material Search and Acquisition
[0445] Based on the results of sentiment analysis and data analysis, the server searches for and obtains the necessary video and audio materials from a material database (e.g., Shutterstock API).
[0446] 5. Payment Information Generation and Verification
[0447] The server generates payment information based on the information of the materials used and presents it to the advertiser. Payment processing is performed using an electronic payment service (e.g., Stripe API). Once the advertiser completes payment, the information is recorded on the server.
[0448] 6. Video Generation
[0449] Once payment is confirmed, the server uses the acquired materials and analysis results to automatically generate a video using an AI video generation service (e.g., Synthesia). The generated video reflects narration and direction based on the results of the emotion analysis.
[0450] 7. Video Provision and Distribution
[0451] The server provides the generated video advertisement to the advertiser, who then sends a link to the advertisement via a smartphone application, allowing the advertiser to distribute the video to target users.
[0452] Specific examples
[0453] Example 1: Advertising video for new product “Express Coffee”
[0454] Advertisers enter text and sketch images into the application that explain the features of their new product, "Express Coffee."
[0455] Example sentence: "Express Coffee provides fast, delicious coffee for busy mornings."
[0456] Storyboard: "Image of an Express Coffee package and coffee being poured into a cup"
[0457] Emotion: We want to convey a sense of comfort and trust to the user.
[0458] The server analyzes the input data and searches for and retrieves appropriate content from a content database based on the results of sentiment analysis. Once payment is completed, an AI video generation service is used to generate a video ad that reflects the specified sentiment. The generated video ad link is provided to the advertiser through the application, and the advertiser uses this link to deliver the video to target users.
[0459] This system enables advertisers to quickly generate high-quality, emotionally relevant video ads and deliver them effectively to target users.
[0460] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0461] Step 1:
[0462] Input and Data Reception
[0463] Users use a smartphone application to input text and storyboards to be used in advertising videos. The input data is collected through a form in the application and sent to the server as an HTTP request. The server temporarily stores the received data.
[0464] Specific operation: The user uploads text data (e.g., advertising copy) and image data (e.g., product sketch) into the application's input form and sends it to a dedicated API endpoint.
[0465] Step 2:
[0466] Data analysis
[0467] The server analyzes the received text and storyboard data. Specifically, it analyzes the text using a natural language processing library (e.g., SpaCy) and the storyboard using an image analysis library (e.g., OpenCV). Text analysis is used to understand the meaning and structure of the text, and image analysis is used to extract the features of the storyboard.
[0468] Input and Output: The input data are the received text and storyboard, and the output is the analyzed text data and image data features.
[0469] Specific operation: Using the SpaCy library, tokenize text data, assign POS tags, and analyze dependencies. Using the OpenCV library, detect edges and extract feature points from storyboard images.
[0470] Step 3:
[0471] emotion recognition
[0472] The server uses an emotion engine based on the analyzed data to analyze emotions from the text and storyboard. It uses an emotion analysis API (e.g., IBM Watson Natural Language Understanding, Google Cloud Natural Language API) to identify the user's emotions.
[0473] Input and Output: The input is the analyzed text data and image data features, and the output is the sentiment analysis results.
[0474] Specific operation: The analyzed text data and image data features are sent to the emotion analysis API, and emotions are categorized (e.g., joy, sadness, surprise, etc.) and received.
[0475] Step 4:
[0476] Material Search and Acquisition
[0477] Based on the results of emotion analysis and data analysis, the server searches and retrieves appropriate video and audio materials from a material database (e.g., Shutterstock API).
[0478] Input and output: The input is the emotion analysis results and data analysis results, and the output is the acquired video and audio materials.
[0479] Specific operation: Creates a search query for the Shutterstock API based on the results of sentiment analysis and data analysis, and retrieves the necessary materials (images, audio, video).
[0480] Step 5:
[0481] Payment information generation and verification
[0482] The server generates payment information based on the information about the used materials and presents it to the user. It processes the payment using an electronic payment service (e.g., Stripe API) and confirms that the payment has been completed.
[0483] Input and Output: The input is the information of the material acquired and the output is the status of payment confirmation.
[0484] Specific operation: Calculate the total amount based on the price information of the obtained materials, send a payment request to the user through the Stripe API, and check the status of whether the payment has been completed.
[0485] Step 6:
[0486] Video Generation
[0487] Once payment is confirmed, the server uses the acquired materials and analysis results to automatically generate a video using an AI video generation service (e.g., Synthesia), including narration and direction based on the results of the emotion analysis.
[0488] Input and Output: The input is the acquired material and analysis results, and the output is the generated video file.
[0489] Specific operation: The acquired material, analysis results, and a prompt for a video to be generated based on the emotion analysis results are sent to the Synthesia API, and the generated video is retrieved.
[0490] Step 7:
[0491] Video provision and distribution
[0492] The generated video advertisement is provided to the user (advertiser) from the server, and a link is sent to the user via a smartphone application. The user can then distribute the video to target users.
[0493] Input and Output: The input is the generated video file and the output is the video link provided to the user and the delivery status to the target user.
[0494] Specific operation: The generated video is stored in the server storage, and a download link is generated and sent to the user. The user receives the link and the video is distributed to the target user.
[0495] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0496] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0497] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0498] [Second embodiment]
[0499] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0500] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0501] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0502] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0503] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0504] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0505] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0506] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0507] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0508] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0509] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0510] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0511] The present invention relates to a system that automatically generates high-quality animation based on text and storyboards input by a user. This system has the following main functions:
[0512] 1. Data reception function
[0513] The user uses a terminal to access a dedicated web page or application, input text and storyboards, and this data is sent to the server via an HTTP request.
[0514] 2. Data analysis function
[0515] The server analyzes the received text and storyboard data. Specifically, it uses natural language processing technology to analyze the text and image analysis technology to analyze the storyboard. This allows it to identify the materials needed to generate the video.
[0516] 3. Material search and acquisition function
[0517] The server searches and retrieves the necessary materials, such as illustrations and audio, from a materials database based on the analysis results. This process involves searching the materials database using SQL queries.
[0518] 4. Payment information generation function
[0519] The server generates payment information for the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates total billing information.
[0520] 5. Payment confirmation function
[0521] The generated billing information is presented to the user for payment confirmation, and the payment is processed based on the payment information provided by the user. The server then confirms that the payment has been completed.
[0522] 6. Video generation function
[0523] Once payment is complete, the server uses the analysis results and acquired materials to run the AI video generation service and generate a video. The AI video generation service automatically generates a video according to the provided materials and instructions.
[0524] 7. Video provision function
[0525] Finally, the server provides the generated video to the user by sending a download link for the generated video file to the user's device, which the user can use to download or stream the video.
[0526] Specific examples
[0527] Example 1: Manual video for manufacturing industry
[0528] A user in the manufacturing industry enters the steps for assembling part A in text and uploads a simple sketch as a storyboard. The server receives this and performs text and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The server then searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[0529] Example 2: Educational lesson videos
[0530] Educators enter a text description of the solar system and upload a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio of the solar system. It then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[0531] This system will enable users to easily generate and use high-quality videos, significantly reducing the time and cost required for video production and is expected to be used in a wide range of applications.
[0532] The processing flow will be explained below.
[0533] Step 1:
[0534] The user uses the device to access a dedicated web page or application, enter text, and upload storyboards, thereby providing information about the content of the video.
[0535] Specific behavior:
[0536] The user logs in to a dedicated web page or application.
[0537] The user enters a description of the video in a text box.
[0538] The user opens a file selection dialog and selects an image file to upload a storyboard.
[0539] Once you have completed the entry and upload, click the "Submit" button.
[0540] Step 2:
[0541] The device sends the input text and storyboard data to the server as an HTTP request.
[0542] Specific behavior:
[0543] The device encodes the input text and storyboard as JSON format or multipart form data.
[0544] The encoded data is sent as an HTTP POST request to the server's API endpoint.
[0545] Step 3:
[0546] The server analyzes the received text and storyboard data, specifically using natural language processing technology for the text and image analysis technology for the storyboard.
[0547] Specific behavior:
[0548] The server receives the HTTP request and passes the data to the analysis module.
[0549] A natural language processing engine analyzes the text data and identifies the necessary keywords and context.
[0550] An image analysis engine processes the storyboard and extracts specific elements and scenes.
[0551] Step 4:
[0552] The server searches for and acquires the necessary materials based on the analysis results, including searching for illustrations and audio files from a materials database.
[0553] Specific behavior:
[0554] The server executes an SQL query against the material database and searches for materials using search criteria based on the analysis results.
[0555] The path of the corresponding material file is obtained and temporarily saved.
[0556] Step 5:
[0557] The server generates payment information for the used material, including information about the material provider and a totaling process for the usage fee.
[0558] Specific behavior:
[0559] The server obtains information about the material provider and the usage fee from the material database.
[0560] Billing information is generated based on the acquired information and presented to the user.
[0561] Step 6:
[0562] The user confirms the payment information provided by the server and makes the payment, after which the material becomes officially available to the user.
[0563] Specific behavior:
[0564] The server generates billing information and sends it to the user in HTML format.
[0565] The user visits the payment page and enters the required payment information.
[0566] The terminal sends the payment information to the server and waits for the payment to be completed.
[0567] Step 7:
[0568] The server verifies that payment has been made and prepares the video for generation, thereby confirming permission to legally use the material.
[0569] Specific behavior:
[0570] Your server uses the payment gateway API to confirm the payment is successful.
[0571] Once payment is confirmed, the permission to use the materials will be updated in our internal system.
[0572] Step 8:
[0573] The server uses the analysis results and the acquired materials to send the data to an AI video generation service, which then automatically generates the video.
[0574] Specific behavior:
[0575] The server sends the necessary material files and analysis results to the API of the AI video generation service.
[0576] The AI video generation service generates a video based on the provided data and returns the generated video file to the server.
[0577] Step 9:
[0578] The server sends a download link for the generated video to the user's device, allowing the user to access the video.
[0579] Specific behavior:
[0580] The server creates a URL for the generated video file and generates a link in an HTML email or on the dashboard.
[0581] When users click on the link, they can download or stream the video on their device.
[0582] Example 1
[0583] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0584] Demand for high-quality video content is increasing in many fields today. However, video production requires specialized skills and is time-consuming and costly. Furthermore, there are not enough methods available for users to easily create high-quality videos based on their own ideas and information. Therefore, there is a need for a system that can efficiently create and provide high-quality videos without specialized knowledge.
[0585] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0586] In this invention, the server includes means for receiving information and visual material data input by a user, means for analyzing the information and visual material, and means for searching for and acquiring necessary data based on the analysis results, thereby enabling users without specialized knowledge to easily turn their ideas into high-quality videos.
[0587] "User" refers to an individual or organization that utilizes this system to provide information and visual materials and request the creation of a video.
[0588] "Visual materials" refers to materials that convey information visually, such as storyboards and image files that users upload to the system.
[0589] "Data receiving means" refers to a mechanism by which the server receives information and visual materials entered by the user via the network.
[0590] "Analysis means" refers to the process of breaking down and interpreting received information and visual materials, and extracting and identifying the necessary data.
[0591] "Search and acquisition means" refers to the function of searching the database for the necessary data based on the analysis results and acquiring the appropriate materials.
[0592] The "fee information generating means" is a mechanism for calculating the fee for the user based on the data and materials used and creating billing information.
[0593] "Payment confirmation means" refers to a process for confirming that the user has completed payment.
[0594] "Video generation means" refers to a function that automatically generates high-quality video using AI technology, etc., based on the analysis results and acquired materials.
[0595] The "video providing means" refers to a mechanism for providing the generated video in a form that allows users to download or stream the video.
[0596] The present invention relates to a system for automatically generating and providing high-quality video based on user-provided information and visual materials. The system includes the following main components:
[0597] Data reception
[0598] The user uses a device to access a dedicated web page or application. The user inputs and uploads text information and visual materials such as storyboards for the video they want to create. The device then sends this data to the server via an HTTP POST request. The server then receives the data provided by the user.
[0599] Data analysis
[0600] The server analyzes the received text and visual materials. For text data, it uses a natural language processing engine (e.g., SpaCy, BERT) to analyze the text and extract keywords and important location information. For image data, it uses image analysis technology (e.g., OpenCV, TensorFlow) to analyze the storyboard and identify objects. This identifies the materials needed to generate the video.
[0601] Material Search and Acquisition
[0602] Based on the analysis results, the server searches for the necessary data (illustrations, audio files, etc.) from the material database and retrieves them using SQL queries. These materials are used as components necessary for video generation.
[0603] Payment information generation
[0604] The server generates fee information for the use of the material. Specifically, it retrieves information about the material provider and the usage fee from the database, adds them up, and generates an invoice to be presented to the user.
[0605] Payment confirmation
[0606] Provide the user with a link to the payment page and confirm that they wish to complete the payment. The server receives notification that the payment has been completed and proceeds to the next processing step.
[0607] Image Generation
[0608] The server generates videos using AI video generation services (e.g., OpenAI's DALL-E, GPT-4) based on the analysis results and acquired materials. The video generation process is fully automated, and high-quality videos are generated according to the instructions provided by the user.
[0609] Video provided by
[0610] Finally, the server provides the user with a download link for the generated video, which the user can use to download or stream the video, allowing the user to obtain high-quality video in a convenient way.
[0611] Specific examples
[0612] Example 1: Manual video for manufacturing industry
[0613] The user inputs the assembly instructions and associated storyboard for a specific product from their device. The server analyzes the instructions, identifies and acquires the necessary illustrations and audio materials, and generates payment information. Once the user completes payment, the server generates a video using an AI video generation service and provides the user with a download link for the final video.
[0614] prompt:
[0615] "Assembly steps for part A:
[0616] 1. Take out part A.
[0617] 2. Connect to part B.
[0618] 3. Tighten the screws.
[0619] "
[0620] Image: [Illustration: Part A, Part B, Screw diagram]
[0621] Example 2: Educational lesson videos
[0622] The user, an educator, enters a text description of the solar system and uploads a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio of the solar system. It then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user.
[0623] prompt:
[0624] "Solar System Description:
[0625] 1. The sun is at the center and the other planets revolve around it.
[0626] 2. Mercury is the planet closest to the sun.
[0627] 3. Venus is the brightest star known.
[0628] "
[0629] Image: [Diagram showing the solar system, with diagrams of each planet]
[0630] This allows users to easily create high-quality videos without specialized knowledge or skills, and provides a system that can be used for a variety of purposes.
[0631] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0632] Step 1: Receiving data
[0633] Users use their devices to access a dedicated web page or application and input and upload written information about the video they want to create and visual materials such as storyboards.
[0634] Input: Text information and storyboard data entered by the user
[0635] The terminal sends this data to the server as an HTTP POST request.
[0636] Output: Text information and storyboard data received by the server
[0637] Step 2: Data analysis
[0638] The server analyzes the received text information using a natural language processing engine (e.g., SpaCy, BERT) to extract keywords and important location information.
[0639] Input: Received text information
[0640] Output: Extracted keywords and location information
[0641] The server analyzes the storyboard using image analysis technology (e.g., OpenCV, TensorFlow) and identifies the objects.
[0642] Input: Received storyboard data
[0643] Output: Identified objects
[0644] Step 3: Search and acquire materials
[0645] Based on the analysis results, the server searches for the necessary data (illustrations, audio files, etc.) from the material database.
[0646] Input: Analysis results (extracted keywords and identified objects)
[0647] Output: Searched material data
[0648] The server uses an SQL query to search the materials database and retrieve the appropriate materials.
[0649] Input: SQL query
[0650] Output: Acquired material data
[0651] Step 4: Generate payment information
[0652] The server generates fee information based on the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates a combined invoice.
[0653] Input: Acquired material data
[0654] Output: Generated invoice
[0655] Step 5: Payment confirmation
[0656] The server presents the generated invoice to the user and provides a link to a payment page.
[0657] Input: Generated Invoice
[0658] Output: Payment page link
[0659] The user enters payment information and completes the payment.
[0660] Input: User's payment information
[0661] Once the server receives notification that payment has been completed, it proceeds to the next processing step.
[0662] Output: Confirmation of successful payment
[0663] Step 6: Image generation
[0664] The server generates video using an AI video generation service (e.g., OpenAI's DALL-E, GPT-4) based on the analysis results and acquired materials.
[0665] Input: Analysis results and acquired materials
[0666] Output: Generated video file
[0667] The server stores the generated video files.
[0668] Input: Generated video file
[0669] Output: Saved video file
[0670] Step 7: Provide footage
[0671] The server provides the user with a download link for the generated video.
[0672] Input: Saved video file
[0673] Output: Video download link
[0674] Users can use this link to download or stream the video.
[0675] Input: Video download link
[0676] Output: Downloaded or streamed video
[0677] Through these processing steps, users can easily create high-quality videos for a variety of uses, even without specialized knowledge or skills.
[0678] (Application example 1)
[0679] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0680] Conventional video production systems require a lot of time and effort for the entire process of collecting, editing, and creating the materials needed to create a video, making it difficult for average users to easily create high-quality videos. Furthermore, the payment process for using the materials is complicated, placing a burden on users. The present invention aims to solve these problems by providing a system that allows users to easily create high-quality videos and complete payments smoothly.
[0681] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0682] In this invention, the server includes means for receiving text and storyboard data entered by a user, means for analyzing the text and storyboard, means for searching for and acquiring necessary materials based on the analysis results, means for generating payment information for the use of the materials, means for confirming that the user has completed payment, means for automatically generating a video using the materials and the analysis results, means for providing the generated video to the user, means for providing a user interface integrated into a smartphone application, and means for using AI to generate a video based on materials acquired from a materials database and confirm payment. This allows users to generate high-quality videos in a short amount of time and complete payment procedures easily and quickly.
[0683] "User-input text" refers to data based on text provided by a user through the system's input interface.
[0684] A "storyboard" is drawing data based on a visual guide entered by the user.
[0685] A "receiving means" is a part of the system that has the function of capturing data sent by a user.
[0686] The "analyzing means" is a part of the system that has the function of analyzing the received data and extracting the necessary information.
[0687] "Necessary materials" refers to data such as images and audio required to generate video.
[0688] A "search and retrieval means" is a part of a system that has the ability to search for information in a database and retrieve requested material.
[0689] A "means for generating payment information" is a part of the system that has the functionality to calculate fees for use of material and generate that information.
[0690] A "means for confirming payment completion" is a part of the system that has the function of confirming that a user's payment has been successfully made.
[0691] The "means for automatically generating a video" is a part of a system that has the function of automatically creating a video using the acquired materials and analysis results.
[0692] A "means for providing videos to users" is a part of the system that has the function of making the generated videos available to users.
[0693] A "smartphone application" is a program that runs on a smartphone and provides an interface that can be operated directly by the user.
[0694] A "user interface" is a feature that serves as an entry point for a user to interact with a system.
[0695] A "material database" is a data storage that stores materials necessary for video generation.
[0696] "AI-based means" refers to a part of a system that has the function of generating videos using artificial intelligence technology.
[0697] The embodiment of the present invention is a system that automatically generates high-quality videos based on text and storyboards entered by the user. This system is composed of a server and a user's terminal (smartphone application).
[0698] System Program
[0699] The system includes the following key features:
[0700] 1. Data reception function: The server receives text and storyboards sent from the user's device. The user inputs and sends this data via a smartphone application.
[0701] 2. Data analysis function: The server analyzes the received text using natural language processing technology (e.g., spaCy or NLTK) and analyzes the storyboard using image analysis technology (e.g., OpenCV or TensorFlow).
[0702] 3. Material search and retrieval function: Based on the analysis results, the server searches and retrieves related images, audio, and other materials from the material database. Here, the database search is performed using SQL queries.
[0703] 4. Payment information generation function: The server calculates the usage fee for the acquired material and presents payment information to the user. At this time, information about the material provider is also obtained.
[0704] 5. Payment confirmation function: The server processes the payment based on the payment information provided by the user and confirms that the payment has been completed. Possible payment services used include Stripe and PayPal.
[0705] 6. Video generation function: The server automatically generates videos using AI based on the acquired materials and analysis results. AI video generation uses services such as DeepArt and Runway ML.
[0706] 7. Video provision function: The generated video is provided to the user from the server. The user can stream or download the video via the URL of the generated video.
[0707] Natural language description of the process
[0708] Data reception and analysis: The text and storyboards entered by the user using the smartphone application are sent to the server via HTTP requests. The server stores the received data, analyzes the text using natural language processing (NLP) technology, and analyzes the storyboards using image analysis technology.
[0709] Material search and retrieval: Based on the analysis results, the server performs a database search to retrieve the required image and audio materials. The material database stores materials by multiple categories, so the required materials can be quickly retrieved using the appropriate SQL query.
[0710] Payment processing: Calculate the usage fee for the acquired materials and present the payment information to the user. The server confirms that the user has completed the payment and proceeds to the next step.
[0711] Video generation and provision: Using AI technology, a video is generated by combining the necessary materials and analysis results. The generated video is provided to the user as a URL link. The user can view or download the generated video via the link.
[0712] Examples of concrete examples and prompts
[0713] Examples:
[0714] In the education field, users can enter a text description of each planet in the solar system and upload a corresponding hand-drawn sketch of the planet as a storyboard. This data is then analyzed to retrieve images of the corresponding planet from, for example, a NASA image database, and audio material from LibriVox. The resulting video can then be used in the user's educational activities.
[0715] Example prompt sentence:
[0716] "Generate a video about each planet in the solar system. Create a high-quality video based on the text and storyboard below.
[0717] Text: The solar system has the sun at its center and the planets that orbit it are Mercury, Venus, Earth, Mars, Jupiter, Saturn, Uranus, and Neptune.
[0718] Storyboard: Hand-drawn planet sketch (image data)
[0719] The above is a specific embodiment of the present invention, which allows users to easily generate high-quality videos and use them immediately.
[0720] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0721] Step 1:
[0722] The user inputs text and storyboards using a smartphone application and sends them.
[0723] Input: Text data, storyboard data
[0724] Output: HTTP request data
[0725] Specific operation: The user enters text into the application's input form and uploads image files that will serve as storyboards. This data is sent from the device to the server as an HTTP request.
[0726] Step 2:
[0727] The server receives the HTTP request and stores the data.
[0728] Input: HTTP request data
[0729] Output: Saved text data, saved storyboard data
[0730] Specific operation: The server stores the received data in temporary storage and prepares it for the next analysis step.
[0731] Step 3:
[0732] The server analyzes the received data.
[0733] Input: Saved text data, saved storyboard data
[0734] Output: Analyzed text data, analyzed storyboard data
[0735] Specific operation: The server analyzes the text using natural language processing technology (e.g., spaCy or NLTK) to extract keywords and structural information, and also analyzes the storyboard using image analysis technology (e.g., OpenCV or TensorFlow) to extract the objects and layout information contained therein.
[0736] Step 4:
[0737] The server searches the material database based on the analysis results and retrieves the necessary images and audio materials.
[0738] Input: Analyzed text data, analyzed storyboard data
[0739] Output: Acquired image material data, acquired audio material data
[0740] Specific operation: The server uses an SQL query to search the material database based on the analysis results and retrieve relevant images and audio materials.
[0741] Step 5:
[0742] The server calculates the usage fee for the acquired material and generates payment information.
[0743] Input: Acquired image material data, acquired audio material data
[0744] Output: Payment information data
[0745] Specific operation: The server retrieves the information of the material provider and the usage fee data, combines them, and generates payment information, which is then presented to the user.
[0746] Step 6:
[0747] The server processes the payment based on the user's payment information and confirms that the payment has been completed.
[0748] Input: Payment information data, user's payment information
[0749] Output: Payment confirmation data
[0750] Specific operation: The server uses a payment service such as Stripe or PayPal to check whether the user's payment has been successfully completed. If the confirmation is successful, it proceeds to the next step.
[0751] Step 7:
[0752] Using the materials acquired by the server and the analysis results, videos are automatically generated using AI.
[0753] Input: Acquired image material data, acquired audio material data, analyzed text data, analyzed storyboard data
[0754] Output: Generated video data
[0755] Specific operation: The server uses an AI video generation service (e.g., DeepArt or Runway ML) to generate a video that combines the acquired materials and analysis results.
[0756] Step 8:
[0757] The server provides the generated video to the user.
[0758] Input: Generated video data
[0759] Output: Video data provided to users (URL link, etc.)
[0760] Specific operation: The server saves the generated video in cloud storage and provides the URL to the user, who can use this URL to stream or download the video.
[0761] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0762] This invention relates to a system that automatically generates high-quality videos based on text and storyboards entered by the user, as well as a system that recognizes the user's emotions and reflects them in the videos. This system has the following main functions:
[0763] 1. Data reception function
[0764] The user uses a terminal to access a dedicated web page or application, input text and storyboards, and this data is sent to the server via an HTTP request.
[0765] 2. Data analysis function
[0766] The server analyzes the received text and storyboard data using natural language processing and image analysis technologies.
[0767] 3. Emotion engine function
[0768] The server is equipped with an emotion engine for recognizing emotions contained in text and storyboard data. The emotion engine performs emotion analysis when analyzing text, and also performs emotion analysis on storyboards.
[0769] 4. Material search and acquisition function
[0770] The server searches and retrieves the necessary materials from a materials database based on the analysis and emotion analysis results, including appropriate illustrations, audio, and other materials.
[0771] 5. Payment information generation function
[0772] The server generates payment information for the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates total billing information.
[0773] 6. Payment confirmation function
[0774] The generated billing information is presented to the user for payment confirmation, and the payment is processed based on the payment information provided by the user. The server then confirms that the payment has been completed.
[0775] 7. Video generation function
[0776] Once payment is complete, the server uses the analysis results and acquired materials to run the AI video generation service and generate a video, including a function to adjust the expressions in the video based on the emotion engine.
[0777] 8. Video provision function
[0778] Finally, the server provides the generated video to the user by sending a download link for the generated video file to the user's device, which the user can use to download or stream the video.
[0779] Specific examples
[0780] Example 1: Manual video for manufacturing industry
[0781] A user working in the manufacturing industry enters the steps for assembling part A in text and uploads a simple sketch as a storyboard. The server receives this and performs natural language processing and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The emotion engine then analyzes the user's intentions and emotions from the input text and generates a narration that reflects kindness or strictness. The server searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[0782] Example 2: Educational lesson videos
[0783] Educators enter a text description of the solar system and upload a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio for the solar system. The emotion engine then performs an emotional analysis of the entered text and selects passionate narration or a calm tone. The system then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[0784] This system enables users to easily generate and use high-quality, emotionally-reflective videos, significantly reducing the time and cost required for video production and is expected to be used in a wide range of applications.
[0785] The processing flow will be explained below.
[0786] Step 1:
[0787] The user uses the device to access a dedicated web page or application, enter text, and upload storyboards, thereby providing information about the content of the video.
[0788] Specific behavior:
[0789] The user logs in to a dedicated web page or application.
[0790] The user enters a description of the video in a text box.
[0791] The user opens a file selection dialog and selects an image file to upload a storyboard.
[0792] Once you have completed the entry and upload, click the "Submit" button.
[0793] Step 2:
[0794] The device sends the input text and storyboard data to the server as an HTTP request.
[0795] Specific behavior:
[0796] The device encodes the input text and storyboard as JSON format or multipart form data.
[0797] The encoded data is sent as an HTTP POST request to the server's API endpoint.
[0798] Step 3:
[0799] The server analyzes the received text and storyboard data, specifically using natural language processing technology for the text and image analysis technology for the storyboard.
[0800] Specific behavior:
[0801] The server receives the HTTP request and passes the data to the analysis module.
[0802] A natural language processing engine analyzes the text data and identifies the necessary keywords and context.
[0803] An image analysis engine processes the storyboard and extracts specific elements and scenes.
[0804] Step 4:
[0805] The server uses an emotion engine to recognize emotions contained in the received text and storyboard data.
[0806] Specific behavior:
[0807] The emotion engine performs sentiment analysis on the text to identify the user's intentions and emotions.
[0808] The emotion engine performs an emotion analysis of the storyboard and identifies the emotional nuances of the depicted scene.
[0809] Step 5:
[0810] The server searches for and acquires the necessary materials based on the analysis and emotion analysis results, including searching for illustrations and audio files from a materials database.
[0811] Specific behavior:
[0812] The server executes an SQL query against the material database and searches for materials using search criteria based on the analysis results and emotion analysis results.
[0813] The path of the corresponding material file is obtained and temporarily saved.
[0814] Step 6:
[0815] The server generates payment information for the used material, including information about the material provider and a totaling process for the usage fee.
[0816] Specific behavior:
[0817] The server obtains information about the material provider and the usage fee from the material database.
[0818] Billing information is generated based on the acquired information and presented to the user.
[0819] Step 7:
[0820] The user confirms the payment information provided by the server and makes the payment, after which the material becomes officially available to the user.
[0821] Specific behavior:
[0822] The server generates billing information and sends it to the user in HTML format.
[0823] The user visits the payment page and enters the required payment information.
[0824] The terminal sends the payment information to the server and waits for the payment to be completed.
[0825] Step 8:
[0826] The server verifies that payment has been made and prepares the video for generation, thereby confirming permission to legally use the material.
[0827] Specific behavior:
[0828] Your server uses the payment gateway API to confirm the payment is successful.
[0829] Once payment is confirmed, the permission to use the materials will be updated in our internal system.
[0830] Step 9:
[0831] The server uses the analysis results and the acquired materials to send the data to an AI video generation service, which then automatically generates the video.
[0832] Specific behavior:
[0833] The server sends the necessary material files and analysis results to the API of the AI video generation service.
[0834] The AI video generation service generates a video based on the provided data and returns the generated video file to the server.
[0835] Based on the emotion engine, the tone of the video's narration and visual expression are adjusted.
[0836] Step 10:
[0837] The server sends a download link for the generated video to the user's device, allowing the user to access the video.
[0838] Specific behavior:
[0839] The server creates a URL for the generated video file and generates a link in an HTML email or on the dashboard.
[0840] When users click on the link, they can download or stream the video on their device.
[0841] Example 2
[0842] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0843] Conventional video generation systems require a lot of manual work when generating videos based on user-entered text and storyboards. Furthermore, they lacked a means to reflect the user's emotions in the video, making it difficult to generate personalized, high-quality videos. This increased the time and cost required for video production, making them unusable for many users.
[0844] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving text and storyboard data input by the user, means for analyzing the text and storyboard, means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results, means for generating payment information for use of the materials, means for confirming that the user has completed payment, means for automatically generating a video using the materials and the analysis results and emotion analysis results, and means for providing the generated video to the user. This enables the user to automatically generate high-quality videos that reflect emotions with little effort and quickly use them.
[0845] "User" refers to the person who inputs text and storyboards to use the system.
[0846] "Data receiving means" refers to the function of receiving text and storyboard data entered by the user.
[0847] "Data analysis means" refers to a function for analyzing received text and storyboards.
[0848] "Emotion analysis means" refers to a function that identifies the user's emotions contained in text or storyboards based on the analysis results.
[0849] "Material search means" refers to a function for searching for and acquiring necessary materials based on the analysis results and emotion analysis results.
[0850] "Material acquisition means" refers to a function for acquiring materials identified by the search means from a database or external service.
[0851] The "payment information generating means" refers to a function for generating payment information for the materials used.
[0852] "Payment Verification Method" refers to the function that verifies that a user has completed a payment.
[0853] "Video generation means" refers to a function that automatically generates videos based on acquired materials and analysis results.
[0854] "Video providing means" refers to the function of providing the generated video to the user.
[0855] The present invention relates to a system that automatically generates high-quality videos based on text and storyboard data entered by a user, and further enables personalized video creation by recognizing the user's emotions and reflecting them in the videos. This system includes the following components.
[0856] First, the user accesses a dedicated web page or application using their device and enters text and storyboards for the video they want to create. This input data is then sent from the device to the server via an HTTP request.
[0857] The server temporarily stores the received data and then begins processing to analyze it. Natural language processing (NLP) technology is used to analyze the text data, specifically using Python's NLTK and Spacy libraries. Image analysis technology is used to analyze the storyboards, specifically using the OpenCV library for contour detection and shape recognition.
[0858] Furthermore, we use sentiment analysis tools to identify the emotions contained in the analyzed text and storyboards. For sentiment analysis of text, we use natural language processing services such as Google Cloud NLP, and for sentiment analysis of storyboards, we use machine learning frameworks such as TensorFlow.
[0859] Based on the analysis results and sentiment analysis results, the necessary materials (e.g., illustrations and audio) are searched and acquired from a stock material database. The material database is connected to the API of an online stock material site, allowing necessary materials to be searched and acquired programmatically.
[0860] Next, the server generates payment information for the materials used. This information is generated by retrieving information about the material provider and the usage fee from a database and adding them up. The database used here is MySQL, a common relational database.
[0861] Once the user's payment is confirmed, the server generates a video using AI video generation techniques, adjusting the expressions in the video based on the results of sentiment analysis. The video is generated using OpenAI's CLIP model and DeepAI's API.
[0862] Finally, the generated video is served to the user: the server sends a download link for the generated video to the user's device, which the user can use to download or stream the video.
[0863] Specific examples
[0864] Example 1: Manual video for manufacturing industry
[0865] The user enters the "steps to assemble part A" in text on their device and uploads a simple sketch as a storyboard. The server receives this and performs natural language processing and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The emotion engine then analyzes the user's intentions and emotions from the entered text and generates a narration that reflects kindness or strictness. The server searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[0866] Example 2: Educational lesson videos
[0867] A user who is an educator enters a text description of the solar system and uploads a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio for the solar system. The emotion engine also performs emotional analysis of the entered text and selects passionate narration or a calm tone. The server searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[0868] Prompt Sentence Examples
[0869] Example 1: Manual video for manufacturing industry
[0870] "Please upload a written description of the steps to assemble part A and a simple sketch as a storyboard. After analysis, we will identify the illustrations of the parts and the audio material of the assembly steps and generate a narration that reflects kindness or strictness."
[0871] Example 2: Educational lesson videos
[0872] "Enter a description of the solar system in text and upload a simple storyboard for each planet. After analysis, we'll generate a video with illustrations of the solar system and a choice of passionate or calming narration."
[0873] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0874] Step 1: Data entry
[0875] The user uses a terminal to access a dedicated web page or application and inputs text and storyboards related to the video they want to generate. Specifically, the user inputs the text "Assembly procedure for part A" and uploads the storyboard for part A. The input format is a text file and an image file. This results in the text and storyboards as input data.
[0876] Step 2: Send data
[0877] Data entered by the user is sent from the terminal to the server via an HTTP request. For example, a POST request is used to send text data and image data to the server's endpoint. The data is received by the server and temporarily stored in a database or system memory.
[0878] Step 3: Text analysis
[0879] The server analyzes the received text data using natural language processing (NLP) techniques, specifically Python's NLTK and Spacy libraries. For example, it tokenizes the text, tags parts of speech, and performs syntactic analysis. This process extracts syntactic and semantic data from the input text.
[0880] Step 4: Storyboard Analysis
[0881] The server analyzes the storyboard it receives using image analysis technology. The OpenCV library is used for image analysis, specifically contour detection and shape recognition. For example, it can detect a specific shape (e.g., the shape of part A) from the uploaded storyboard and extract the corresponding data.
[0882] Step 5: Sentiment Analysis
[0883] The server uses an emotion engine to identify emotions contained in the text and storyboards. It uses machine learning frameworks such as Google Cloud NLP for emotion analysis of text and TensorFlow for emotion analysis of storyboards. Specifically, it uses an emotion model to identify emotions such as "joy," "sadness," and "anger."
[0884] Step 6: Material Search
[0885] Based on the analysis and emotion analysis results, the server searches for the necessary materials from a material database. The material database is connected to an online stock material site using an API. For example, it searches for "illustration of part A" or "gentle-toned voice narration" to identify appropriate materials.
[0886] Step 7: Obtaining Materials
[0887] The server retrieves the searched material. This involves downloading data from the stock material site via API. The retrieved material is then stored in a directory or database managed by the server.
[0888] Step 8: Generate payment information
[0889] The server generates payment information for the materials used. The payment information is generated by retrieving information about the material provider and the usage fee from a database and adding them up. Specifically, a specific SQL query is executed using a MySQL database to calculate the usage fee.
[0890] Step 9: Confirm payment
[0891] The server presents the generated billing information to the user. The user enters the payment information (e.g., credit card information). The server processes the payment using an online payment system (e.g., PayPal or Stripe). If the payment is successful, the server confirms it.
[0892] Step 10: Video Generation
[0893] After the user's payment is confirmed, the server uses the analysis results and the acquired materials to execute the AI video generation service. The emotion analysis results are taken into account and the expressions in the video are adjusted. The video is generated using OpenAI's CLIP model and DeepAI's API. The generated video file is saved in the appropriate format.
[0894] Step 11: Submit your video
[0895] The server provides the generated video to the user by generating a download link for the video and sending it to the user's device. The user can then use the link to download or stream the video.
[0896] (Application example 2)
[0897] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0898] Conventional video generation systems simply generate videos based on user-entered text and storyboards, making it difficult to generate personalized videos that reflect the user's emotions. Furthermore, there is a lack of systems that efficiently execute the entire process of delivering generated videos to target users. For this reason, the advertising industry needs a way to quickly generate and deliver effective video ads that match the interests and emotions of target users.
[0899] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0900] In this invention, the server includes means for receiving text and storyboard data input by the user, means for analyzing the text and storyboard, and means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results. This makes it possible to generate high-quality videos that reflect emotions based on the user's input data and distribute them to target users.
[0901] "Means for receiving text and storyboard data entered by the user" refers to a function that allows a user to use a dedicated application or web page via a terminal to enter text and storyboard data and send it to the server.
[0902] The "means for analyzing text and storyboards" refers to a means for analyzing received text and storyboard data using natural language processing and image analysis techniques to understand the content.
[0903] "Means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results" refers to means for searching for and acquiring necessary video and audio materials from a materials database based on the analyzed data and the results of emotion analysis.
[0904] The "means for generating payment information for use of material" is a function for calculating the fee for the used material and generating corresponding payment information.
[0905] The "means for confirming that the user has completed the payment" refers to a means for confirming that the user has made the payment based on the generated payment information.
[0906] "Means for automatically generating videos using materials, analysis results, and emotion analysis results" refers to means for automatically generating videos using AI video generation technology based on the acquired materials and analysis results.
[0907] "Means for providing the generated video to the user" means means for providing the generated video to the user and enabling the user to download or stream the video.
[0908] "Means for distributing the generated video to target users" refers to means for distributing the generated video to target users designated by the advertiser, thereby increasing the effectiveness of the advertisement.
[0909] The embodiments for carrying out the present invention are described in detail below.
[0910] The system aims to enable advertisers to generate personalized video ads using smartphone applications and efficiently deliver them to target users.
[0911] 1. User Input and Data Receipt
[0912] Advertisers use a dedicated smartphone application to input text and storyboards for their advertisements, which are then sent to the server via HTTP requests.
[0913] 2. Data Analysis
[0914] The server analyzes the received text and storyboard data using a natural language processing library (e.g., SpaCy) and an image analysis library (e.g., OpenCV). The analysis includes syntactic analysis of the text data and feature extraction of the image data.
[0915] 3. Emotion recognition
[0916] The emotion engine analyzes emotions based on the analyzed data. A common emotion analysis API (e.g., IBM Watson Natural Language Understanding, Google Cloud Natural Language API) is used as the emotion engine.
[0917] 4. Material Search and Acquisition
[0918] Based on the results of sentiment analysis and data analysis, the server searches for and obtains the necessary video and audio materials from a material database (e.g., Shutterstock API).
[0919] 5. Payment Information Generation and Verification
[0920] The server generates payment information based on the information of the materials used and presents it to the advertiser. Payment processing is performed using an electronic payment service (e.g., Stripe API). Once the advertiser completes payment, the information is recorded on the server.
[0921] 6. Video Generation
[0922] Once payment is confirmed, the server uses the acquired materials and analysis results to automatically generate a video using an AI video generation service (e.g., Synthesia). The generated video reflects narration and direction based on the results of the emotion analysis.
[0923] 7. Video Provision and Distribution
[0924] The server provides the generated video advertisement to the advertiser, who then sends a link to the advertisement via a smartphone application, allowing the advertiser to distribute the video to target users.
[0925] Specific examples
[0926] Example 1: Advertising video for new product "Express Coffee"
[0927] Advertisers enter text and sketch images into the application that explain the features of their new product, "Express Coffee."
[0928] Example sentence: "Express Coffee provides fast, delicious coffee for busy mornings."
[0929] Storyboard: "Image of an Express Coffee package and coffee being poured into a cup"
[0930] Emotion: We want to convey a sense of comfort and trust to the user.
[0931] The server analyzes the input data and searches for and retrieves appropriate content from a content database based on the results of sentiment analysis. Once payment is completed, an AI video generation service is used to generate a video ad that reflects the specified sentiment. The generated video ad link is provided to the advertiser through the application, and the advertiser uses this link to deliver the video to target users.
[0932] This system enables advertisers to quickly generate high-quality, emotionally relevant video ads and deliver them effectively to target users.
[0933] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0934] Step 1:
[0935] Input and Data Reception
[0936] Users use a smartphone application to input text and storyboards to be used in advertising videos. The input data is collected through a form in the application and sent to the server as an HTTP request. The server temporarily stores the received data.
[0937] Specific operation: The user uploads text data (e.g., advertising copy) and image data (e.g., product sketch) into the application's input form and sends it to a dedicated API endpoint.
[0938] Step 2:
[0939] Data analysis
[0940] The server analyzes the received text and storyboard data. Specifically, it analyzes the text using a natural language processing library (e.g., SpaCy) and the storyboard using an image analysis library (e.g., OpenCV). Text analysis is used to understand the meaning and structure of the text, and image analysis is used to extract the features of the storyboard.
[0941] Input and Output: The input data are the received text and storyboard, and the output is the analyzed text data and image data features.
[0942] Specific operation: Using the SpaCy library, tokenize text data, assign POS tags, and analyze dependencies. Using the OpenCV library, detect edges and extract feature points from storyboard images.
[0943] Step 3:
[0944] emotion recognition
[0945] The server uses an emotion engine based on the analyzed data to analyze emotions from the text and storyboard. It uses an emotion analysis API (e.g., IBM Watson Natural Language Understanding, Google Cloud Natural Language API) to identify the user's emotions.
[0946] Input and Output: The input is the analyzed text data and image data features, and the output is the sentiment analysis results.
[0947] Specific operation: The analyzed text data and image data features are sent to the emotion analysis API, and emotions are categorized (e.g., joy, sadness, surprise, etc.) and received.
[0948] Step 4:
[0949] Material Search and Acquisition
[0950] Based on the results of emotion analysis and data analysis, the server searches and retrieves appropriate video and audio materials from a material database (e.g., Shutterstock API).
[0951] Input and output: The input is the emotion analysis results and data analysis results, and the output is the acquired video and audio materials.
[0952] Specific operation: Creates a search query for the Shutterstock API based on the results of sentiment analysis and data analysis, and retrieves the necessary materials (images, audio, video).
[0953] Step 5:
[0954] Payment information generation and verification
[0955] The server generates payment information based on the information about the used materials and presents it to the user. It processes the payment using an electronic payment service (e.g., Stripe API) and confirms that the payment has been completed.
[0956] Input and Output: The input is the information of the material acquired and the output is the status of payment confirmation.
[0957] Specific operation: Calculate the total amount based on the price information of the obtained materials, send a payment request to the user through the Stripe API, and check the status of whether the payment has been completed.
[0958] Step 6:
[0959] Video Generation
[0960] Once payment is confirmed, the server uses the acquired materials and analysis results to automatically generate a video using an AI video generation service (e.g., Synthesia), including narration and direction based on the results of the emotion analysis.
[0961] Input and Output: The input is the acquired material and analysis results, and the output is the generated video file.
[0962] Specific operation: The acquired material, analysis results, and a prompt for a video to be generated based on the emotion analysis results are sent to the Synthesia API, and the generated video is retrieved.
[0963] Step 7:
[0964] Video provision and distribution
[0965] The generated video advertisement is provided to the user (advertiser) from the server, and a link is sent to the user via a smartphone application. The user can then distribute the video to target users.
[0966] Input and Output: The input is the generated video file and the output is the video link provided to the user and the delivery status to the target user.
[0967] Specific operation: The generated video is stored in the server storage, and a download link is generated and sent to the user. The user receives the link and the video is distributed to the target user.
[0968] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0969] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0970] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0971] [Third embodiment]
[0972] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0973] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0974] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0975] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0976] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0977] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0978] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0979] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0980] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0981] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0982] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0983] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0984] The present invention relates to a system that automatically generates high-quality animation based on text and storyboards input by a user. This system has the following main functions:
[0985] 1. Data reception function
[0986] The user uses a terminal to access a dedicated web page or application, input text and storyboards, and this data is sent to the server via an HTTP request.
[0987] 2. Data analysis function
[0988] The server analyzes the received text and storyboard data. Specifically, it uses natural language processing technology to analyze the text and image analysis technology to analyze the storyboard. This allows it to identify the materials needed to generate the video.
[0989] 3. Material search and acquisition function
[0990] The server searches and retrieves the necessary materials, such as illustrations and audio, from a materials database based on the analysis results. This process involves searching the materials database using SQL queries.
[0991] 4. Payment information generation function
[0992] The server generates payment information for the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates total billing information.
[0993] 5. Payment confirmation function
[0994] The generated billing information is presented to the user for payment confirmation, and the payment is processed based on the payment information provided by the user. The server then confirms that the payment has been completed.
[0995] 6. Video generation function
[0996] Once payment is complete, the server uses the analysis results and acquired materials to run the AI video generation service and generate a video. The AI video generation service automatically generates a video according to the provided materials and instructions.
[0997] 7. Video provision function
[0998] Finally, the server provides the generated video to the user by sending a download link for the generated video file to the user's device, which the user can use to download or stream the video.
[0999] Specific examples
[1000] Example 1: Manual video for manufacturing industry
[1001] A user in the manufacturing industry enters the steps for assembling part A in text and uploads a simple sketch as a storyboard. The server receives this and performs text and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The server then searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[1002] Example 2: Educational lesson videos
[1003] Educators enter a text description of the solar system and upload a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio of the solar system. It then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[1004] This system will enable users to easily generate and use high-quality videos, significantly reducing the time and cost required for video production and is expected to be used in a wide range of applications.
[1005] The processing flow will be explained below.
[1006] Step 1:
[1007] The user uses the device to access a dedicated web page or application, enter text, and upload storyboards, thereby providing information about the content of the video.
[1008] Specific behavior:
[1009] The user logs in to a dedicated web page or application.
[1010] The user enters a description of the video in a text box.
[1011] The user opens a file selection dialog and selects an image file to upload a storyboard.
[1012] Once you have completed the entry and upload, click the "Submit" button.
[1013] Step 2:
[1014] The device sends the input text and storyboard data to the server as an HTTP request.
[1015] Specific behavior:
[1016] The device encodes the input text and storyboard as JSON format or multipart form data.
[1017] The encoded data is sent as an HTTP POST request to the server's API endpoint.
[1018] Step 3:
[1019] The server analyzes the received text and storyboard data, specifically using natural language processing technology for the text and image analysis technology for the storyboard.
[1020] Specific behavior:
[1021] The server receives the HTTP request and passes the data to the analysis module.
[1022] A natural language processing engine analyzes the text data and identifies the necessary keywords and context.
[1023] An image analysis engine processes the storyboard and extracts specific elements and scenes.
[1024] Step 4:
[1025] The server searches for and acquires the necessary materials based on the analysis results, including searching for illustrations and audio files from a materials database.
[1026] Specific behavior:
[1027] The server executes an SQL query against the material database and searches for materials using search criteria based on the analysis results.
[1028] The path of the corresponding material file is obtained and temporarily saved.
[1029] Step 5:
[1030] The server generates payment information for the used material, including information about the material provider and a totaling process for the usage fee.
[1031] Specific behavior:
[1032] The server obtains information about the material provider and the usage fee from the material database.
[1033] Billing information is generated based on the acquired information and presented to the user.
[1034] Step 6:
[1035] The user confirms the payment information provided by the server and makes the payment, after which the material becomes officially available to the user.
[1036] Specific behavior:
[1037] The server generates billing information and sends it to the user in HTML format.
[1038] The user visits the payment page and enters the required payment information.
[1039] The terminal sends the payment information to the server and waits for the payment to be completed.
[1040] Step 7:
[1041] The server verifies that payment has been made and prepares the video for generation, thereby confirming permission to legally use the material.
[1042] Specific behavior:
[1043] Your server uses the payment gateway API to confirm the payment is successful.
[1044] Once payment is confirmed, the permission to use the materials will be updated in our internal system.
[1045] Step 8:
[1046] The server uses the analysis results and the acquired materials to send the data to an AI video generation service, which then automatically generates the video.
[1047] Specific behavior:
[1048] The server sends the necessary material files and analysis results to the API of the AI video generation service.
[1049] The AI video generation service generates a video based on the provided data and returns the generated video file to the server.
[1050] Step 9:
[1051] The server sends a download link for the generated video to the user's device, allowing the user to access the video.
[1052] Specific behavior:
[1053] The server creates a URL for the generated video file and generates a link in an HTML email or on the dashboard.
[1054] When users click on the link, they can download or stream the video on their device.
[1055] Example 1
[1056] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1057] Demand for high-quality video content is increasing in many fields today. However, video production requires specialized skills and is time-consuming and costly. Furthermore, there are not enough methods available for users to easily create high-quality videos based on their own ideas and information. Therefore, there is a need for a system that can efficiently create and provide high-quality videos without specialized knowledge.
[1058] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1059] In this invention, the server includes means for receiving information and visual material data input by a user, means for analyzing the information and visual material, and means for searching for and acquiring necessary data based on the analysis results, thereby enabling users without specialized knowledge to easily turn their ideas into high-quality videos.
[1060] "User" refers to an individual or organization that utilizes this system to provide information and visual materials and request the creation of a video.
[1061] "Visual materials" refers to materials that convey information visually, such as storyboards and image files that users upload to the system.
[1062] "Data receiving means" refers to a mechanism by which the server receives information and visual materials entered by the user via the network.
[1063] "Analysis means" refers to the process of breaking down and interpreting received information and visual materials, and extracting and identifying the necessary data.
[1064] "Search and acquisition means" refers to the function of searching the database for the necessary data based on the analysis results and acquiring the appropriate materials.
[1065] The "fee information generating means" is a mechanism for calculating the fee for the user based on the data and materials used and creating billing information.
[1066] "Payment confirmation means" refers to a process for confirming that the user has completed payment.
[1067] "Video generation means" refers to a function that automatically generates high-quality video using AI technology, etc., based on the analysis results and acquired materials.
[1068] The "video providing means" refers to a mechanism for providing the generated video in a form that allows users to download or stream the video.
[1069] The present invention relates to a system for automatically generating and providing high-quality video based on user-provided information and visual materials. The system includes the following main components:
[1070] Data reception
[1071] The user uses a device to access a dedicated web page or application. The user inputs and uploads text information and visual materials such as storyboards for the video they want to create. The device then sends this data to the server via an HTTP POST request. The server then receives the data provided by the user.
[1072] Data analysis
[1073] The server analyzes the received text and visual materials. For text data, it uses a natural language processing engine (e.g., SpaCy, BERT) to analyze the text and extract keywords and important location information. For image data, it uses image analysis technology (e.g., OpenCV, TensorFlow) to analyze the storyboard and identify objects. This identifies the materials needed to generate the video.
[1074] Material Search and Acquisition
[1075] Based on the analysis results, the server searches for the necessary data (illustrations, audio files, etc.) from the material database and retrieves them using SQL queries. These materials are used as components necessary for video generation.
[1076] Payment information generation
[1077] The server generates fee information for the use of the material. Specifically, it retrieves information about the material provider and the usage fee from the database, adds them up, and generates an invoice to be presented to the user.
[1078] Payment confirmation
[1079] Provide the user with a link to the payment page and confirm that they wish to complete the payment. The server receives notification that the payment has been completed and proceeds to the next processing step.
[1080] Image Generation
[1081] The server generates videos using AI video generation services (e.g., OpenAI's DALL-E, GPT-4) based on the analysis results and acquired materials. The video generation process is fully automated, and high-quality videos are generated according to the instructions provided by the user.
[1082] Video provided by
[1083] Finally, the server provides the user with a download link for the generated video, which the user can use to download or stream the video, allowing the user to obtain high-quality video in a convenient way.
[1084] Specific examples
[1085] Example 1: Manual video for manufacturing industry
[1086] The user inputs the assembly instructions and associated storyboard for a specific product from their device. The server analyzes the instructions, identifies and acquires the necessary illustrations and audio materials, and generates payment information. Once the user completes payment, the server generates a video using an AI video generation service and provides the user with a download link for the final video.
[1087] prompt:
[1088] "Assembly steps for part A:
[1089] 1. Take out part A.
[1090] 2. Connect to part B.
[1091] 3. Tighten the screws.
[1092] "
[1093] Image: [Illustration: Part A, Part B, Screw diagram]
[1094] Example 2: Educational lesson videos
[1095] The user, an educator, enters a text description of the solar system and uploads a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio of the solar system. It then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user.
[1096] prompt:
[1097] "Solar System Description:
[1098] 1. The sun is at the center and the other planets revolve around it.
[1099] 2. Mercury is the planet closest to the sun.
[1100] 3. Venus is the brightest star known.
[1101] "
[1102] Image: [Diagram showing the solar system, with diagrams of each planet]
[1103] This allows users to easily create high-quality videos without specialized knowledge or skills, and provides a system that can be used for a variety of purposes.
[1104] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1105] Step 1: Receiving data
[1106] Users use their devices to access a dedicated web page or application and input and upload written information about the video they want to create and visual materials such as storyboards.
[1107] Input: Text information and storyboard data entered by the user
[1108] The terminal sends this data to the server as an HTTP POST request.
[1109] Output: Text information and storyboard data received by the server
[1110] Step 2: Data analysis
[1111] The server analyzes the received text information using a natural language processing engine (e.g., SpaCy, BERT) to extract keywords and important location information.
[1112] Input: Received text information
[1113] Output: Extracted keywords and location information
[1114] The server analyzes the storyboard using image analysis technology (e.g., OpenCV, TensorFlow) and identifies the objects.
[1115] Input: Received storyboard data
[1116] Output: Identified objects
[1117] Step 3: Search and acquire materials
[1118] Based on the analysis results, the server searches for the necessary data (illustrations, audio files, etc.) from the material database.
[1119] Input: Analysis results (extracted keywords and identified objects)
[1120] Output: Searched material data
[1121] The server uses an SQL query to search the materials database and retrieve the appropriate materials.
[1122] Input: SQL query
[1123] Output: Acquired material data
[1124] Step 4: Generate payment information
[1125] The server generates fee information based on the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates a combined invoice.
[1126] Input: Acquired material data
[1127] Output: Generated invoice
[1128] Step 5: Payment confirmation
[1129] The server presents the generated invoice to the user and provides a link to a payment page.
[1130] Input: Generated Invoice
[1131] Output: Payment page link
[1132] The user enters payment information and completes the payment.
[1133] Input: User's payment information
[1134] Once the server receives notification that payment has been completed, it proceeds to the next processing step.
[1135] Output: Confirmation of successful payment
[1136] Step 6: Image generation
[1137] The server generates video using an AI video generation service (e.g., OpenAI's DALL-E, GPT-4) based on the analysis results and acquired materials.
[1138] Input: Analysis results and acquired materials
[1139] Output: Generated video file
[1140] The server stores the generated video files.
[1141] Input: Generated video file
[1142] Output: Saved video file
[1143] Step 7: Provide footage
[1144] The server provides the user with a download link for the generated video.
[1145] Input: Saved video file
[1146] Output: Video download link
[1147] Users can use this link to download or stream the video.
[1148] Input: Video download link
[1149] Output: Downloaded or streamed video
[1150] Through these processing steps, users can easily create high-quality videos for a variety of uses, even without specialized knowledge or skills.
[1151] (Application example 1)
[1152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1153] Conventional video production systems require a lot of time and effort for the entire process of collecting, editing, and creating the materials needed to create a video, making it difficult for average users to easily create high-quality videos. Furthermore, the payment process for using the materials is complicated, placing a burden on users. The present invention aims to solve these problems by providing a system that allows users to easily create high-quality videos and complete payments smoothly.
[1154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1155] In this invention, the server includes means for receiving text and storyboard data entered by a user, means for analyzing the text and storyboard, means for searching for and acquiring necessary materials based on the analysis results, means for generating payment information for the use of the materials, means for confirming that the user has completed payment, means for automatically generating a video using the materials and the analysis results, means for providing the generated video to the user, means for providing a user interface integrated into a smartphone application, and means for using AI to generate a video based on materials acquired from a materials database and confirm payment. This allows users to generate high-quality videos in a short amount of time and complete payment procedures easily and quickly.
[1156] "User-input text" refers to data based on text provided by a user through the system's input interface.
[1157] A "storyboard" is drawing data based on a visual guide entered by the user.
[1158] A "receiving means" is a part of the system that has the function of capturing data sent by a user.
[1159] The "analyzing means" is a part of the system that has the function of analyzing the received data and extracting the necessary information.
[1160] "Necessary materials" refers to data such as images and audio required to generate video.
[1161] A "search and retrieval means" is a part of a system that has the ability to search for information in a database and retrieve requested material.
[1162] A "means for generating payment information" is a part of the system that has the functionality to calculate fees for use of material and generate that information.
[1163] A "means for confirming payment completion" is a part of the system that has the function of confirming that a user's payment has been successfully made.
[1164] The "means for automatically generating a video" is a part of a system that has the function of automatically creating a video using the acquired materials and analysis results.
[1165] A "means for providing videos to users" is a part of the system that has the function of making the generated videos available to users.
[1166] A "smartphone application" is a program that runs on a smartphone and provides an interface that can be operated directly by the user.
[1167] A "user interface" is a feature that serves as an entry point for a user to interact with a system.
[1168] A "material database" is a data storage that stores materials necessary for video generation.
[1169] "AI-based means" refers to a part of a system that has the function of generating videos using artificial intelligence technology.
[1170] The embodiment of the present invention is a system that automatically generates high-quality videos based on text and storyboards entered by the user. This system is composed of a server and a user's terminal (smartphone application).
[1171] System Program
[1172] The system includes the following key features:
[1173] 1. Data reception function: The server receives text and storyboards sent from the user's device. The user inputs and sends this data via a smartphone application.
[1174] 2. Data analysis function: The server analyzes the received text using natural language processing technology (e.g., spaCy or NLTK) and analyzes the storyboard using image analysis technology (e.g., OpenCV or TensorFlow).
[1175] 3. Material search and retrieval function: Based on the analysis results, the server searches and retrieves related images, audio, and other materials from the material database. Here, the database search is performed using SQL queries.
[1176] 4. Payment information generation function: The server calculates the usage fee for the acquired material and presents payment information to the user. At this time, information about the material provider is also obtained.
[1177] 5. Payment confirmation function: The server processes the payment based on the payment information provided by the user and confirms that the payment has been completed. Possible payment services used include Stripe and PayPal.
[1178] 6. Video generation function: The server automatically generates videos using AI based on the acquired materials and analysis results. AI video generation uses services such as DeepArt and Runway ML.
[1179] 7. Video provision function: The generated video is provided to the user from the server. The user can stream or download the video via the URL of the generated video.
[1180] Natural language description of the process
[1181] Data reception and analysis: The text and storyboards entered by the user using the smartphone application are sent to the server via HTTP requests. The server stores the received data, analyzes the text using natural language processing (NLP) technology, and analyzes the storyboards using image analysis technology.
[1182] Material search and retrieval: Based on the analysis results, the server performs a database search to retrieve the required image and audio materials. The material database stores materials by multiple categories, so the required materials can be quickly retrieved using the appropriate SQL query.
[1183] Payment processing: Calculate the usage fee for the acquired materials and present the payment information to the user. The server confirms that the user has completed the payment and proceeds to the next step.
[1184] Video generation and provision: Using AI technology, a video is generated by combining the necessary materials and analysis results. The generated video is provided to the user as a URL link. The user can view or download the generated video via the link.
[1185] Examples of concrete examples and prompts
[1186] Examples:
[1187] In the education field, users can enter a text description of each planet in the solar system and upload a corresponding hand-drawn sketch of the planet as a storyboard. This data is then analyzed to retrieve images of the corresponding planet from, for example, a NASA image database, and audio material from LibriVox. The resulting video can then be used in the user's educational activities.
[1188] Example prompt sentence:
[1189] "Generate a video about each planet in the solar system. Create a high-quality video based on the text and storyboard below.
[1190] Text: The solar system has the sun at its center and the planets that orbit it are Mercury, Venus, Earth, Mars, Jupiter, Saturn, Uranus, and Neptune.
[1191] Storyboard: Hand-drawn planet sketch (image data)
[1192] The above is a specific embodiment of the present invention, which allows users to easily generate high-quality videos and use them immediately.
[1193] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1194] Step 1:
[1195] The user inputs text and storyboards using a smartphone application and sends them.
[1196] Input: Text data, storyboard data
[1197] Output: HTTP request data
[1198] Specific operation: The user enters text into the application's input form and uploads image files that will serve as storyboards. This data is sent from the device to the server as an HTTP request.
[1199] Step 2:
[1200] The server receives the HTTP request and stores the data.
[1201] Input: HTTP request data
[1202] Output: Saved text data, saved storyboard data
[1203] Specific operation: The server stores the received data in temporary storage and prepares it for the next analysis step.
[1204] Step 3:
[1205] The server analyzes the received data.
[1206] Input: Saved text data, saved storyboard data
[1207] Output: Analyzed text data, analyzed storyboard data
[1208] Specific operation: The server analyzes the text using natural language processing technology (e.g., spaCy or NLTK) to extract keywords and structural information, and also analyzes the storyboard using image analysis technology (e.g., OpenCV or TensorFlow) to extract the objects and layout information contained therein.
[1209] Step 4:
[1210] The server searches the material database based on the analysis results and retrieves the necessary images and audio materials.
[1211] Input: Analyzed text data, analyzed storyboard data
[1212] Output: Acquired image material data, acquired audio material data
[1213] Specific operation: The server uses an SQL query to search the material database based on the analysis results and retrieve relevant images and audio materials.
[1214] Step 5:
[1215] The server calculates the usage fee for the acquired material and generates payment information.
[1216] Input: Acquired image material data, acquired audio material data
[1217] Output: Payment information data
[1218] Specific operation: The server retrieves the information of the material provider and the usage fee data, combines them, and generates payment information, which is then presented to the user.
[1219] Step 6:
[1220] The server processes the payment based on the user's payment information and confirms that the payment has been completed.
[1221] Input: Payment information data, user's payment information
[1222] Output: Payment confirmation data
[1223] Specific operation: The server uses a payment service such as Stripe or PayPal to check whether the user's payment has been successfully completed. If the confirmation is successful, it proceeds to the next step.
[1224] Step 7:
[1225] Using the materials acquired by the server and the analysis results, videos are automatically generated using AI.
[1226] Input: Acquired image material data, acquired audio material data, analyzed text data, analyzed storyboard data
[1227] Output: Generated video data
[1228] Specific operation: The server uses an AI video generation service (e.g., DeepArt or Runway ML) to generate a video that combines the acquired materials and analysis results.
[1229] Step 8:
[1230] The server provides the generated video to the user.
[1231] Input: Generated video data
[1232] Output: Video data provided to users (URL link, etc.)
[1233] Specific operation: The server saves the generated video in cloud storage and provides the URL to the user, who can use this URL to stream or download the video.
[1234] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1235] This invention relates to a system that automatically generates high-quality videos based on text and storyboards entered by the user, as well as a system that recognizes the user's emotions and reflects them in the videos. This system has the following main functions:
[1236] 1. Data reception function
[1237] The user uses a terminal to access a dedicated web page or application, input text and storyboards, and this data is sent to the server via an HTTP request.
[1238] 2. Data analysis function
[1239] The server analyzes the received text and storyboard data using natural language processing and image analysis technologies.
[1240] 3. Emotion engine function
[1241] The server is equipped with an emotion engine for recognizing emotions contained in text and storyboard data. The emotion engine performs emotion analysis when analyzing text, and also performs emotion analysis on storyboards.
[1242] 4. Material search and acquisition function
[1243] The server searches and retrieves the necessary materials from a materials database based on the analysis and emotion analysis results, including appropriate illustrations, audio, and other materials.
[1244] 5. Payment information generation function
[1245] The server generates payment information for the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates total billing information.
[1246] 6. Payment confirmation function
[1247] The generated billing information is presented to the user for payment confirmation, and the payment is processed based on the payment information provided by the user. The server then confirms that the payment has been completed.
[1248] 7. Video generation function
[1249] Once payment is complete, the server uses the analysis results and acquired materials to run the AI video generation service and generate a video, including a function to adjust the expressions in the video based on the emotion engine.
[1250] 8. Video provision function
[1251] Finally, the server provides the generated video to the user by sending a download link for the generated video file to the user's device, which the user can use to download or stream the video.
[1252] Specific examples
[1253] Example 1: Manual video for manufacturing industry
[1254] A user working in the manufacturing industry enters the steps for assembling part A in text and uploads a simple sketch as a storyboard. The server receives this and performs natural language processing and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The emotion engine then analyzes the user's intentions and emotions from the input text and generates a narration that reflects kindness or strictness. The server searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[1255] Example 2: Educational lesson videos
[1256] Educators enter a text description of the solar system and upload a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio for the solar system. The emotion engine then performs an emotional analysis of the entered text and selects passionate narration or a calm tone. The system then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[1257] This system enables users to easily generate and use high-quality, emotionally-reflective videos, significantly reducing the time and cost required for video production and is expected to be used in a wide range of applications.
[1258] The processing flow will be explained below.
[1259] Step 1:
[1260] The user uses the device to access a dedicated web page or application, enter text, and upload storyboards, thereby providing information about the content of the video.
[1261] Specific behavior:
[1262] The user logs in to a dedicated web page or application.
[1263] The user enters a description of the video in a text box.
[1264] The user opens a file selection dialog and selects an image file to upload a storyboard.
[1265] Once you have completed the entry and upload, click the "Submit" button.
[1266] Step 2:
[1267] The device sends the input text and storyboard data to the server as an HTTP request.
[1268] Specific behavior:
[1269] The device encodes the input text and storyboard as JSON format or multipart form data.
[1270] The encoded data is sent as an HTTP POST request to the server's API endpoint.
[1271] Step 3:
[1272] The server analyzes the received text and storyboard data, specifically using natural language processing technology for the text and image analysis technology for the storyboard.
[1273] Specific behavior:
[1274] The server receives the HTTP request and passes the data to the analysis module.
[1275] A natural language processing engine analyzes the text data and identifies the necessary keywords and context.
[1276] An image analysis engine processes the storyboard and extracts specific elements and scenes.
[1277] Step 4:
[1278] The server uses an emotion engine to recognize emotions contained in the received text and storyboard data.
[1279] Specific behavior:
[1280] The emotion engine performs sentiment analysis on the text to identify the user's intentions and emotions.
[1281] The emotion engine performs an emotion analysis of the storyboard and identifies the emotional nuances of the depicted scene.
[1282] Step 5:
[1283] The server searches for and acquires the necessary materials based on the analysis and emotion analysis results, including searching for illustrations and audio files from a materials database.
[1284] Specific behavior:
[1285] The server executes an SQL query against the material database and searches for materials using search criteria based on the analysis results and emotion analysis results.
[1286] The path of the corresponding material file is obtained and temporarily saved.
[1287] Step 6:
[1288] The server generates payment information for the used material, including information about the material provider and a totaling process for the usage fee.
[1289] Specific behavior:
[1290] The server obtains information about the material provider and the usage fee from the material database.
[1291] Billing information is generated based on the acquired information and presented to the user.
[1292] Step 7:
[1293] The user confirms the payment information provided by the server and makes the payment, after which the material becomes officially available to the user.
[1294] Specific behavior:
[1295] The server generates billing information and sends it to the user in HTML format.
[1296] The user visits the payment page and enters the required payment information.
[1297] The terminal sends the payment information to the server and waits for the payment to be completed.
[1298] Step 8:
[1299] The server verifies that payment has been made and prepares the video for generation, thereby confirming permission to legally use the material.
[1300] Specific behavior:
[1301] Your server uses the payment gateway API to confirm the payment is successful.
[1302] Once payment is confirmed, the permission to use the materials will be updated in our internal system.
[1303] Step 9:
[1304] The server uses the analysis results and the acquired materials to send the data to an AI video generation service, which then automatically generates the video.
[1305] Specific behavior:
[1306] The server sends the necessary material files and analysis results to the API of the AI video generation service.
[1307] The AI video generation service generates a video based on the provided data and returns the generated video file to the server.
[1308] Based on the emotion engine, the tone of the video's narration and visual expression are adjusted.
[1309] Step 10:
[1310] The server sends a download link for the generated video to the user's device, allowing the user to access the video.
[1311] Specific behavior:
[1312] The server creates a URL for the generated video file and generates a link in an HTML email or on the dashboard.
[1313] When users click on the link, they can download or stream the video on their device.
[1314] Example 2
[1315] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1316] Conventional video generation systems require a lot of manual work when generating videos based on user-entered text and storyboards. Furthermore, they lacked a means to reflect the user's emotions in the video, making it difficult to generate personalized, high-quality videos. This increased the time and cost required for video production, making them unusable for many users.
[1317] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving text and storyboard data input by the user, means for analyzing the text and storyboard, means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results, means for generating payment information for use of the materials, means for confirming that the user has completed payment, means for automatically generating a video using the materials and the analysis results and emotion analysis results, and means for providing the generated video to the user. This enables the user to automatically generate high-quality videos that reflect emotions with little effort and quickly use them.
[1318] "User" refers to the person who inputs text and storyboards to use the system.
[1319] "Data receiving means" refers to the function of receiving text and storyboard data entered by the user.
[1320] "Data analysis means" refers to a function for analyzing received text and storyboards.
[1321] "Emotion analysis means" refers to a function that identifies the user's emotions contained in text or storyboards based on the analysis results.
[1322] "Material search means" refers to a function for searching for and acquiring necessary materials based on the analysis results and emotion analysis results.
[1323] "Material acquisition means" refers to a function for acquiring materials identified by the search means from a database or external service.
[1324] The "payment information generating means" refers to a function for generating payment information for the materials used.
[1325] "Payment Verification Method" refers to the function that verifies that a user has completed a payment.
[1326] "Video generation means" refers to a function that automatically generates videos based on acquired materials and analysis results.
[1327] "Video providing means" refers to the function of providing the generated video to the user.
[1328] The present invention relates to a system that automatically generates high-quality videos based on text and storyboard data entered by a user, and further enables personalized video creation by recognizing the user's emotions and reflecting them in the videos. This system includes the following components.
[1329] First, the user accesses a dedicated web page or application using their device and enters text and storyboards for the video they want to create. This input data is then sent from the device to the server via an HTTP request.
[1330] The server temporarily stores the received data and then begins processing to analyze it. Natural language processing (NLP) technology is used to analyze the text data, specifically using Python's NLTK and Spacy libraries. Image analysis technology is used to analyze the storyboards, specifically using the OpenCV library for contour detection and shape recognition.
[1331] Furthermore, we use sentiment analysis tools to identify the emotions contained in the analyzed text and storyboards. For sentiment analysis of text, we use natural language processing services such as Google Cloud NLP, and for sentiment analysis of storyboards, we use machine learning frameworks such as TensorFlow.
[1332] Based on the analysis results and sentiment analysis results, the necessary materials (e.g., illustrations and audio) are searched and acquired from a stock material database. The material database is connected to the API of an online stock material site, allowing necessary materials to be searched and acquired programmatically.
[1333] Next, the server generates payment information for the materials used. This information is generated by retrieving information about the material provider and the usage fee from a database and adding them up. The database used here is MySQL, a common relational database.
[1334] Once the user's payment is confirmed, the server generates a video using AI video generation techniques, adjusting the expressions in the video based on the results of sentiment analysis. The video is generated using OpenAI's CLIP model and DeepAI's API.
[1335] Finally, the generated video is served to the user: the server sends a download link for the generated video to the user's device, which the user can use to download or stream the video.
[1336] Specific examples
[1337] Example 1: Manual video for manufacturing industry
[1338] The user enters the "steps to assemble part A" in text on their device and uploads a simple sketch as a storyboard. The server receives this and performs natural language processing and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The emotion engine then analyzes the user's intentions and emotions from the entered text and generates a narration that reflects kindness or strictness. The server searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[1339] Example 2: Educational lesson videos
[1340] A user who is an educator enters a text description of the solar system and uploads a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio for the solar system. The emotion engine also performs emotional analysis of the entered text and selects passionate narration or a calm tone. The server searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[1341] Prompt Sentence Examples
[1342] Example 1: Manual video for manufacturing industry
[1343] "Please upload a written description of the steps to assemble part A and a simple sketch as a storyboard. After analysis, we will identify the illustrations of the parts and the audio material of the assembly steps and generate a narration that reflects kindness or strictness."
[1344] Example 2: Educational lesson videos
[1345] "Enter a description of the solar system in text and upload a simple storyboard for each planet. After analysis, we'll generate a video with illustrations of the solar system and a choice of passionate or calming narration."
[1346] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1347] Step 1: Data entry
[1348] The user uses a terminal to access a dedicated web page or application and inputs text and storyboards related to the video they want to generate. Specifically, the user inputs the text "Assembly procedure for part A" and uploads the storyboard for part A. The input format is a text file and an image file. This results in the text and storyboards as input data.
[1349] Step 2: Send data
[1350] Data entered by the user is sent from the terminal to the server via an HTTP request. For example, a POST request is used to send text data and image data to the server's endpoint. The data is received by the server and temporarily stored in a database or system memory.
[1351] Step 3: Text analysis
[1352] The server analyzes the received text data using natural language processing (NLP) techniques, specifically Python's NLTK and Spacy libraries. For example, it tokenizes the text, tags parts of speech, and performs syntactic analysis. This process extracts syntactic and semantic data from the input text.
[1353] Step 4: Storyboard Analysis
[1354] The server analyzes the storyboard it receives using image analysis technology. The OpenCV library is used for image analysis, specifically contour detection and shape recognition. For example, it can detect a specific shape (e.g., the shape of part A) from the uploaded storyboard and extract the corresponding data.
[1355] Step 5: Sentiment Analysis
[1356] The server uses an emotion engine to identify emotions contained in the text and storyboards. It uses machine learning frameworks such as Google Cloud NLP for emotion analysis of text and TensorFlow for emotion analysis of storyboards. Specifically, it uses an emotion model to identify emotions such as "joy," "sadness," and "anger."
[1357] Step 6: Material Search
[1358] Based on the analysis and emotion analysis results, the server searches for the necessary materials from a material database. The material database is connected to an online stock material site using an API. For example, it searches for "illustration of part A" or "gentle-toned voice narration" to identify appropriate materials.
[1359] Step 7: Obtaining Materials
[1360] The server retrieves the searched material. This involves downloading data from the stock material site via API. The retrieved material is then stored in a directory or database managed by the server.
[1361] Step 8: Generate payment information
[1362] The server generates payment information for the materials used. The payment information is generated by retrieving information about the material provider and the usage fee from a database and adding them up. Specifically, a specific SQL query is executed using a MySQL database to calculate the usage fee.
[1363] Step 9: Confirm payment
[1364] The server presents the generated billing information to the user. The user enters the payment information (e.g., credit card information). The server processes the payment using an online payment system (e.g., PayPal or Stripe). If the payment is successful, the server confirms it.
[1365] Step 10: Video Generation
[1366] After the user's payment is confirmed, the server uses the analysis results and the acquired materials to execute the AI video generation service. The emotion analysis results are taken into account and the expressions in the video are adjusted. The video is generated using OpenAI's CLIP model and DeepAI's API. The generated video file is saved in the appropriate format.
[1367] Step 11: Submit your video
[1368] The server provides the generated video to the user by generating a download link for the video and sending it to the user's device. The user can then use the link to download or stream the video.
[1369] (Application example 2)
[1370] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1371] Conventional video generation systems simply generate videos based on user-entered text and storyboards, making it difficult to generate personalized videos that reflect the user's emotions. Furthermore, there is a lack of systems that efficiently execute the entire process of delivering generated videos to target users. For this reason, the advertising industry needs a way to quickly generate and deliver effective video ads that match the interests and emotions of target users.
[1372] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1373] In this invention, the server includes means for receiving text and storyboard data input by the user, means for analyzing the text and storyboard, and means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results. This makes it possible to generate high-quality videos that reflect emotions based on the user's input data and distribute them to target users.
[1374] "Means for receiving text and storyboard data entered by the user" refers to a function that allows a user to use a dedicated application or web page via a terminal to enter text and storyboard data and send it to the server.
[1375] The "means for analyzing text and storyboards" refers to a means for analyzing received text and storyboard data using natural language processing and image analysis techniques to understand the content.
[1376] "Means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results" refers to means for searching for and acquiring necessary video and audio materials from a materials database based on the analyzed data and the results of emotion analysis.
[1377] The "means for generating payment information for use of material" is a function for calculating the fee for the used material and generating corresponding payment information.
[1378] The "means for confirming that the user has completed the payment" refers to a means for confirming that the user has made the payment based on the generated payment information.
[1379] "Means for automatically generating videos using materials, analysis results, and emotion analysis results" refers to means for automatically generating videos using AI video generation technology based on the acquired materials and analysis results.
[1380] "Means for providing the generated video to the user" means means for providing the generated video to the user and enabling the user to download or stream the video.
[1381] "Means for distributing the generated video to target users" refers to means for distributing the generated video to target users designated by the advertiser, thereby increasing the effectiveness of the advertisement.
[1382] The embodiments for carrying out the present invention are described in detail below.
[1383] The system aims to enable advertisers to generate personalized video ads using smartphone applications and efficiently deliver them to target users.
[1384] 1. User Input and Data Receipt
[1385] Advertisers use a dedicated smartphone application to input text and storyboards for their advertisements, which are then sent to the server via HTTP requests.
[1386] 2. Data Analysis
[1387] The server analyzes the received text and storyboard data using a natural language processing library (e.g., SpaCy) and an image analysis library (e.g., OpenCV). The analysis includes syntactic analysis of the text data and feature extraction of the image data.
[1388] 3. Emotion recognition
[1389] The emotion engine analyzes emotions based on the analyzed data. A common emotion analysis API (e.g., IBM Watson Natural Language Understanding, Google Cloud Natural Language API) is used as the emotion engine.
[1390] 4. Material Search and Acquisition
[1391] Based on the results of sentiment analysis and data analysis, the server searches for and obtains the necessary video and audio materials from a material database (e.g., Shutterstock API).
[1392] 5. Payment Information Generation and Verification
[1393] The server generates payment information based on the information of the materials used and presents it to the advertiser. Payment processing is performed using an electronic payment service (e.g., Stripe API). Once the advertiser completes payment, the information is recorded on the server.
[1394] 6. Video Generation
[1395] Once payment is confirmed, the server uses the acquired materials and analysis results to automatically generate a video using an AI video generation service (e.g., Synthesia). The generated video reflects narration and direction based on the results of the emotion analysis.
[1396] 7. Video Provision and Distribution
[1397] The server provides the generated video advertisement to the advertiser, who then sends a link to the advertisement via a smartphone application, allowing the advertiser to distribute the video to target users.
[1398] Specific examples
[1399] Example 1: Advertising video for new product "Express Coffee"
[1400] Advertisers enter text and sketch images into the application that explain the features of their new product, "Express Coffee."
[1401] Example sentence: "Express Coffee provides fast, delicious coffee for busy mornings."
[1402] Storyboard: "Image of an Express Coffee package and coffee being poured into a cup"
[1403] Emotion: We want to convey a sense of comfort and trust to the user.
[1404] The server analyzes the input data and searches for and retrieves appropriate content from a content database based on the results of sentiment analysis. Once payment is completed, an AI video generation service is used to generate a video ad that reflects the specified sentiment. The generated video ad link is provided to the advertiser through the application, and the advertiser uses this link to deliver the video to target users.
[1405] This system enables advertisers to quickly generate high-quality, emotionally relevant video ads and deliver them effectively to target users.
[1406] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1407] Step 1:
[1408] Input and Data Reception
[1409] Users use a smartphone application to input text and storyboards to be used in advertising videos. The input data is collected through a form in the application and sent to the server as an HTTP request. The server temporarily stores the received data.
[1410] Specific operation: The user uploads text data (e.g., advertising copy) and image data (e.g., product sketch) into the application's input form and sends it to a dedicated API endpoint.
[1411] Step 2:
[1412] Data analysis
[1413] The server analyzes the received text and storyboard data. Specifically, it analyzes the text using a natural language processing library (e.g., SpaCy) and the storyboard using an image analysis library (e.g., OpenCV). Text analysis is used to understand the meaning and structure of the text, and image analysis is used to extract the features of the storyboard.
[1414] Input and Output: The input data are the received text and storyboard, and the output is the analyzed text data and image data features.
[1415] Specific operation: Using the SpaCy library, tokenize text data, assign POS tags, and analyze dependencies. Using the OpenCV library, detect edges and extract feature points from storyboard images.
[1416] Step 3:
[1417] emotion recognition
[1418] The server uses an emotion engine based on the analyzed data to analyze emotions from the text and storyboard. It uses an emotion analysis API (e.g., IBM Watson Natural Language Understanding, Google Cloud Natural Language API) to identify the user's emotions.
[1419] Input and Output: The input is the analyzed text data and image data features, and the output is the sentiment analysis results.
[1420] Specific operation: The analyzed text data and image data features are sent to the emotion analysis API, and emotions are categorized (e.g., joy, sadness, surprise, etc.) and received.
[1421] Step 4:
[1422] Material Search and Acquisition
[1423] Based on the results of emotion analysis and data analysis, the server searches and retrieves appropriate video and audio materials from a material database (e.g., Shutterstock API).
[1424] Input and output: The input is the emotion analysis results and data analysis results, and the output is the acquired video and audio materials.
[1425] Specific operation: Creates a search query for the Shutterstock API based on the results of sentiment analysis and data analysis, and retrieves the necessary materials (images, audio, video).
[1426] Step 5:
[1427] Payment information generation and verification
[1428] The server generates payment information based on the information about the used materials and presents it to the user. It processes the payment using an electronic payment service (e.g., Stripe API) and confirms that the payment has been completed.
[1429] Input and Output: The input is the information of the material acquired and the output is the status of payment confirmation.
[1430] Specific operation: Calculate the total amount based on the price information of the obtained materials, send a payment request to the user through the Stripe API, and check the status of whether the payment has been completed.
[1431] Step 6:
[1432] Video Generation
[1433] Once payment is confirmed, the server uses the acquired materials and analysis results to automatically generate a video using an AI video generation service (e.g., Synthesia), including narration and direction based on the results of the emotion analysis.
[1434] Input and Output: The input is the acquired material and analysis results, and the output is the generated video file.
[1435] Specific operation: The acquired material, analysis results, and a prompt for a video to be generated based on the emotion analysis results are sent to the Synthesia API, and the generated video is retrieved.
[1436] Step 7:
[1437] Video provision and distribution
[1438] The generated video advertisement is provided to the user (advertiser) from the server, and a link is sent to the user via a smartphone application. The user can then distribute the video to target users.
[1439] Input and Output: The input is the generated video file and the output is the video link provided to the user and the delivery status to the target user.
[1440] Specific operation: The generated video is stored in the server storage, and a download link is generated and sent to the user. The user receives the link and the video is distributed to the target user.
[1441] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1442] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1443] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1444] [Fourth embodiment]
[1445] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1446] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1447] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1448] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1449] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1451] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1452] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1453] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1454] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1455] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1456] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1457] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1458] The present invention relates to a system that automatically generates high-quality animation based on text and storyboards input by a user. This system has the following main functions:
[1459] 1. Data reception function
[1460] The user uses a terminal to access a dedicated web page or application, input text and storyboards, and this data is sent to the server via an HTTP request.
[1461] 2. Data analysis function
[1462] The server analyzes the received text and storyboard data. Specifically, it uses natural language processing technology to analyze the text and image analysis technology to analyze the storyboard. This allows it to identify the materials needed to generate the video.
[1463] 3. Material search and acquisition function
[1464] The server searches and retrieves the necessary materials, such as illustrations and audio, from a materials database based on the analysis results. This process involves searching the materials database using SQL queries.
[1465] 4. Payment information generation function
[1466] The server generates payment information for the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates total billing information.
[1467] 5. Payment confirmation function
[1468] The generated billing information is presented to the user for payment confirmation, and the payment is processed based on the payment information provided by the user. The server then confirms that the payment has been completed.
[1469] 6. Video generation function
[1470] Once payment is complete, the server uses the analysis results and acquired materials to run the AI video generation service and generate a video. The AI video generation service automatically generates a video according to the provided materials and instructions.
[1471] 7. Video provision function
[1472] Finally, the server provides the generated video to the user by sending a download link for the generated video file to the user's device, which the user can use to download or stream the video.
[1473] Specific examples
[1474] Example 1: Manual video for manufacturing industry
[1475] A user in the manufacturing industry enters the steps for assembling part A in text and uploads a simple sketch as a storyboard. The server receives this and performs text and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The server then searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[1476] Example 2: Educational lesson videos
[1477] Educators enter a text description of the solar system and upload a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio of the solar system. It then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[1478] This system will enable users to easily generate and use high-quality videos, significantly reducing the time and cost required for video production and is expected to be used in a wide range of applications.
[1479] The processing flow will be explained below.
[1480] Step 1:
[1481] The user uses the device to access a dedicated web page or application, enter text, and upload storyboards, thereby providing information about the content of the video.
[1482] Specific behavior:
[1483] The user logs in to a dedicated web page or application.
[1484] The user enters a description of the video in a text box.
[1485] The user opens a file selection dialog and selects an image file to upload a storyboard.
[1486] Once you have completed the entry and upload, click the "Submit" button.
[1487] Step 2:
[1488] The device sends the input text and storyboard data to the server as an HTTP request.
[1489] Specific behavior:
[1490] The device encodes the input text and storyboard as JSON format or multipart form data.
[1491] The encoded data is sent as an HTTP POST request to the server's API endpoint.
[1492] Step 3:
[1493] The server analyzes the received text and storyboard data, specifically using natural language processing technology for the text and image analysis technology for the storyboard.
[1494] Specific behavior:
[1495] The server receives the HTTP request and passes the data to the analysis module.
[1496] A natural language processing engine analyzes the text data and identifies the necessary keywords and context.
[1497] An image analysis engine processes the storyboard and extracts specific elements and scenes.
[1498] Step 4:
[1499] The server searches for and acquires the necessary materials based on the analysis results, including searching for illustrations and audio files from a materials database.
[1500] Specific behavior:
[1501] The server executes an SQL query against the material database and searches for materials using search criteria based on the analysis results.
[1502] The path of the corresponding material file is obtained and temporarily saved.
[1503] Step 5:
[1504] The server generates payment information for the used material, including information about the material provider and a totaling process for the usage fee.
[1505] Specific behavior:
[1506] The server obtains information about the material provider and the usage fee from the material database.
[1507] Billing information is generated based on the acquired information and presented to the user.
[1508] Step 6:
[1509] The user confirms the payment information provided by the server and makes the payment, after which the material becomes officially available to the user.
[1510] Specific behavior:
[1511] The server generates billing information and sends it to the user in HTML format.
[1512] The user visits the payment page and enters the required payment information.
[1513] The terminal sends the payment information to the server and waits for the payment to be completed.
[1514] Step 7:
[1515] The server verifies that payment has been made and prepares the video for generation, thereby confirming permission to legally use the material.
[1516] Specific behavior:
[1517] Your server uses the payment gateway API to confirm the payment is successful.
[1518] Once payment is confirmed, the permission to use the materials will be updated in our internal system.
[1519] Step 8:
[1520] The server uses the analysis results and the acquired materials to send the data to an AI video generation service, which then automatically generates the video.
[1521] Specific behavior:
[1522] The server sends the necessary material files and analysis results to the API of the AI video generation service.
[1523] The AI video generation service generates a video based on the provided data and returns the generated video file to the server.
[1524] Step 9:
[1525] The server sends a download link for the generated video to the user's device, allowing the user to access the video.
[1526] Specific behavior:
[1527] The server creates a URL for the generated video file and generates a link in an HTML email or on the dashboard.
[1528] When users click on the link, they can download or stream the video on their device.
[1529] Example 1
[1530] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1531] Demand for high-quality video content is increasing in many fields today. However, video production requires specialized skills and is time-consuming and costly. Furthermore, there are not enough methods available for users to easily create high-quality videos based on their own ideas and information. Therefore, there is a need for a system that can efficiently create and provide high-quality videos without specialized knowledge.
[1532] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1533] In this invention, the server includes means for receiving information and visual material data input by a user, means for analyzing the information and visual material, and means for searching for and acquiring necessary data based on the analysis results, thereby enabling users without specialized knowledge to easily turn their ideas into high-quality videos.
[1534] "User" refers to an individual or organization that utilizes this system to provide information and visual materials and request the creation of a video.
[1535] "Visual materials" refers to materials that convey information visually, such as storyboards and image files that users upload to the system.
[1536] "Data receiving means" refers to a mechanism by which the server receives information and visual materials entered by the user via the network.
[1537] "Analysis means" refers to the process of breaking down and interpreting received information and visual materials, and extracting and identifying the necessary data.
[1538] "Search and acquisition means" refers to the function of searching the database for the necessary data based on the analysis results and acquiring the appropriate materials.
[1539] The "fee information generating means" is a mechanism for calculating the fee for the user based on the data and materials used and creating billing information.
[1540] "Payment confirmation means" refers to a process for confirming that the user has completed payment.
[1541] "Video generation means" refers to a function that automatically generates high-quality video using AI technology, etc., based on the analysis results and acquired materials.
[1542] The "video providing means" refers to a mechanism for providing the generated video in a form that allows users to download or stream the video.
[1543] The present invention relates to a system for automatically generating and providing high-quality video based on user-provided information and visual materials. The system includes the following main components:
[1544] Data reception
[1545] The user uses a device to access a dedicated web page or application. The user inputs and uploads text information and visual materials such as storyboards for the video they want to create. The device then sends this data to the server via an HTTP POST request. The server then receives the data provided by the user.
[1546] Data analysis
[1547] The server analyzes the received text and visual materials. For text data, it uses a natural language processing engine (e.g., SpaCy, BERT) to analyze the text and extract keywords and important location information. For image data, it uses image analysis technology (e.g., OpenCV, TensorFlow) to analyze the storyboard and identify objects. This identifies the materials needed to generate the video.
[1548] Material Search and Acquisition
[1549] Based on the analysis results, the server searches for the necessary data (illustrations, audio files, etc.) from the material database and retrieves them using SQL queries. These materials are used as components necessary for video generation.
[1550] Payment information generation
[1551] The server generates fee information for the use of the material. Specifically, it retrieves information about the material provider and the usage fee from the database, adds them up, and generates an invoice to be presented to the user.
[1552] Payment confirmation
[1553] Provide the user with a link to the payment page and confirm that they wish to complete the payment. The server receives notification that the payment has been completed and proceeds to the next processing step.
[1554] Image Generation
[1555] The server generates videos using AI video generation services (e.g., OpenAI's DALL-E, GPT-4) based on the analysis results and acquired materials. The video generation process is fully automated, and high-quality videos are generated according to the instructions provided by the user.
[1556] Video provided by
[1557] Finally, the server provides the user with a download link for the generated video, which the user can use to download or stream the video, allowing the user to obtain high-quality video in a convenient way.
[1558] Specific examples
[1559] Example 1: Manual video for manufacturing industry
[1560] The user inputs the assembly instructions and associated storyboard for a specific product from their device. The server analyzes the instructions, identifies and acquires the necessary illustrations and audio materials, and generates payment information. Once the user completes payment, the server generates a video using an AI video generation service and provides the user with a download link for the final video.
[1561] prompt:
[1562] "Assembly steps for part A:
[1563] 1. Take out part A.
[1564] 2. Connect to part B.
[1565] 3. Tighten the screws.
[1566] "
[1567] Image: [Illustration: Part A, Part B, Screw diagram]
[1568] Example 2: Educational lesson videos
[1569] The user, an educator, enters a text description of the solar system and uploads a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio of the solar system. It then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user.
[1570] prompt:
[1571] "Solar System Description:
[1572] 1. The sun is at the center and the other planets revolve around it.
[1573] 2. Mercury is the planet closest to the sun.
[1574] 3. Venus is the brightest star known.
[1575] "
[1576] Image: [Diagram showing the solar system, with diagrams of each planet]
[1577] This allows users to easily create high-quality videos without specialized knowledge or skills, and provides a system that can be used for a variety of purposes.
[1578] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1579] Step 1: Receiving data
[1580] Users use their devices to access a dedicated web page or application and input and upload written information about the video they want to create and visual materials such as storyboards.
[1581] Input: Text information and storyboard data entered by the user
[1582] The terminal sends this data to the server as an HTTP POST request.
[1583] Output: Text information and storyboard data received by the server
[1584] Step 2: Data analysis
[1585] The server analyzes the received text information using a natural language processing engine (e.g., SpaCy, BERT) to extract keywords and important location information.
[1586] Input: Received text information
[1587] Output: Extracted keywords and location information
[1588] The server analyzes the storyboard using image analysis technology (e.g., OpenCV, TensorFlow) and identifies the objects.
[1589] Input: Received storyboard data
[1590] Output: Identified objects
[1591] Step 3: Search and acquire materials
[1592] Based on the analysis results, the server searches for the necessary data (illustrations, audio files, etc.) from the material database.
[1593] Input: Analysis results (extracted keywords and identified objects)
[1594] Output: Searched material data
[1595] The server uses an SQL query to search the materials database and retrieve the appropriate materials.
[1596] Input: SQL query
[1597] Output: Acquired material data
[1598] Step 4: Generate payment information
[1599] The server generates fee information based on the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates a combined invoice.
[1600] Input: Acquired material data
[1601] Output: Generated invoice
[1602] Step 5: Payment confirmation
[1603] The server presents the generated invoice to the user and provides a link to a payment page.
[1604] Input: Generated Invoice
[1605] Output: Payment page link
[1606] The user enters payment information and completes the payment.
[1607] Input: User's payment information
[1608] Once the server receives notification that payment has been completed, it proceeds to the next processing step.
[1609] Output: Confirmation of successful payment
[1610] Step 6: Image generation
[1611] The server generates video using an AI video generation service (e.g., OpenAI's DALL-E, GPT-4) based on the analysis results and acquired materials.
[1612] Input: Analysis results and acquired materials
[1613] Output: Generated video file
[1614] The server stores the generated video files.
[1615] Input: Generated video file
[1616] Output: Saved video file
[1617] Step 7: Provide footage
[1618] The server provides the user with a download link for the generated video.
[1619] Input: Saved video file
[1620] Output: Video download link
[1621] Users can use this link to download or stream the video.
[1622] Input: Video download link
[1623] Output: Downloaded or streamed video
[1624] Through these processing steps, users can easily create high-quality videos for a variety of uses, even without specialized knowledge or skills.
[1625] (Application example 1)
[1626] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1627] Conventional video production systems require a lot of time and effort for the entire process of collecting, editing, and creating the materials needed to create a video, making it difficult for average users to easily create high-quality videos. Furthermore, the payment process for using the materials is complicated, placing a burden on users. The present invention aims to solve these problems by providing a system that allows users to easily create high-quality videos and complete payments smoothly.
[1628] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1629] In this invention, the server includes means for receiving text and storyboard data entered by a user, means for analyzing the text and storyboard, means for searching for and acquiring necessary materials based on the analysis results, means for generating payment information for the use of the materials, means for confirming that the user has completed payment, means for automatically generating a video using the materials and the analysis results, means for providing the generated video to the user, means for providing a user interface integrated into a smartphone application, and means for using AI to generate a video based on materials acquired from a materials database and confirm payment. This allows users to generate high-quality videos in a short amount of time and complete payment procedures easily and quickly.
[1630] "User-input text" refers to data based on text provided by a user through the system's input interface.
[1631] A "storyboard" is drawing data based on a visual guide entered by the user.
[1632] A "receiving means" is a part of the system that has the function of capturing data sent by a user.
[1633] The "analyzing means" is a part of the system that has the function of analyzing the received data and extracting the necessary information.
[1634] "Necessary materials" refers to data such as images and audio required to generate video.
[1635] A "search and retrieval means" is a part of a system that has the ability to search for information in a database and retrieve requested material.
[1636] A "means for generating payment information" is a part of the system that has the functionality to calculate fees for use of material and generate that information.
[1637] A "means for confirming payment completion" is a part of the system that has the function of confirming that a user's payment has been successfully made.
[1638] The "means for automatically generating a video" is a part of a system that has the function of automatically creating a video using the acquired materials and analysis results.
[1639] A "means for providing videos to users" is a part of the system that has the function of making the generated videos available to users.
[1640] A "smartphone application" is a program that runs on a smartphone and provides an interface that can be operated directly by the user.
[1641] A "user interface" is a feature that serves as an entry point for a user to interact with a system.
[1642] A "material database" is a data storage that stores materials necessary for video generation.
[1643] "AI-based means" refers to a part of a system that has the function of generating videos using artificial intelligence technology.
[1644] The embodiment of the present invention is a system that automatically generates high-quality videos based on text and storyboards entered by the user. This system is composed of a server and a user's terminal (smartphone application).
[1645] System Program
[1646] The system includes the following key features:
[1647] 1. Data reception function: The server receives text and storyboards sent from the user's device. The user inputs and sends this data via a smartphone application.
[1648] 2. Data analysis function: The server analyzes the received text using natural language processing technology (e.g., spaCy or NLTK) and analyzes the storyboard using image analysis technology (e.g., OpenCV or TensorFlow).
[1649] 3. Material search and retrieval function: Based on the analysis results, the server searches and retrieves related images, audio, and other materials from the material database. Here, the database search is performed using SQL queries.
[1650] 4. Payment information generation function: The server calculates the usage fee for the acquired material and presents payment information to the user. At this time, information about the material provider is also obtained.
[1651] 5. Payment confirmation function: The server processes the payment based on the payment information provided by the user and confirms that the payment has been completed. Possible payment services used include Stripe and PayPal.
[1652] 6. Video generation function: The server automatically generates videos using AI based on the acquired materials and analysis results. AI video generation uses services such as DeepArt and Runway ML.
[1653] 7. Video provision function: The generated video is provided to the user from the server. The user can stream or download the video via the URL of the generated video.
[1654] Natural language description of the process
[1655] Data reception and analysis: The text and storyboards entered by the user using the smartphone application are sent to the server via HTTP requests. The server stores the received data, analyzes the text using natural language processing (NLP) technology, and analyzes the storyboards using image analysis technology.
[1656] Material search and retrieval: Based on the analysis results, the server performs a database search to retrieve the required image and audio materials. The material database stores materials by multiple categories, so the required materials can be quickly retrieved using the appropriate SQL query.
[1657] Payment processing: Calculate the usage fee for the acquired materials and present the payment information to the user. The server confirms that the user has completed the payment and proceeds to the next step.
[1658] Video generation and provision: Using AI technology, a video is generated by combining the necessary materials and analysis results. The generated video is provided to the user as a URL link. The user can view or download the generated video via the link.
[1659] Examples of concrete examples and prompts
[1660] Examples:
[1661] In the education field, users can enter a text description of each planet in the solar system and upload a corresponding hand-drawn sketch of the planet as a storyboard. This data is then analyzed to retrieve images of the corresponding planet from, for example, a NASA image database, and audio material from LibriVox. The resulting video can then be used in the user's educational activities.
[1662] Example prompt sentence:
[1663] "Generate a video about each planet in the solar system. Create a high-quality video based on the text and storyboard below.
[1664] Text: The solar system has the sun at its center and the planets that orbit it are Mercury, Venus, Earth, Mars, Jupiter, Saturn, Uranus, and Neptune.
[1665] Storyboard: Hand-drawn planet sketch (image data)
[1666] The above is a specific embodiment of the present invention, which allows users to easily generate high-quality videos and use them immediately.
[1667] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1668] Step 1:
[1669] The user inputs text and storyboards using a smartphone application and sends them.
[1670] Input: Text data, storyboard data
[1671] Output: HTTP request data
[1672] Specific operation: The user enters text into the application's input form and uploads image files that will serve as storyboards. This data is sent from the device to the server as an HTTP request.
[1673] Step 2:
[1674] The server receives the HTTP request and stores the data.
[1675] Input: HTTP request data
[1676] Output: Saved text data, saved storyboard data
[1677] Specific operation: The server stores the received data in temporary storage and prepares it for the next analysis step.
[1678] Step 3:
[1679] The server analyzes the received data.
[1680] Input: Saved text data, saved storyboard data
[1681] Output: Analyzed text data, analyzed storyboard data
[1682] Specific operation: The server analyzes the text using natural language processing technology (e.g., spaCy or NLTK) to extract keywords and structural information, and also analyzes the storyboard using image analysis technology (e.g., OpenCV or TensorFlow) to extract the objects and layout information contained therein.
[1683] Step 4:
[1684] The server searches the material database based on the analysis results and retrieves the necessary images and audio materials.
[1685] Input: Analyzed text data, analyzed storyboard data
[1686] Output: Acquired image material data, acquired audio material data
[1687] Specific operation: The server uses an SQL query to search the material database based on the analysis results and retrieve relevant images and audio materials.
[1688] Step 5:
[1689] The server calculates the usage fee for the acquired material and generates payment information.
[1690] Input: Acquired image material data, acquired audio material data
[1691] Output: Payment information data
[1692] Specific operation: The server retrieves the information of the material provider and the usage fee data, combines them, and generates payment information, which is then presented to the user.
[1693] Step 6:
[1694] The server processes the payment based on the user's payment information and confirms that the payment has been completed.
[1695] Input: Payment information data, user's payment information
[1696] Output: Payment confirmation data
[1697] Specific operation: The server uses a payment service such as Stripe or PayPal to check whether the user's payment has been successfully completed. If the confirmation is successful, it proceeds to the next step.
[1698] Step 7:
[1699] Using the materials acquired by the server and the analysis results, videos are automatically generated using AI.
[1700] Input: Acquired image material data, acquired audio material data, analyzed text data, analyzed storyboard data
[1701] Output: Generated video data
[1702] Specific operation: The server uses an AI video generation service (e.g., DeepArt or Runway ML) to generate a video that combines the acquired materials and analysis results.
[1703] Step 8:
[1704] The server provides the generated video to the user.
[1705] Input: Generated video data
[1706] Output: Video data provided to users (URL link, etc.)
[1707] Specific operation: The server saves the generated video in cloud storage and provides the URL to the user, who can use this URL to stream or download the video.
[1708] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1709] This invention relates to a system that automatically generates high-quality videos based on text and storyboards entered by the user, as well as a system that recognizes the user's emotions and reflects them in the videos. This system has the following main functions:
[1710] 1. Data reception function
[1711] The user uses a terminal to access a dedicated web page or application, input text and storyboards, and this data is sent to the server via an HTTP request.
[1712] 2. Data analysis function
[1713] The server analyzes the received text and storyboard data using natural language processing and image analysis technologies.
[1714] 3. Emotion engine function
[1715] The server is equipped with an emotion engine for recognizing emotions contained in text and storyboard data. The emotion engine performs emotion analysis when analyzing text, and also performs emotion analysis on storyboards.
[1716] 4. Material search and acquisition function
[1717] The server searches and retrieves the necessary materials from a materials database based on the analysis and emotion analysis results, including appropriate illustrations, audio, and other materials.
[1718] 5. Payment information generation function
[1719] The server generates payment information for the materials used. Specifically, it retrieves information about the material provider and the usage fee from a database and generates total billing information.
[1720] 6. Payment confirmation function
[1721] The generated billing information is presented to the user for payment confirmation, and the payment is processed based on the payment information provided by the user. The server then confirms that the payment has been completed.
[1722] 7. Video generation function
[1723] Once payment is complete, the server uses the analysis results and acquired materials to run the AI video generation service and generate a video, including a function to adjust the expressions in the video based on the emotion engine.
[1724] 8. Video provision function
[1725] Finally, the server provides the generated video to the user by sending a download link for the generated video file to the user's device, which the user can use to download or stream the video.
[1726] Specific examples
[1727] Example 1: Manual video for manufacturing industry
[1728] A user working in the manufacturing industry enters the steps for assembling part A in text and uploads a simple sketch as a storyboard. The server receives this and performs natural language processing and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The emotion engine then analyzes the user's intentions and emotions from the input text and generates a narration that reflects kindness or strictness. The server searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[1729] Example 2: Educational lesson videos
[1730] Educators enter a text description of the solar system and upload a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio for the solar system. The emotion engine then performs an emotional analysis of the entered text and selects passionate narration or a calm tone. The system then searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[1731] This system enables users to easily generate and use high-quality, emotionally-reflective videos, significantly reducing the time and cost required for video production and is expected to be used in a wide range of applications.
[1732] The processing flow will be explained below.
[1733] Step 1:
[1734] The user uses the device to access a dedicated web page or application, enter text, and upload storyboards, thereby providing information about the content of the video.
[1735] Specific behavior:
[1736] The user logs in to a dedicated web page or application.
[1737] The user enters a description of the video in a text box.
[1738] The user opens a file selection dialog and selects an image file to upload a storyboard.
[1739] Once you have completed the entry and upload, click the "Submit" button.
[1740] Step 2:
[1741] The device sends the input text and storyboard data to the server as an HTTP request.
[1742] Specific behavior:
[1743] The device encodes the input text and storyboard as JSON format or multipart form data.
[1744] The encoded data is sent as an HTTP POST request to the server's API endpoint.
[1745] Step 3:
[1746] The server analyzes the received text and storyboard data, specifically using natural language processing technology for the text and image analysis technology for the storyboard.
[1747] Specific behavior:
[1748] The server receives the HTTP request and passes the data to the analysis module.
[1749] A natural language processing engine analyzes the text data and identifies the necessary keywords and context.
[1750] An image analysis engine processes the storyboard and extracts specific elements and scenes.
[1751] Step 4:
[1752] The server uses an emotion engine to recognize emotions contained in the received text and storyboard data.
[1753] Specific behavior:
[1754] The emotion engine performs sentiment analysis on the text to identify the user's intentions and emotions.
[1755] The emotion engine performs an emotion analysis of the storyboard and identifies the emotional nuances of the depicted scene.
[1756] Step 5:
[1757] The server searches for and acquires the necessary materials based on the analysis and emotion analysis results, including searching for illustrations and audio files from a materials database.
[1758] Specific behavior:
[1759] The server executes an SQL query against the material database and searches for materials using search criteria based on the analysis results and emotion analysis results.
[1760] The path of the corresponding material file is obtained and temporarily saved.
[1761] Step 6:
[1762] The server generates payment information for the used material, including information about the material provider and a totaling process for the usage fee.
[1763] Specific behavior:
[1764] The server obtains information about the material provider and the usage fee from the material database.
[1765] Billing information is generated based on the acquired information and presented to the user.
[1766] Step 7:
[1767] The user confirms the payment information provided by the server and makes the payment, after which the material becomes officially available to the user.
[1768] Specific behavior:
[1769] The server generates billing information and sends it to the user in HTML format.
[1770] The user visits the payment page and enters the required payment information.
[1771] The terminal sends the payment information to the server and waits for the payment to be completed.
[1772] Step 8:
[1773] The server verifies that payment has been made and prepares the video for generation, thereby confirming permission to legally use the material.
[1774] Specific behavior:
[1775] Your server uses the payment gateway API to confirm the payment is successful.
[1776] Once payment is confirmed, the permission to use the materials will be updated in our internal system.
[1777] Step 9:
[1778] The server uses the analysis results and the acquired materials to send the data to an AI video generation service, which then automatically generates the video.
[1779] Specific behavior:
[1780] The server sends the necessary material files and analysis results to the API of the AI video generation service.
[1781] The AI video generation service generates a video based on the provided data and returns the generated video file to the server.
[1782] Based on the emotion engine, the tone of the video's narration and visual expression are adjusted.
[1783] Step 10:
[1784] The server sends a download link for the generated video to the user's device, allowing the user to access the video.
[1785] Specific behavior:
[1786] The server creates a URL for the generated video file and generates a link in an HTML email or on the dashboard.
[1787] When users click on the link, they can download or stream the video on their device.
[1788] Example 2
[1789] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1790] Conventional video generation systems require a lot of manual work when generating videos based on user-entered text and storyboards. Furthermore, they lacked a means to reflect the user's emotions in the video, making it difficult to generate personalized, high-quality videos. This increased the time and cost required for video production, making them unusable for many users.
[1791] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving text and storyboard data input by the user, means for analyzing the text and storyboard, means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results, means for generating payment information for use of the materials, means for confirming that the user has completed payment, means for automatically generating a video using the materials and the analysis results and emotion analysis results, and means for providing the generated video to the user. This enables the user to automatically generate high-quality videos that reflect emotions with little effort and quickly use them.
[1792] "User" refers to the person who inputs text and storyboards to use the system.
[1793] "Data receiving means" refers to the function of receiving text and storyboard data entered by the user.
[1794] "Data analysis means" refers to a function for analyzing received text and storyboards.
[1795] "Emotion analysis means" refers to a function that identifies the user's emotions contained in text or storyboards based on the analysis results.
[1796] "Material search means" refers to a function for searching for and acquiring necessary materials based on the analysis results and emotion analysis results.
[1797] "Material acquisition means" refers to a function for acquiring materials identified by the search means from a database or external service.
[1798] The "payment information generating means" refers to a function for generating payment information for the materials used.
[1799] "Payment Verification Method" refers to the function that verifies that a user has completed a payment.
[1800] "Video generation means" refers to a function that automatically generates videos based on acquired materials and analysis results.
[1801] "Video providing means" refers to the function of providing the generated video to the user.
[1802] The present invention relates to a system that automatically generates high-quality videos based on text and storyboard data entered by a user, and further enables personalized video creation by recognizing the user's emotions and reflecting them in the videos. This system includes the following components.
[1803] First, the user accesses a dedicated web page or application using their device and enters text and storyboards for the video they want to create. This input data is then sent from the device to the server via an HTTP request.
[1804] The server temporarily stores the received data and then begins processing to analyze it. Natural language processing (NLP) technology is used to analyze the text data, specifically using Python's NLTK and Spacy libraries. Image analysis technology is used to analyze the storyboards, specifically using the OpenCV library for contour detection and shape recognition.
[1805] Furthermore, we use sentiment analysis tools to identify the emotions contained in the analyzed text and storyboards. For sentiment analysis of text, we use natural language processing services such as Google Cloud NLP, and for sentiment analysis of storyboards, we use machine learning frameworks such as TensorFlow.
[1806] Based on the analysis results and sentiment analysis results, the necessary materials (e.g., illustrations and audio) are searched and acquired from a stock material database. The material database is connected to the API of an online stock material site, allowing necessary materials to be searched and acquired programmatically.
[1807] Next, the server generates payment information for the materials used. This information is generated by retrieving information about the material provider and the usage fee from a database and adding them up. The database used here is MySQL, a common relational database.
[1808] Once the user's payment is confirmed, the server generates a video using AI video generation techniques, adjusting the expressions in the video based on the results of sentiment analysis. The video is generated using OpenAI's CLIP model and DeepAI's API.
[1809] Finally, the generated video is served to the user: the server sends a download link for the generated video to the user's device, which the user can use to download or stream the video.
[1810] Specific examples
[1811] Example 1: Manual video for manufacturing industry
[1812] The user enters the "steps to assemble part A" in text on their device and uploads a simple sketch as a storyboard. The server receives this and performs natural language processing and image analysis to identify the necessary illustrations of the parts and audio material for the assembly steps. The emotion engine then analyzes the user's intentions and emotions from the entered text and generates a narration that reflects kindness or strictness. The server searches these in a materials database and generates payment information for the acquired materials. Once the user completes payment, the server uses an AI video generation service to generate a video and provides the user with a download link for the generated video. The user opens the link on their device and downloads and uses the video.
[1813] Example 2: Educational lesson videos
[1814] A user who is an educator enters a text description of the solar system and uploads a simple storyboard for each planet. The server receives and analyzes the text and storyboard to identify illustrations and audio for the solar system. The emotion engine also performs emotional analysis of the entered text and selects passionate narration or a calm tone. The server searches for the relevant material in a material database, calculates the usage fee, and generates payment information. Once the user completes payment, the server uses an AI video generation service to generate a video and provides it to the user. The user can use the provided video in educational settings.
[1815] Prompt Sentence Examples
[1816] Example 1: Manual video for manufacturing industry
[1817] "Please upload a written description of the steps to assemble part A and a simple sketch as a storyboard. After analysis, we will identify the illustrations of the parts and the audio material of the assembly steps and generate a narration that reflects kindness or strictness."
[1818] Example 2: Educational lesson videos
[1819] "Enter a description of the solar system in text and upload a simple storyboard for each planet. After analysis, we'll generate a video with illustrations of the solar system and a choice of passionate or calming narration."
[1820] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1821] Step 1: Data entry
[1822] The user uses a terminal to access a dedicated web page or application and inputs text and storyboards related to the video they want to generate. Specifically, the user inputs the text "Assembly procedure for part A" and uploads the storyboard for part A. The input format is a text file and an image file. This results in the text and storyboards as input data.
[1823] Step 2: Send data
[1824] Data entered by the user is sent from the terminal to the server via an HTTP request. For example, a POST request is used to send text data and image data to the server's endpoint. The data is received by the server and temporarily stored in a database or system memory.
[1825] Step 3: Text analysis
[1826] The server analyzes the received text data using natural language processing (NLP) techniques, specifically Python's NLTK and Spacy libraries. For example, it tokenizes the text, tags parts of speech, and performs syntactic analysis. This process extracts syntactic and semantic data from the input text.
[1827] Step 4: Storyboard Analysis
[1828] The server analyzes the storyboard it receives using image analysis technology. The OpenCV library is used for image analysis, specifically contour detection and shape recognition. For example, it can detect a specific shape (e.g., the shape of part A) from the uploaded storyboard and extract the corresponding data.
[1829] Step 5: Sentiment Analysis
[1830] The server uses an emotion engine to identify emotions contained in the text and storyboards. It uses machine learning frameworks such as Google Cloud NLP for emotion analysis of text and TensorFlow for emotion analysis of storyboards. Specifically, it uses an emotion model to identify emotions such as "joy," "sadness," and "anger."
[1831] Step 6: Material Search
[1832] Based on the analysis and emotion analysis results, the server searches for the necessary materials from a material database. The material database is connected to an online stock material site using an API. For example, it searches for "illustration of part A" or "gentle-toned voice narration" to identify appropriate materials.
[1833] Step 7: Obtaining Materials
[1834] The server retrieves the searched material. This involves downloading data from the stock material site via API. The retrieved material is then stored in a directory or database managed by the server.
[1835] Step 8: Generate payment information
[1836] The server generates payment information for the materials used. The payment information is generated by retrieving information about the material provider and the usage fee from a database and adding them up. Specifically, a specific SQL query is executed using a MySQL database to calculate the usage fee.
[1837] Step 9: Confirm payment
[1838] The server presents the generated billing information to the user. The user enters the payment information (e.g., credit card information). The server processes the payment using an online payment system (e.g., PayPal or Stripe). If the payment is successful, the server confirms it.
[1839] Step 10: Video Generation
[1840] After the user's payment is confirmed, the server uses the analysis results and the acquired materials to execute the AI video generation service. The emotion analysis results are taken into account and the expressions in the video are adjusted. The video is generated using OpenAI's CLIP model and DeepAI's API. The generated video file is saved in the appropriate format.
[1841] Step 11: Submit your video
[1842] The server provides the generated video to the user by generating a download link for the video and sending it to the user's device. The user can then use the link to download or stream the video.
[1843] (Application example 2)
[1844] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1845] Conventional video generation systems simply generate videos based on user-entered text and storyboards, making it difficult to generate personalized videos that reflect the user's emotions. Furthermore, there is a lack of systems that efficiently execute the entire process of delivering generated videos to target users. For this reason, the advertising industry needs a way to quickly generate and deliver effective video ads that match the interests and emotions of target users.
[1846] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1847] In this invention, the server includes means for receiving text and storyboard data input by the user, means for analyzing the text and storyboard, and means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results. This makes it possible to generate high-quality videos that reflect emotions based on the user's input data and distribute them to target users.
[1848] "Means for receiving text and storyboard data entered by the user" refers to a function that allows a user to use a dedicated application or web page via a terminal to enter text and storyboard data and send it to the server.
[1849] The "means for analyzing text and storyboards" refers to a means for analyzing received text and storyboard data using natural language processing and image analysis techniques to understand the content.
[1850] "Means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results" refers to means for searching for and acquiring necessary video and audio materials from a materials database based on the analyzed data and the results of emotion analysis.
[1851] The "means for generating payment information for use of material" is a function for calculating the fee for the used material and generating corresponding payment information.
[1852] The "means for confirming that the user has completed the payment" refers to a means for confirming that the user has made the payment based on the generated payment information.
[1853] "Means for automatically generating videos using materials, analysis results, and emotion analysis results" refers to means for automatically generating videos using AI video generation technology based on the acquired materials and analysis results.
[1854] "Means for providing the generated video to the user" means means for providing the generated video to the user and enabling the user to download or stream the video.
[1855] "Means for distributing the generated video to target users" refers to means for distributing the generated video to target users designated by the advertiser, thereby increasing the effectiveness of the advertisement.
[1856] The embodiments for carrying out the present invention are described in detail below.
[1857] The system aims to enable advertisers to generate personalized video ads using smartphone applications and efficiently deliver them to target users.
[1858] 1. User Input and Data Receipt
[1859] Advertisers use a dedicated smartphone application to input text and storyboards for their advertisements, which are then sent to the server via HTTP requests.
[1860] 2. Data Analysis
[1861] The server analyzes the received text and storyboard data using a natural language processing library (e.g., SpaCy) and an image analysis library (e.g., OpenCV). The analysis includes syntactic analysis of the text data and feature extraction of the image data.
[1862] 3. Emotion recognition
[1863] The emotion engine analyzes emotions based on the analyzed data. A common emotion analysis API (e.g., IBM Watson Natural Language Understanding, Google Cloud Natural Language API) is used as the emotion engine.
[1864] 4. Material Search and Acquisition
[1865] Based on the results of sentiment analysis and data analysis, the server searches for and obtains the necessary video and audio materials from a material database (e.g., Shutterstock API).
[1866] 5. Payment Information Generation and Verification
[1867] The server generates payment information based on the information of the materials used and presents it to the advertiser. Payment processing is performed using an electronic payment service (e.g., Stripe API). Once the advertiser completes payment, the information is recorded on the server.
[1868] 6. Video Generation
[1869] Once payment is confirmed, the server uses the acquired materials and analysis results to automatically generate a video using an AI video generation service (e.g., Synthesia). The generated video reflects narration and direction based on the results of the emotion analysis.
[1870] 7. Video Provision and Distribution
[1871] The server provides the generated video advertisement to the advertiser, who then sends a link to the advertisement via a smartphone application, allowing the advertiser to distribute the video to target users.
[1872] Specific examples
[1873] Example 1: Advertising video for new product "Express Coffee"
[1874] Advertisers enter text and sketch images into the application that explain the features of their new product, "Express Coffee."
[1875] Example sentence: "Express Coffee provides fast, delicious coffee for busy mornings."
[1876] Storyboard: "Image of an Express Coffee package and coffee being poured into a cup"
[1877] Emotion: We want to convey a sense of comfort and trust to the user.
[1878] The server analyzes the input data and searches for and retrieves appropriate content from a content database based on the results of sentiment analysis. Once payment is completed, an AI video generation service is used to generate a video ad that reflects the specified sentiment. The generated video ad link is provided to the advertiser through the application, and the advertiser uses this link to deliver the video to target users.
[1879] This system enables advertisers to quickly generate high-quality, emotionally relevant video ads and deliver them effectively to target users.
[1880] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1881] Step 1:
[1882] Input and Data Reception
[1883] Users use a smartphone application to input text and storyboards to be used in advertising videos. The input data is collected through a form in the application and sent to the server as an HTTP request. The server temporarily stores the received data.
[1884] Specific operation: The user uploads text data (e.g., advertising copy) and image data (e.g., product sketch) into the application's input form and sends it to a dedicated API endpoint.
[1885] Step 2:
[1886] Data analysis
[1887] The server analyzes the received text and storyboard data. Specifically, it analyzes the text using a natural language processing library (e.g., SpaCy) and the storyboard using an image analysis library (e.g., OpenCV). Text analysis is used to understand the meaning and structure of the text, and image analysis is used to extract the features of the storyboard.
[1888] Input and Output: The input data are the received text and storyboard, and the output is the analyzed text data and image data features.
[1889] Specific operation: Using the SpaCy library, tokenize text data, assign POS tags, and analyze dependencies. Using the OpenCV library, detect edges and extract feature points from storyboard images.
[1890] Step 3:
[1891] emotion recognition
[1892] The server uses an emotion engine based on the analyzed data to analyze emotions from the text and storyboard. It uses an emotion analysis API (e.g., IBM Watson Natural Language Understanding, Google Cloud Natural Language API) to identify the user's emotions.
[1893] Input and Output: The input is the analyzed text data and image data features, and the output is the sentiment analysis results.
[1894] Specific operation: The analyzed text data and image data features are sent to the emotion analysis API, and emotions are categorized (e.g., joy, sadness, surprise, etc.) and received.
[1895] Step 4:
[1896] Material Search and Acquisition
[1897] Based on the results of emotion analysis and data analysis, the server searches and retrieves appropriate video and audio materials from a material database (e.g., Shutterstock API).
[1898] Input and output: The input is the emotion analysis results and data analysis results, and the output is the acquired video and audio materials.
[1899] Specific operation: Creates a search query for the Shutterstock API based on the results of sentiment analysis and data analysis, and retrieves the necessary materials (images, audio, video).
[1900] Step 5:
[1901] Payment information generation and verification
[1902] The server generates payment information based on the information about the used materials and presents it to the user. It processes the payment using an electronic payment service (e.g., Stripe API) and confirms that the payment has been completed.
[1903] Input and Output: The input is the information of the material acquired and the output is the status of payment confirmation.
[1904] Specific operation: Calculate the total amount based on the price information of the obtained materials, send a payment request to the user through the Stripe API, and check the status of whether the payment has been completed.
[1905] Step 6:
[1906] Video Generation
[1907] Once payment is confirmed, the server uses the acquired materials and analysis results to automatically generate a video using an AI video generation service (e.g., Synthesia), including narration and direction based on the results of the emotion analysis.
[1908] Input and Output: The input is the acquired material and analysis results, and the output is the generated video file.
[1909] Specific operation: The acquired material, analysis results, and a prompt for a video to be generated based on the emotion analysis results are sent to the Synthesia API, and the generated video is retrieved.
[1910] Step 7:
[1911] Video provision and distribution
[1912] The generated video advertisement is provided to the user (advertiser) from the server, and a link is sent to the user via a smartphone application. The user can then distribute the video to target users.
[1913] Input and Output: The input is the generated video file and the output is the video link provided to the user and the delivery status to the target user.
[1914] Specific operation: The generated video is stored in the server storage, and a download link is generated and sent to the user. The user receives the link and the video is distributed to the target user.
[1915] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1916] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1917] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1918] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1919] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1920] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1921] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1922] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1923] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1924] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1925] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1926] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1927] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1928] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1929] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1930] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1931] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1932] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1933] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1934] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1935] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1936] The following is further disclosed regarding the above embodiment.
[1937] (Claim 1)
[1938] means for receiving user-entered text and storyboard data;
[1939] means for analyzing the text and storyboard;
[1940] a means for searching for and acquiring necessary materials based on the analysis results;
[1941] means for generating payment information for use of said material;
[1942] A means to confirm that the user will complete the payment;
[1943] means for automatically generating a video using the material and the analysis results;
[1944] means for providing the generated video to a user;
[1945] A system including:
[1946] (Claim 2)
[1947] 2. The system of claim 1, wherein the analysis means includes text analysis and image analysis.
[1948] (Claim 3)
[1949] 2. The system according to claim 1, wherein the means for generating payment information acquires and sums information about the material provider and the usage fee.
[1950] "Example 1"
[1951] (Claim 1)
[1952] means for receiving user-entered information and visual material data;
[1953] means for analyzing said information and visual material;
[1954] a means for searching and acquiring necessary data based on the analysis results;
[1955] means for generating fee information for use of said data;
[1956] a means for verifying that the user has completed payment;
[1957] means for automatically generating an image based on the data and the analysis results;
[1958] means for providing the generated video to a user;
[1959] A system including:
[1960] (Claim 2)
[1961] 10. The system of claim 1, wherein the analysis means includes text analysis and image analysis.
[1962] (Claim 3)
[1963] 2. The system according to claim 1, wherein the means for generating fee information acquires and sums information about data providers and usage fees.
[1964] "Application Example 1"
[1965] (Claim 1)
[1966] means for receiving user-entered text and storyboard data;
[1967] means for analyzing the text and storyboard;
[1968] a means for searching for and acquiring necessary materials based on the analysis results;
[1969] means for generating payment information for use of said material;
[1970] A means to confirm that the user will complete the payment;
[1971] means for automatically generating a video using the material and the analysis results;
[1972] means for providing the generated video to a user;
[1973] a means for providing an integrated user interface for a smartphone application;
[1974] AI-based methods to generate videos based on materials retrieved from a materials database, along with payment confirmation;
[1975] A system including:
[1976] (Claim 2)
[1977] 2. The system of claim 1, wherein the analysis means includes natural language processing and image analysis.
[1978] (Claim 3)
[1979] 2. The system according to claim 1, wherein the means for generating payment information acquires and sums information about the material provider and the usage fee.
[1980] "Example 2: Combining Emotion Engines"
[1981] (Claim 1)
[1982] means for receiving user-entered text and storyboard data;
[1983] means for analyzing the text and storyboard;
[1984] a means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results;
[1985] means for generating payment information for use of said material;
[1986] A means to confirm that the user will complete the payment;
[1987] means for automatically generating a video using the material, the analysis results, and the emotion analysis results;
[1988] means for providing the generated video to a user;
[1989] A system including:
[1990] (Claim 2)
[1991] 2. The system of claim 1, wherein the analysis means includes text analysis and image analysis.
[1992] (Claim 3)
[1993] 2. The system according to claim 1, wherein the means for generating payment information acquires and sums information about the material provider and the usage fee.
[1994] "Application example 2 when combining emotion engines"
[1995] (Claim 1)
[1996] means for receiving user-entered text and storyboard data;
[1997] means for analyzing the text and storyboard;
[1998] a means for searching for and acquiring necessary materials based on the analysis results and emotion analysis results;
[1999] means for generating payment information for use of said material;
[2000] A means to confirm that the user will complete the payment;
[2001] means for automatically generating a video using the material, the a...
Claims
1. means for receiving user-entered text and storyboard data; means for analyzing the text and storyboard; a means for searching for and acquiring necessary materials based on the analysis results; means for generating payment information for use of said material; A means to confirm that the user will complete the payment; means for automatically generating a video using the material and the analysis results; means for providing the generated video to a user; A system including:
2. The system of claim 1 , wherein the analysis means includes text analysis and image analysis.
3. 2. The system according to claim 1, wherein said means for generating payment information acquires and sums information on material providers and usage fees.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A