System
The system allows users to easily generate and modify high-quality videos by using a terminal for input and a server for deep learning-based image and video generation, addressing the complexity of conventional video generation technologies.
Patent Information
- Application Number
- JP2024141336
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-06
AI Technical Summary
Conventional image and video generation technologies require specialized skills, making it difficult for ordinary users to easily realize their own ideas and images, and the process of generating a video according to user instructions and incorporating modifications is complex and time-consuming.
A system that includes a terminal for receiving user input data in the form of text or sketches, a server for generating initial images and successive frames using deep learning models, and allowing users to specify corrections for video regeneration without specialized skills.
Enables users to intuitively and easily create high-quality videos by automatically generating, modifying, and regenerating moving images based on user input, reducing the complexity and time required for video creation and modification.
Smart Images

Figure 2026038002000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional image and video generation technologies require specialized skills, making it difficult for ordinary users to easily realize their own ideas and images. Furthermore, the process of generating a video according to a user's instructions and then incorporating modifications into the video is complex and time-consuming. The present invention aims to solve these problems and provide a system that allows users to intuitively and easily generate high-quality videos. [Means for solving the problem]
[0005] The present invention provides a system including a means for receiving user input data, a means for generating an initial image based on the user input data, a video generation means for generating successive frames based on the initial image, a means for transmitting the generated video to a user, a means for receiving corrections specified by the user, and a means for regenerating the video based on the corrections. This allows users to easily embody their ideas and images as videos and further modify the videos without requiring specialized skills. Specifically, the system is configured such that the user provides input data in the form of text or sketches, and the video generation means generates a video including continuous movements using a deep learning model.
[0006] Below are definitions of important terms contained in the claims.
[0007] "User input data" refers to text information or sketches that a user provides to the system.
[0008] "Initial image" refers to a still image generated based on user input data.
[0009] "Sequential frames" refers to a series of images generated to represent a moving scene.
[0010] "Video generator" refers to a system or device for generating a sequence of frames based on an initial image.
[0011] "User-specified modifications" refers to changes or additional information that a user requests for the generated video.
[0012] "Regeneration" refers to the process of regenerating a video based on modifications specified by the user.
[0013] A "deep learning model" refers to an algorithm that uses neural network technology to generate images and videos. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The embodiment of the present invention will now be described in detail.
[0036] System Structure
[0037] The present invention provides a system that allows users to easily embody their own ideas and images as animations. The system consists of a terminal that receives input data from the user and a server that processes the data based on the input data.
[0038] Program processing flow
[0039] 1. User Input
[0040] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if a user wants a video of a deer walking through a deep forest, they can write that in text or input a simple sketch. This input data is sent to the device.
[0041] 2. Sending input data
[0042] The terminal sends the user's input data to the server, which receives it and prepares to generate an initial image based on the user's wishes.
[0043] 3. Generate initial images
[0044] The server uses an image generation AI model to generate an initial image based on the user's input data. This model uses the latest neural network technology, for example. For example, an initial image representing the user's input scene of "a deer walking in a deep forest" is generated.
[0045] 4. Video Generation
[0046] After the initial images are generated, the server then uses a deep learning-based video generation AI model to generate a series of frames to create a video. The server automatically generates frames to smoothly express the required movement between images.
[0047] 5. Send and review your video
[0048] The generated video is sent to the device and can be viewed by the user, who can then review the video and input any corrections or changes they wish to make.
[0049] 6. Reflecting the changes and regenerating
[0050] The device sends the user-specified corrections to the server, and the server reflects the corrections and regenerates the video. For example, if the user instructs the server to "make the deer walk faster," the correction is reflected.
[0051] 7. Final check and output
[0052] The regenerated video is sent back to the user. The user finally reviews the video to their satisfaction and proceeds to save or share. The device provides the final video file and generates a sharing link if necessary.
[0053] Specific examples
[0054] Specific examples are shown below.
[0055] Example: A user requests a video of a puppy running on a beach at sunset and enters that in the text.
[0056] Initial image: The server generates an initial image of a beach at sunset and a video of a puppy running.
[0057] Example fix: The user inputs the command "Make the puppy faster and add the sound of waves."
[0058] Final output: A video is generated that reflects the edits, and the user is happy to save and share the video.
[0059] As described above, the present invention enables users to intuitively and easily create high-quality moving images by automatically generating, modifying, and regenerating moving images based on user input.
[0060] The processing flow will be explained below.
[0061] Step 1:
[0062] The user opens a dedicated application or web interface, enters the desired video content in text or draws a rough sketch, and is ready to send their specific request to the device.
[0063] Step 2:
[0064] The device checks the user's input data and sends it to the server, including text information and sketches entered by the user, as well as optional settings.
[0065] Step 3:
[0066] The server receives the input data, analyzes it, and prepares to generate an initial image based on the user's request.
[0067] Step 4:
[0068] The server uses an image generation AI model to generate an initial image based on the user's input data. For example, if the user's text input is "A deer is walking in a deep forest," it generates an initial image that represents that scene.
[0069] Step 5:
[0070] The server prepares the video based on the generated initial images, which includes preprocessing to generate successive frames.
[0071] Step 6:
[0072] The server uses a video generation AI model to generate a series of frames that smoothly represent the required movement between the initial images to create the video. This model uses deep learning techniques, for example, to reproduce natural movements.
[0073] Step 7:
[0074] The server sends the generated video to the terminal, which receives the video and notifies the user to confirm it.
[0075] Step 8:
[0076] The user reviews the received video and decides whether they are satisfied. If the user wishes to make any corrections, they input specific corrections (e.g., color adjustments, changing the speed of the motion, etc.) and send them to the device.
[0077] Step 9:
[0078] The device sends the user's modified data to the server, which receives it and prepares to generate a new video that reflects the modifications.
[0079] Step 10:
[0080] The server regenerates the video based on the corrections, and then regenerates successive frames to complete the entire video.
[0081] Step 11:
[0082] The regenerated video is sent to the device again, and the device receives the video and notifies the user for final confirmation.
[0083] Step 12:
[0084] The user reviews the final video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[0085] Through these steps, users can embody their own images as videos and easily modify and regenerate them without requiring any specialized skills.
[0086] Example 1
[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0088] Conventional video generation systems have struggled to easily and intuitively create high-quality videos based on user input data. In particular, it is technically complex to generate videos with natural, continuous movement based on input data in different formats, such as text and sketches. Furthermore, regenerating videos based on user correction requests requires cumbersome procedures and takes a long time.
[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0090] In this invention, the server includes a terminal that receives user input data, a means for transmitting the user input data to the server, a means for utilizing an image generation AI model that generates an initial image based on the user input data, a means for utilizing a video generation AI model that generates successive frames based on the initial image, a means for transmitting the generated video to the terminal, a terminal that receives corrections specified by the user, a means for regenerating the video based on the corrections, and a means for transmitting the regenerated video to the terminal. This allows users to easily and intuitively create high-quality videos and quickly respond to correction requests.
[0091] "User" refers to an individual or organization that uses this system to generate moving images.
[0092] "Input Data" refers to information required for video generation that is provided by a user through a dedicated application or web interface, and may include formats such as text or sketches.
[0093] "Terminal" means a device used by a user to provide input data and view the generated video, including a smartphone, tablet, computer, etc.
[0094] The "server" is a central processing unit for generating moving images based on data input by the user, and performs the functions of data analysis, image generation, moving image generation, and correction reflection.
[0095] An "image generation AI model" is an artificial intelligence model that uses neural network technology to generate initial images based on user input data.
[0096] A "video generation AI model" is an artificial intelligence model that generates successive frames based on initial images to create videos that include natural movements.
[0097] The "corrections" are specific instructions that the user wishes to add or change to the generated video.
[0098] "Regeneration" is the process of regenerating a video by reflecting the corrections specified by the user.
[0099] MODE FOR CARRYING OUT THE INVENTION
[0100] This invention is a system that allows users to easily embody their own ideas and images as animations. This system consists of a terminal that receives input data from the user and a server that processes the data based on that input.
[0101] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if a user wants a video of a deer walking in a deep forest, they can write that in text or enter a simple sketch. The device receives this input data and sends it to the server.
[0102] The server receives the user's input data and generates an initial image using an image generation AI model. This model uses the latest neural network technology, such as Stable Diffusion. The generated initial image will represent the scene the user input, such as "a deer walking in a deep forest."
[0103] The server then uses a deep learning-based video generation AI model to generate a series of frames to create a video, using video generation technologies such as DALL-E 2. The deep learning model generates frames based on the initial images to create a natural-looking sequence of motion.
[0104] The generated video is sent to the device and can be viewed by the user. The user watches the video and inputs any desired corrections or changes in text. These corrections are then sent back to the server via the device.
[0105] The server receives the corrections specified by the user and regenerates the video using the image generation AI model and video generation AI model. For example, if the user requests that the deer walk faster, the server adjusts the video generation parameters according to the request and regenerates the video.
[0106] The regenerated video is sent back to the device, and this process is repeated until the user is satisfied. Finally, when a video that the user is satisfied with is generated, the device provides the video file and proceeds to the storage or sharing stage. For example, a sharing link for the video can be generated using a storage service or file sharing service.
[0107] Prompt Sentence Examples
[0108] Below is an example of a prompt sentence to input to the generative AI model.
[0109] Prompt: "Create a video of a deer walking through a deep forest."
[0110] Prompt: "Draw a scene of a puppy running on a beach at sunset."
[0111] In this way, the present invention allows users to intuitively and easily create high-quality videos by automatically generating, modifying, and regenerating videos based on user input.
[0112] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0113] Step 1:
[0114] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if they want a video of a deer walking through a deep forest, they can write a specific prompt in text or upload a simple sketch. This input data is sent to the device.
[0115] Input: User input data via text or sketch
[0116] Output: Input data stored on the device
[0117] Step 2:
[0118] The device sends the user's input data to the server using the secure HTTP / HTTPS protocol. The server analyzes the received data and prepares for initial image generation.
[0119] Input: Input data stored on the device
[0120] Output: The input data sent to the server
[0121] Step 3:
[0122] The server invokes an image generation AI model to generate an initial image based on the user's input data. This model uses the latest neural network technology, Stable Diffusion. The server inputs a prompt statement into the model and generates an initial image that embodies the specified scene.
[0123] Input: User input data, prompt text
[0124] Output: The initial image generated
[0125] Step 4:
[0126] The server generates a series of frames based on the initial image to create a video. This task uses DALL-E 2, a deep learning-based video generation AI model. The server inputs the initial image into the AI model and generates a series of frames that naturally express continuous movement to create a video.
[0127] Input: Initial image
[0128] Output: Video with consecutive frames
[0129] Step 5:
[0130] The generated video is sent to the device and can be viewed by the user, who can watch the video and enter any desired corrections or changes in text.
[0131] Input: Generated video
[0132] Output: User-supplied text of corrections and changes
[0133] Step 6:
[0134] The device sends the user-specified corrections to the server. The server analyzes the corrections and regenerates the video using the image generation AI model and video generation AI model. For example, if the user instructs the device to "make the deer walk faster," the server reflects that parameter.
[0135] Input: Text of user corrections or changes
[0136] Output: Regenerated video
[0137] Step 7:
[0138] The regenerated video is sent to the device, and the user repeats this process until they are satisfied. Finally, the device can also generate a sharing link in conjunction with a storage or file sharing service to save or share the final video.
[0139] Input: Regenerated video
[0140] Output: Final video file, share link (optional)
[0141] The above is the specific processing flow of the program of this system.
[0142] (Application example 1)
[0143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0144] Currently, creating advertisements on the market requires a high level of expertise, time, and resources. Small businesses and individual marketers in particular face challenges in creating advertising videos quickly and effectively. There is a need for a way for users to easily realize their ideas and create, edit, and share videos suitable for advertising media.
[0145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0146] In this invention, the server includes means for receiving user input data, means for generating an initial image based on the user input data, video generation means for generating successive frames based on the initial image, means for transmitting the generated video to the user, means for receiving corrections specified by the user, means for regenerating the video based on the corrections, and means for saving and sharing the user-generated video on an advertising medium, thereby enabling users to easily create high-quality advertising videos and quickly edit and share them.
[0147] "Means for receiving user input data" refers to a device or interface through which a user inputs their ideas or desired images in the form of text, sketches, or the like.
[0148] "Means for generating an initial image" refers to the algorithm or software that creates the initial still image based on user input data.
[0149] "Video generation means for generating successive frames" refers to a technology for creating frames that change continuously from a generated initial image and outputting them as a video.
[0150] "Means for transmitting the generated video to the user" refers to a protocol or system for transferring the generated video file to the user terminal via a network.
[0151] The "means for receiving corrections specified by the user" refers to an interface that allows the user to instruct changes or corrections to the generated video, and a system that receives the content of those changes or corrections.
[0152] "Means for regenerating a video based on the modifications" refers to algorithms or software for regenerating a video that reflects the modifications specified by the user.
[0153] "Means for storing and sharing on advertising media" refers to systems and technologies for storing the generated videos on advertising platforms such as websites, social media, and email, and sharing them with other users.
[0154] This invention provides a system that allows users to easily generate high-quality advertising videos and modify and share them as needed. This system includes means for receiving user input data, means for generating an initial image, means for generating a video that generates a series of frames, means for sending the generated video to the user, means for receiving modifications specified by the user, means for regenerating the video based on the modifications, and means for saving and sharing the generated video on an advertising medium.
[0155] Hardware and software used
[0156] Hardware: The device (smartphone, tablet, smart glasses, etc.) and server through which the user manipulates input data.
[0157] Software: Generative AI models (e.g., Stable Diffusion and DALL-E) and deep learning-based video generation models (e.g., DeepMind and OpenAI® Video GPT) are implemented on cloud servers.
[0158] Processing flow and data calculation
[0159] 1. User Input
[0160] Users input their ideas for the ad video they want to generate in text or sketch form. For example, they can input a prompt in text such as "A new smartphone flying through the air."
[0161] 2. Data transmission and initial image generation
[0162] The device sends input data to a cloud server, which uses a generative AI model to generate an initial image based on the input data, such as a scene of a new smartphone flying through the air.
[0163] 3. Video Generation
[0164] An initial image is input to a deep learning-based video generation model, which then generates successive frames based on the initial image, generating a continuous video from the initial image.
[0165] 4. Sending videos and receiving corrections
[0166] The generated video is sent to the device for the user to review. An interface is provided for the user to input desired corrections or changes, such as "increase the smartphone's flight speed and change the background to a sunset."
[0167] 5. Modify and Regenerate
[0168] The server then reflects the received corrections and regenerates the video. The corrections are made automatically by the AI model, significantly reducing the user's workload.
[0169] 6. Final confirmation, saving and sharing
[0170] The regenerated video is sent to the device for final confirmation by the user. If the user is satisfied, the video is saved and shared on advertising media (website, social media, email, etc.).
[0171] Specific examples
[0172] Prompt Sentence Examples
[0173] "I want to create a video of my new smartphone flying through the air."
[0174] Correction prompt: "Make my phone fly faster and change the background to a sunset."
[0175] This system allows users to easily realize their ideas and quickly create, edit, and share high-quality advertising videos.
[0176] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0177] Step 1:
[0178] Getting user input data
[0179] - Input: The user inputs the desired video content in the form of text or a sketch. For example, the user inputs "a video of a new smartphone flying in the air."
[0180] - Action: The user provides input data using a dedicated application or a web interface.
[0181] - Output: The input data is stored inside the application and is ready to be sent to the cloud server.
[0182] Step 2:
[0183] Generate the initial image
[0184] - Input: User input data sent to the cloud server.
[0185] - Operation: The server passes the received input data to a generative AI model, which generates an initial image from the text or sketch. The generative AI model uses, for example, Stable Diffusion or DALL-E.
[0186] - Output: The generative AI model generates an initial image, for example, an image of a new smartphone flying through the air.
[0187] Step 3:
[0188] Video frame generation
[0189] - Input: The generated initial image.
[0190] Operation: The server inputs the initial image into a deep learning-based video generation model to generate successive frames. The deep learning model used is, for example, Video GPT.
[0191] - Output: A series of frames are generated from the initial image and compiled into a video.
[0192] Step 4:
[0193] Submitting and reviewing videos
[0194] - Input: The generated video file.
[0195] - Operation: The server sends the generated video file to the user's device. The user checks the received video and specifies any corrections.
[0196] - Output: The video sent to the user's device and the modifications specified by the user (e.g., "increase the smartphone's flying speed and change the background to a sunset") are obtained.
[0197] Step 5:
[0198] Applying the changes and regenerating
[0199] - Input: User-specified corrections.
[0200] - Operation: The server uses the AI model and video generation model to modify the video based on the user's modifications. Specifically, it reflects instructions such as flight speed and background changes.
[0201] - Output: A new video file is generated with the modifications reflected.
[0202] Step 6:
[0203] Final confirmation, saving and sharing
[0204] - Input: The modified video file.
[0205] - Operation: The server resends the modified video file to the user's device. The user performs a final check and saves or shares the video on the advertising medium if satisfied.
[0206] - Output: The final video file is saved on the user's device, and a sharing link is generated to the specified advertising medium.
[0207] By following these steps, users can easily and quickly create, modify, and share high-quality advertising videos.
[0208] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0209] The embodiment of the present invention will now be described in detail.
[0210] System Structure
[0211] This invention provides a system that allows users to easily embody their ideas and images as videos. The system consists of a terminal that receives user input data and a server that processes the data. Furthermore, by incorporating an emotion engine that recognizes the user's emotions, it becomes possible to generate videos based on the user's emotions.
[0212] Program processing flow
[0213] 1. User Input
[0214] Using a dedicated application or web interface, users input the desired video content using text or sketches. For example, if a user wants to add the emotion "calm" to a video of a deer walking in a deep forest, they input that information. The input data is then sent to the device.
[0215] 2. Sending input data and emotion data
[0216] The device sends the user's input data and emotion data together to the server. The transmitted data includes the text information and sketches entered by the user, as well as emotion tags.
[0217] 3. Emotion Recognition and Initial Image Generation
[0218] The server receives the input data and emotion data. The server uses an emotion engine to analyze and identify the user's emotion. The recognized emotion is reflected in the initial image generation means. For example, if the user's emotion is "calm," an initial image with a calm atmosphere is generated to reflect this.
[0219] 4. Video Generation
[0220] After the initial image is generated, the server uses a video generation AI model to generate a series of frames to create a video. The model also incorporates the user's emotional data to generate videos that are more emotionally relevant. For example, a calm scene would use gentle movements and soft lighting.
[0221] 5. Send and review your video
[0222] The generated video is sent to the device and can be viewed by the user. The user can review the video and input corrections or changes as necessary.
[0223] 6. Reflecting and regenerating emotionally based modifications
[0224] When the user inputs corrections, the device sends the correction data to the server. The server then generates a new video based on the correction data and emotion data. The corrections are optimized based on the emotion data.
[0225] 7. Final check and output
[0226] The regenerated video is sent back to the device. The user finally reviews the video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[0227] Specific examples
[0228] Specific examples are shown below.
[0229] Example: A user inputs a video of a puppy running on a beach at sunset and adds the emotion "happiness."
[0230] Initial image: The server generates an initial image of a beach at sunset, and the emotion engine creates a scene with bright colors and a feeling of happiness.
[0231] Video Generation: A video of a puppy running on the beach is generated, with lively and happy movements.
[0232] Example fix: The user suggests making fixes like "Make the puppy a little faster and add the sound of the ocean."
[0233] Final output: A video is generated that reflects the edits and the user is happy to save and share the video.
[0234] As described above, the present invention is a system that incorporates a user's emotional data and generates videos that match the user's emotions, allowing the user to easily generate more specific and emotionally rich videos.
[0235] The processing flow will be explained below.
[0236] Step 1:
[0237] The user opens a dedicated application or web interface, enters the desired video content in text or draws a simple image using a sketch, and also enters the emotion they want the video to reflect (e.g., "calm" or "happy"). This prepares the device to send the specific request and emotional data.
[0238] Step 2:
[0239] The device checks the user's input data and emotion data and transmits them to the server. The transmitted data includes the user's input text information, sketches, and emotion tags.
[0240] Step 3:
[0241] The server receives the input data and emotion data, analyzes the received data, and prepares to generate an initial image based on the user's request and emotion.
[0242] Step 4:
[0243] The server uses an emotion engine to analyze the emotions contained in the user's text data and sketches. For example, if the user's emotion is analyzed as "calm," it will recognize it.
[0244] Step 5:
[0245] The server uses an image generation AI model to generate an initial image based on the user's input data and emotion data. For example, based on the user's input text "A deer walking in a deep forest" and the emotion "Calm," it generates an initial image with a calm atmosphere.
[0246] Step 6:
[0247] The server uses a video generation AI model based on the generated initial images to generate a series of frames to create a video. This model also reflects the user's emotional data, generating videos that are in line with their emotions. For example, gentle movements and soft lighting are used for calm scenes.
[0248] Step 7:
[0249] The server sends the generated video to the terminal, which receives the video and notifies the user to confirm it.
[0250] Step 8:
[0251] The user reviews the received video and decides whether they are satisfied. If the user wishes to make any corrections, they input specific corrections (e.g., changing the speed of the puppy's movements, adjusting the color) and send them to the device.
[0252] Step 9:
[0253] The device sends the user's modified data to the server, which receives it and prepares to generate a new video that reflects the modifications.
[0254] Step 10:
[0255] The server regenerates the video based on the modifications. Based on the modifications and emotion data, it regenerates consecutive frames to complete the entire video. For example, if a user requests the puppy to move a little faster, a video reflecting the speed change is generated.
[0256] Step 11:
[0257] The regenerated video is sent to the device again, and the device receives the video and notifies the user for final confirmation.
[0258] Step 12:
[0259] The user reviews the final video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[0260] Through these steps, users can embody their own images into emotive videos without requiring any specialized skills, and can easily modify and regenerate them.
[0261] Example 2
[0262] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0263] Currently, many video generation systems generate videos based on user input data, but lack the technology to adjust the content and atmosphere of the video based on the user's emotions. This makes it difficult to easily generate videos that embody the user's intended emotions and atmosphere. Furthermore, there is no way to reflect emotions when modifying generated videos, making it difficult to obtain results that satisfy the user.
[0264] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0265] In this invention, the server includes means for generating an initial image based on user input data and emotion data, moving image generation means for generating successive frames based on the initial image, and means for transmitting the generated moving image to the user. This makes it possible to generate a moving image that reflects the emotion data input by the user, and to regenerate the moving image while taking emotion into consideration even in subsequent modifications.
[0266] "User input data" refers to information about text, sketches, and video content entered by a user using a dedicated application or web interface.
[0267] "Emotion data" is information that represents the emotion or atmosphere that the user wants to impart to the content of the video.
[0268] An "initial image" is the first still image of a video generated based on the user's input data and emotion data.
[0269] "Video generation means" refers to the function or technology for generating a series of frames based on initial images and emotion data to create a video.
[0270] "Correction points" refer to requests for changes or corrections that a user specifies for a generated video.
[0271] "Regeneration" is the process of recreating a video based on user-specified modifications and emotion data.
[0272] An "emotion engine" is a software module or algorithm for analyzing and recognizing a user's emotional data.
[0273] A "deep learning model" is an artificial intelligence technology that uses large amounts of data to learn and generate sophisticated videos from the input data.
[0274] The present invention provides a system for generating videos that reflect a user's emotions by incorporating the user's emotional data. This system allows users to easily embody their own ideas and images as videos.
[0275] System Structure
[0276] The system consists of a terminal that receives user input data and a server that processes the data. Furthermore, by incorporating an emotion engine, it becomes possible to generate videos based on the user's emotions.
[0277] Hardware and software used
[0278] Device: The device on which the user provides input data (e.g., smartphone, tablet, computer)
[0279] Server: A server that processes input data and emotion data and generates videos.
[0280] Emotion engine: A software module that analyzes and recognizes user emotion data.
[0281] Generative AI model: A deep learning model that generates videos based on input data and emotion data
[0282] A concrete example of the video generation flow
[0283] User Input
[0284] The user inputs the desired video content using a dedicated application or web interface. For example, if the user wants to add the emotion "calm" to a video of a deer walking in a deep forest, the user inputs the information in text and selects "calm" as the emotion tag. This input is sent to the device.
[0285] Example prompt sentence:
[0286] Add the emotion "Happiness" to a video of a puppy running on a beach at sunset.
[0287] Sending input data and emotion data
[0288] The device transmits the user's input data and emotion data to the server. The transmitted data includes the text information and sketches entered by the user, as well as emotion tags.
[0289] Emotion Recognition and Early Image Generation
[0290] The server analyzes the received input data and emotion data and recognizes the emotion using the emotion engine. The recognized emotion is reflected in the initial image generation means. For example, if the user adds the emotion "calm," the server generates an initial image with a calm atmosphere.
[0291] Video generation
[0292] The server then uses a generative AI model to generate a series of frames based on the initial images, creating a video. This process also incorporates the user's emotional data, generating videos that are more in line with their emotions. For example, calm scenes will use gentle movements and soft lighting.
[0293] Submitting and reviewing videos
[0294] The generated video is sent to the device and can be viewed by the user, who can review the video and enter corrections or changes as necessary.
[0295] Reflecting and regenerating emotion-based modifications
[0296] When the user inputs corrections, the device sends the correction data to the server. The server then generates a new video based on the correction data and emotion data. The corrections are also optimized based on the emotion data.
[0297] Final check and output
[0298] The regenerated video is then sent back to the device, where the user can finally review the video to their satisfaction and save or share it. The device will provide the final video file and generate a sharing link if necessary.
[0299] As described above, the present invention provides a system that allows users to easily create and edit videos that reflect their own emotions and intentions. This system utilizes emotion data to provide more specific and emotionally rich videos.
[0300] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0301] Step 1:
[0302] Using a dedicated application or web interface, users input the content of the desired video and related emotional data. This input includes text information, sketches, and emotion tags. This input data is sent to the device. Specifically, the user inputs "a video of a deer walking in a deep forest" and the emotion "calm," and the data is sent to the device.
[0303] Input: Text information, sketch, emotion tag
[0304] Output: Input data sent to the terminal
[0305] Step 2:
[0306] The device combines the user's input data and emotion data into a single data packet and sends it to the server. Specifically, it combines the text information "A deer is walking in a deep forest" and the emotion tag "Calm" in JSON format. The device creates a data packet and sends it to the server.
[0307] Input: User input data and emotion data
[0308] Output: Data packet sent to the server
[0309] Step 3:
[0310] The server analyzes the received data packets and uses the emotion engine to recognize and analyze the user's emotions. Specifically, the server analyzes the received data and recognizes the emotion "calm." This emotion data is reflected in the initial image generation means, as it affects subsequent processing.
[0311] Input: Data packet (user input data and emotion data)
[0312] Output: Recognized emotion data
[0313] Step 4:
[0314] The server uses the initial image generation means to generate an initial image based on the user's input data and the recognized emotion data. Specifically, an initial image depicting a forest landscape in gentle colors is created. This initial image is used as the first step in video generation.
[0315] Input: User input data, recognized emotion data
[0316] Output: Initial image
[0317] Step 5:
[0318] The server uses a generative AI model to generate a series of frames based on the generated initial image, creating a video. Specifically, the video generation AI model receives the initial image and emotion data as input and generates a series of frames of a deer in a tranquil scene. This series of frames is then combined to form a complete video.
[0319] Input: Initial image, recognized emotion data
[0320] Output: Video with continuous frames
[0321] Step 6:
[0322] The generated video is sent from the server to the device, where the user can view it. The user plays the video through the application and checks the content. Specifically, the video data from the server arrives at the device, and the user clicks the "play" button to watch the video.
[0323] Input: Generated video
[0324] Output: Video sent to device, user confirmation
[0325] Step 7:
[0326] If a user wishes to make corrections to the video content, they input the corrections and send the correction data from their device to the server. Specifically, the user might input, "Make the deer move a little faster and add the sound of birds singing in the background." This correction data is then sent from the device to the server.
[0327] Input: User modifications
[0328] Output: Corrected data sent to the server
[0329] Step 8:
[0330] The server regenerates the initial image and video based on the received correction data and emotion data. Specifically, the server generates new initial images and consecutive frames, and regenerates a video that reflects the corrections and emotion data.
[0331] Input: Corrected data, recognized emotion data
[0332] Output: Regenerated video
[0333] Step 9:
[0334] The regenerated video is sent from the server to the device, and the user is finally confirmed to be satisfied. The user watches the video again, and if satisfied, can save or share it. Specifically, the server sends the regenerated video to the device, and if the user is satisfied, they click the "Save" or "Share" button.
[0335] Input: Regenerated video
[0336] Output: User final review, save or share
[0337] (Application example 2)
[0338] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0339] Current video generation systems make it difficult for users to easily create videos that reflect their desired emotions. They also lack the functionality to easily share generated videos on online platforms. This makes it difficult to quickly create and share emotionally rich, effective videos, especially in the advertising field.
[0340] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0341] In this invention, the server includes means for receiving user input data, means for generating an initial image based on the user input data, video generation means for generating successive frames based on the initial image, means for transmitting the generated video to the user, means for receiving corrections specified by the user, means for regenerating the video based on the corrections, means for the video generation means to generate a video reflecting the user's emotional data, an emotion engine for analyzing the user's emotional data, and means for generating a link for sharing the generated video on an online platform. This enables users to easily generate effective videos that reflect emotions and quickly share them online.
[0342] "User input data" refers to information such as text or sketches that a user provides to the system.
[0343] "Initial image" refers to the first still image generated based on user input data.
[0344] "Video generator" refers to a mechanism for generating a sequence of video frames based on an initial image.
[0345] "Correction points" refer to requests for changes or corrections that a user specifies for a video.
[0346] "Regeneration" refers to regenerating a video by reflecting the corrections.
[0347] "Emotional data" refers to information about emotions that a user provides to the system.
[0348] An "emotion engine" refers to software or algorithms that analyze the emotional data provided by the user and use that data to generate videos.
[0349] "Link for sharing on online platforms" refers to a URL or sharing link for sharing the generated video with other users on the Internet.
[0350] "Consecutive frames" refers to a number of still images that make up a moving image.
[0351] "Means for transmitting to the user" refers to a communication means for providing the generated video to the user.
[0352] The present invention relates to a system for enabling users to quickly and effectively create emotional videos and easily share them on an online platform, and a specific embodiment of the system will be described below.
[0353] System Configuration
[0354] The system of the present invention is composed of a terminal that receives user input data, a server that processes the data, and a communication means that provides the generated video to the user. It also includes an emotion engine for emotion recognition and an AI model for video generation.
[0355] Hardware and software used
[0356] Terminal: A device that allows users to input text and sketches necessary for video generation. This includes smartphones and PCs.
[0357] Server: Provides the computing resources to analyze user input data and emotion data and generate videos. Cloud-based servers are often used.
[0358] Emotion engine: Software for analyzing emotion data and generating initial images based on it.
[0359] Video generation AI model: A deep learning model that generates a series of frames based on initial images to create emotionally relevant videos.
[0360] Communication means: A network that transmits the generated video data to the terminal and allows the user to view the video.
[0361] Video generation process
[0362] 1. Getting User Input
[0363] Users can use a dedicated application or web interface to input the content of the video and the desired emotion using text or a sketch. For example, they can input "a video introducing the latest running shoes" and the emotion "excitement."
[0364] 2. Data transmission and processing
[0365] The device sends the user's input data and emotion data to the server, which receives the data, analyzes the emotion using an emotion engine, and generates an initial image.
[0366] 3. Generate initial images
[0367] The emotion engine analyzes the user's emotional data and generates an initial image that reflects that emotion. If the emotion is "excitement," vivid colors and dynamic compositions are used.
[0368] 4. Video Generation
[0369] The server uses a video generation AI model based on the generated initial images to generate successive frames, which then creates a video that reflects the emotion.
[0370] 5. Send and review your video
[0371] The generated video is sent from the server to the device, where the user can review it and make corrections if necessary.
[0372] 6. Processing and regenerating fixes
[0373] If the user inputs corrections, the data is sent to the server again, and the server generates a new video based on the corrections and emotion data.
[0374] 7. Video sharing
[0375] The final video is provided as a link that users can use to share the video on online platforms.
[0376] Examples and prompts
[0377] As a concrete example, consider a case where a user wants to generate an advertising video of a family picnic scene in a park with spring flowers blooming with the emotion of "happiness." The user inputs the following prompt sentence:
[0378] Example prompt sentence
[0379] Advertising video of a family picnic scene in a park with spring flowers
[0380] Emotion: euphoria
[0381] Based on this prompt, a video is generated that reflects the user's emotions, showing smiling faces and warm colors. The generated video can be easily shared on the user's preferred platform.
[0382] The above is a specific embodiment for carrying out the present invention. This system enables advertisers to quickly create emotionally rich and effective videos and use them in their marketing activities.
[0383] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0384] Step 1:
[0385] Users use a dedicated application or web interface to input text, sketches, and emotional information required for video generation. Input includes "a video introducing the latest running shoes" and the emotion "excitement." The device collects this input data and stores it as part of its data.
[0386] Step 2:
[0387] The device sends the user's input data (text, sketches, emotion information) to the server, which receives it and prepares each data for analysis. The input data is organized for further processing.
[0388] Step 3:
[0389] The server uses an emotion engine to analyze the user's emotional information. It generates an initial image based on the results of this analysis. If the input data is "excitement," an initial image with bright colors and a dynamic composition is generated. This initial image becomes the basis for the next video generation.
[0390] Step 4:
[0391] The server supplies the generated initial images to a video generation AI model, which generates successive frames. The generative AI model generates successive frames based on the initial images and emotion data received as input. The output is video data that reflects the emotion.
[0392] Step 5:
[0393] The server sends the generated video to the terminal. The terminal receives the video data and provides it to the user for confirmation. The user can view the video and input corrections as necessary.
[0394] Step 6:
[0395] If the user inputs corrections, the device sends the correction data back to the server, which then receives the corrections and reuses the emotion engine and video generation AI model to generate a new video that reflects the corrections.
[0396] Step 7:
[0397] The server then sends the new video with the modifications reflected back to the device, and the user repeats this process until they are finally satisfied, at which point they are asked to generate a link to save the video or share it on an online platform.
[0398] As described above, users, devices, and servers work together to efficiently generate videos that reflect emotions, and then modify and optimize them as needed, completing the entire process of sharing them online.
[0399] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0400] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0401] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0402] [Second embodiment]
[0403] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0404] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0405] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0406] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0407] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0408] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0409] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0410] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0411] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0412] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0413] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0414] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0415] The embodiment of the present invention will now be described in detail.
[0416] System Structure
[0417] The present invention provides a system that allows users to easily embody their own ideas and images as animations. The system consists of a terminal that receives input data from the user and a server that processes the data based on the input data.
[0418] Program processing flow
[0419] 1. User Input
[0420] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if a user wants a video of a deer walking through a deep forest, they can write that in text or input a simple sketch. This input data is sent to the device.
[0421] 2. Sending input data
[0422] The terminal sends the user's input data to the server, which receives it and prepares to generate an initial image based on the user's wishes.
[0423] 3. Generate initial images
[0424] The server uses an image generation AI model to generate an initial image based on the user's input data. This model uses the latest neural network technology, for example. For example, an initial image representing the user's input scene of "a deer walking in a deep forest" is generated.
[0425] 4. Video Generation
[0426] After the initial images are generated, the server then uses a deep learning-based video generation AI model to generate a series of frames to create a video. The server automatically generates frames to smoothly express the required movement between images.
[0427] 5. Send and review your video
[0428] The generated video is sent to the device and can be viewed by the user, who can then review the video and input any corrections or changes they wish to make.
[0429] 6. Reflecting the changes and regenerating
[0430] The device sends the user-specified corrections to the server, and the server reflects the corrections and regenerates the video. For example, if the user instructs the server to "make the deer walk faster," the correction is reflected.
[0431] 7. Final check and output
[0432] The regenerated video is sent back to the user. The user finally reviews the video to their satisfaction and proceeds to save or share. The device provides the final video file and generates a sharing link if necessary.
[0433] Specific examples
[0434] Specific examples are shown below.
[0435] Example: A user requests a video of a puppy running on a beach at sunset and enters that in the text.
[0436] Initial image: The server generates an initial image of a beach at sunset and a video of a puppy running.
[0437] Example fix: The user inputs the command "Make the puppy faster and add the sound of waves."
[0438] Final output: A video is generated that reflects the edits, and the user is happy to save and share the video.
[0439] As described above, the present invention enables users to intuitively and easily create high-quality moving images by automatically generating, modifying, and regenerating moving images based on user input.
[0440] The processing flow will be explained below.
[0441] Step 1:
[0442] The user opens a dedicated application or web interface, enters the desired video content in text or draws a rough sketch, and is ready to send their specific request to the device.
[0443] Step 2:
[0444] The device checks the user's input data and sends it to the server, including text information and sketches entered by the user, as well as optional settings.
[0445] Step 3:
[0446] The server receives the input data, analyzes it, and prepares to generate an initial image based on the user's request.
[0447] Step 4:
[0448] The server uses an image generation AI model to generate an initial image based on the user's input data. For example, if the user's text input is "A deer is walking in a deep forest," it generates an initial image that represents that scene.
[0449] Step 5:
[0450] The server prepares the video based on the generated initial images, which includes preprocessing to generate successive frames.
[0451] Step 6:
[0452] The server uses a video generation AI model to generate a series of frames that smoothly represent the required movement between the initial images to create the video. This model uses deep learning techniques, for example, to reproduce natural movements.
[0453] Step 7:
[0454] The server sends the generated video to the terminal, which receives the video and notifies the user to confirm it.
[0455] Step 8:
[0456] The user reviews the received video and decides whether they are satisfied. If the user wishes to make any corrections, they input specific corrections (e.g., color adjustments, changing the speed of the motion, etc.) and send them to the device.
[0457] Step 9:
[0458] The device sends the user's modified data to the server, which receives it and prepares to generate a new video that reflects the modifications.
[0459] Step 10:
[0460] The server regenerates the video based on the corrections, and then regenerates successive frames to complete the entire video.
[0461] Step 11:
[0462] The regenerated video is sent to the device again, and the device receives the video and notifies the user for final confirmation.
[0463] Step 12:
[0464] The user reviews the final video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[0465] Through these steps, users can embody their own images as videos and easily modify and regenerate them without requiring any specialized skills.
[0466] Example 1
[0467] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0468] Conventional video generation systems have struggled to easily and intuitively create high-quality videos based on user input data. In particular, it is technically complex to generate videos with natural, continuous movement based on input data in different formats, such as text and sketches. Furthermore, regenerating videos based on user correction requests requires cumbersome procedures and takes a long time.
[0469] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0470] In this invention, the server includes a terminal that receives user input data, a means for transmitting the user input data to the server, a means for utilizing an image generation AI model that generates an initial image based on the user input data, a means for utilizing a video generation AI model that generates successive frames based on the initial image, a means for transmitting the generated video to the terminal, a terminal that receives corrections specified by the user, a means for regenerating the video based on the corrections, and a means for transmitting the regenerated video to the terminal. This allows users to easily and intuitively create high-quality videos and quickly respond to correction requests.
[0471] "User" refers to an individual or organization that uses this system to generate moving images.
[0472] "Input Data" refers to information required for video generation that is provided by a user through a dedicated application or web interface, and may include formats such as text or sketches.
[0473] "Terminal" means a device used by a user to provide input data and view the generated video, including a smartphone, tablet, computer, etc.
[0474] The "server" is a central processing unit for generating moving images based on data input by the user, and performs the functions of data analysis, image generation, moving image generation, and correction reflection.
[0475] An "image generation AI model" is an artificial intelligence model that uses neural network technology to generate initial images based on user input data.
[0476] A "video generation AI model" is an artificial intelligence model that generates successive frames based on initial images to create videos that include natural movements.
[0477] The "corrections" are specific instructions that the user wishes to add or change to the generated video.
[0478] "Regeneration" is the process of regenerating a video by reflecting the corrections specified by the user.
[0479] MODE FOR CARRYING OUT THE INVENTION
[0480] This invention is a system that allows users to easily embody their own ideas and images as animations. This system consists of a terminal that receives input data from the user and a server that processes the data based on that input.
[0481] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if a user wants a video of a deer walking in a deep forest, they can write that in text or enter a simple sketch. The device receives this input data and sends it to the server.
[0482] The server receives the user's input data and generates an initial image using an image generation AI model. This model uses the latest neural network technology, such as Stable Diffusion. The generated initial image will represent the scene the user input, such as "a deer walking in a deep forest."
[0483] The server then uses a deep learning-based video generation AI model to generate a series of frames to create a video, using video generation technologies such as DALL-E 2. The deep learning model generates frames based on the initial images to create a natural-looking sequence of motion.
[0484] The generated video is sent to the device and can be viewed by the user. The user watches the video and inputs any desired corrections or changes in text. These corrections are then sent back to the server via the device.
[0485] The server receives the corrections specified by the user and regenerates the video using the image generation AI model and video generation AI model. For example, if the user requests that the deer walk faster, the server adjusts the video generation parameters according to the request and regenerates the video.
[0486] The regenerated video is sent back to the device, and this process is repeated until the user is satisfied. Finally, when a video that the user is satisfied with is generated, the device provides the video file and proceeds to the storage or sharing stage. For example, a sharing link for the video can be generated using a storage service or file sharing service.
[0487] Prompt Sentence Examples
[0488] Below is an example of a prompt sentence to input to the generative AI model.
[0489] Prompt: "Create a video of a deer walking through a deep forest."
[0490] Prompt: "Draw a scene of a puppy running on a beach at sunset."
[0491] In this way, the present invention allows users to intuitively and easily create high-quality videos by automatically generating, modifying, and regenerating videos based on user input.
[0492] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0493] Step 1:
[0494] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if they want a video of a deer walking through a deep forest, they can write a specific prompt in text or upload a simple sketch. This input data is sent to the device.
[0495] Input: User input data via text or sketch
[0496] Output: Input data stored on the device
[0497] Step 2:
[0498] The device sends the user's input data to the server using the secure HTTP / HTTPS protocol. The server analyzes the received data and prepares for initial image generation.
[0499] Input: Input data stored on the device
[0500] Output: The input data sent to the server
[0501] Step 3:
[0502] The server invokes an image generation AI model to generate an initial image based on the user's input data. This model uses the latest neural network technology, Stable Diffusion. The server inputs a prompt statement into the model and generates an initial image that embodies the specified scene.
[0503] Input: User input data, prompt text
[0504] Output: The initial image generated
[0505] Step 4:
[0506] The server generates a series of frames based on the initial image to create a video. This task uses DALL-E 2, a deep learning-based video generation AI model. The server inputs the initial image into the AI model and generates a series of frames that naturally express continuous movement to create a video.
[0507] Input: Initial image
[0508] Output: Video with consecutive frames
[0509] Step 5:
[0510] The generated video is sent to the device and can be viewed by the user, who can watch the video and enter any desired corrections or changes in text.
[0511] Input: Generated video
[0512] Output: User-supplied text of corrections and changes
[0513] Step 6:
[0514] The device sends the user-specified corrections to the server. The server analyzes the corrections and regenerates the video using the image generation AI model and video generation AI model. For example, if the user instructs the device to "make the deer walk faster," the server reflects that parameter.
[0515] Input: Text of user corrections or changes
[0516] Output: Regenerated video
[0517] Step 7:
[0518] The regenerated video is sent to the device, and the user repeats this process until they are satisfied. Finally, the device can also generate a sharing link in conjunction with a storage or file sharing service to save or share the final video.
[0519] Input: Regenerated video
[0520] Output: Final video file, share link (optional)
[0521] The above is the specific processing flow of the program of this system.
[0522] (Application example 1)
[0523] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0524] Currently, creating advertisements on the market requires a high level of expertise, time, and resources. Small businesses and individual marketers in particular face challenges in creating advertising videos quickly and effectively. There is a need for a way for users to easily realize their ideas and create, edit, and share videos suitable for advertising media.
[0525] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0526] In this invention, the server includes means for receiving user input data, means for generating an initial image based on the user input data, video generation means for generating successive frames based on the initial image, means for transmitting the generated video to the user, means for receiving corrections specified by the user, means for regenerating the video based on the corrections, and means for saving and sharing the user-generated video on an advertising medium, thereby enabling users to easily create high-quality advertising videos and quickly edit and share them.
[0527] "Means for receiving user input data" refers to a device or interface through which a user inputs their ideas or desired images in the form of text, sketches, or the like.
[0528] "Means for generating an initial image" refers to the algorithm or software that creates the initial still image based on user input data.
[0529] "Video generation means for generating successive frames" refers to a technology for creating frames that change continuously from a generated initial image and outputting them as a video.
[0530] "Means for transmitting the generated video to the user" refers to a protocol or system for transferring the generated video file to the user terminal via a network.
[0531] The "means for receiving corrections specified by the user" refers to an interface that allows the user to instruct changes or corrections to the generated video, and a system that receives the content of those changes or corrections.
[0532] "Means for regenerating a video based on the modifications" refers to algorithms or software for regenerating a video that reflects the modifications specified by the user.
[0533] "Means for storing and sharing on advertising media" refers to systems and technologies for storing the generated videos on advertising platforms such as websites, social media, and email, and sharing them with other users.
[0534] This invention provides a system that allows users to easily generate high-quality advertising videos and modify and share them as needed. This system includes means for receiving user input data, means for generating an initial image, means for generating a video that generates a series of frames, means for sending the generated video to the user, means for receiving modifications specified by the user, means for regenerating the video based on the modifications, and means for saving and sharing the generated video on an advertising medium.
[0535] Hardware and software used
[0536] Hardware: The device (smartphone, tablet, smart glasses, etc.) and server through which the user manipulates input data.
[0537] Software: Generative AI models (e.g., Stable Diffusion and DALL-E) and deep learning-based video generation models (e.g., DeepMind and OpenAI's Video GPT) are implemented on cloud servers.
[0538] Processing flow and data calculation
[0539] 1. User Input
[0540] Users input their ideas for the ad video they want to generate in text or sketch form. For example, they can input a prompt in text such as "A new smartphone flying through the air."
[0541] 2. Data transmission and initial image generation
[0542] The device sends input data to a cloud server, which uses a generative AI model to generate an initial image based on the input data, such as a scene of a new smartphone flying through the air.
[0543] 3. Video Generation
[0544] An initial image is input to a deep learning-based video generation model, which then generates successive frames based on the initial image, generating a continuous video from the initial image.
[0545] 4. Sending videos and receiving corrections
[0546] The generated video is sent to the device for the user to review. An interface is provided for the user to input desired corrections or changes, such as "increase the smartphone's flight speed and change the background to a sunset."
[0547] 5. Modify and Regenerate
[0548] The server then reflects the received corrections and regenerates the video. The corrections are made automatically by the AI model, significantly reducing the user's workload.
[0549] 6. Final confirmation, saving and sharing
[0550] The regenerated video is sent to the device for final confirmation by the user. If the user is satisfied, the video is saved and shared on advertising media (website, social media, email, etc.).
[0551] Specific examples
[0552] Prompt Sentence Examples
[0553] "I want to create a video of my new smartphone flying through the air."
[0554] Correction prompt: "Make my phone fly faster and change the background to a sunset."
[0555] This system allows users to easily realize their ideas and quickly create, edit, and share high-quality advertising videos.
[0556] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0557] Step 1:
[0558] Getting user input data
[0559] - Input: The user inputs the desired video content in the form of text or a sketch. For example, the user inputs "a video of a new smartphone flying in the air."
[0560] - Action: The user provides input data using a dedicated application or a web interface.
[0561] - Output: The input data is stored inside the application and is ready to be sent to the cloud server.
[0562] Step 2:
[0563] Generate the initial image
[0564] - Input: User input data sent to the cloud server.
[0565] - Operation: The server passes the received input data to a generative AI model, which generates an initial image from the text or sketch. The generative AI model uses, for example, Stable Diffusion or DALL-E.
[0566] - Output: The generative AI model generates an initial image, for example, an image of a new smartphone flying through the air.
[0567] Step 3:
[0568] Video frame generation
[0569] - Input: The generated initial image.
[0570] Operation: The server inputs the initial image into a deep learning-based video generation model to generate successive frames. The deep learning model used is, for example, Video GPT.
[0571] - Output: A series of frames are generated from the initial image and compiled into a video.
[0572] Step 4:
[0573] Submitting and reviewing videos
[0574] - Input: The generated video file.
[0575] - Operation: The server sends the generated video file to the user's device. The user checks the received video and specifies any corrections.
[0576] - Output: The video sent to the user's device and the modifications specified by the user (e.g., "increase the smartphone's flying speed and change the background to a sunset") are obtained.
[0577] Step 5:
[0578] Applying the changes and regenerating
[0579] - Input: User-specified corrections.
[0580] - Operation: The server uses the AI model and video generation model to modify the video based on the user's modifications. Specifically, it reflects instructions such as flight speed and background changes.
[0581] - Output: A new video file is generated with the modifications reflected.
[0582] Step 6:
[0583] Final confirmation, saving and sharing
[0584] - Input: The modified video file.
[0585] - Operation: The server resends the modified video file to the user's device. The user performs a final check and saves or shares the video on the advertising medium if satisfied.
[0586] - Output: The final video file is saved on the user's device, and a sharing link is generated to the specified advertising medium.
[0587] By following these steps, users can easily and quickly create, modify, and share high-quality advertising videos.
[0588] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0589] The embodiment of the present invention will now be described in detail.
[0590] System Structure
[0591] This invention provides a system that allows users to easily embody their ideas and images as videos. The system consists of a terminal that receives user input data and a server that processes the data. Furthermore, by incorporating an emotion engine that recognizes the user's emotions, it becomes possible to generate videos based on the user's emotions.
[0592] Program processing flow
[0593] 1. User Input
[0594] Using a dedicated application or web interface, users input the desired video content using text or sketches. For example, if a user wants to add the emotion "calm" to a video of a deer walking in a deep forest, they input that information. The input data is then sent to the device.
[0595] 2. Sending input data and emotion data
[0596] The device sends the user's input data and emotion data together to the server. The transmitted data includes the text information and sketches entered by the user, as well as emotion tags.
[0597] 3. Emotion Recognition and Initial Image Generation
[0598] The server receives the input data and emotion data. The server uses an emotion engine to analyze and identify the user's emotion. The recognized emotion is reflected in the initial image generation means. For example, if the user's emotion is "calm," an initial image with a calm atmosphere is generated to reflect this.
[0599] 4. Video Generation
[0600] After the initial image is generated, the server uses a video generation AI model to generate a series of frames to create a video. The model also incorporates the user's emotional data to generate videos that are more emotionally relevant. For example, a calm scene would use gentle movements and soft lighting.
[0601] 5. Send and review your video
[0602] The generated video is sent to the device and can be viewed by the user. The user can review the video and input corrections or changes as necessary.
[0603] 6. Reflecting and regenerating emotionally based modifications
[0604] When the user inputs corrections, the device sends the correction data to the server. The server then generates a new video based on the correction data and emotion data. The corrections are optimized based on the emotion data.
[0605] 7. Final check and output
[0606] The regenerated video is sent back to the device. The user finally reviews the video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[0607] Specific examples
[0608] Specific examples are shown below.
[0609] Example: A user inputs a video of a puppy running on a beach at sunset and adds the emotion "happiness."
[0610] Initial image: The server generates an initial image of a beach at sunset, and the emotion engine creates a scene with bright colors and a feeling of happiness.
[0611] Video Generation: A video of a puppy running on the beach is generated, with lively and happy movements.
[0612] Example fix: The user suggests making fixes like "Make the puppy a little faster and add the sound of the ocean."
[0613] Final output: A video is generated that reflects the edits and the user is happy to save and share the video.
[0614] As described above, the present invention is a system that incorporates a user's emotional data and generates videos that match the user's emotions, allowing the user to easily generate more specific and emotionally rich videos.
[0615] The processing flow will be explained below.
[0616] Step 1:
[0617] The user opens a dedicated application or web interface, enters the desired video content in text or draws a simple image using a sketch, and also enters the emotion they want the video to reflect (e.g., "calm" or "happy"). This prepares the device to send the specific request and emotional data.
[0618] Step 2:
[0619] The device checks the user's input data and emotion data and transmits them to the server. The transmitted data includes the user's input text information, sketches, and emotion tags.
[0620] Step 3:
[0621] The server receives the input data and emotion data, analyzes the received data, and prepares to generate an initial image based on the user's request and emotion.
[0622] Step 4:
[0623] The server uses an emotion engine to analyze the emotions contained in the user's text data and sketches. For example, if the user's emotion is analyzed as "calm," it will recognize it.
[0624] Step 5:
[0625] The server uses an image generation AI model to generate an initial image based on the user's input data and emotion data. For example, based on the user's input text "A deer walking in a deep forest" and the emotion "Calm," it generates an initial image with a calm atmosphere.
[0626] Step 6:
[0627] The server uses a video generation AI model based on the generated initial images to generate a series of frames to create a video. This model also reflects the user's emotional data, generating videos that are in line with their emotions. For example, gentle movements and soft lighting are used for calm scenes.
[0628] Step 7:
[0629] The server sends the generated video to the terminal, which receives the video and notifies the user to confirm it.
[0630] Step 8:
[0631] The user reviews the received video and decides whether they are satisfied. If the user wishes to make any corrections, they input specific corrections (e.g., changing the speed of the puppy's movements, adjusting the color) and send them to the device.
[0632] Step 9:
[0633] The device sends the user's modified data to the server, which receives it and prepares to generate a new video that reflects the modifications.
[0634] Step 10:
[0635] The server regenerates the video based on the modifications. Based on the modifications and emotion data, it regenerates consecutive frames to complete the entire video. For example, if a user requests the puppy to move a little faster, a video reflecting the speed change is generated.
[0636] Step 11:
[0637] The regenerated video is sent to the device again, and the device receives the video and notifies the user for final confirmation.
[0638] Step 12:
[0639] The user reviews the final video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[0640] Through these steps, users can embody their own images into emotive videos without requiring any specialized skills, and can easily modify and regenerate them.
[0641] Example 2
[0642] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0643] Currently, many video generation systems generate videos based on user input data, but lack the technology to adjust the content and atmosphere of the video based on the user's emotions. This makes it difficult to easily generate videos that embody the user's intended emotions and atmosphere. Furthermore, there is no way to reflect emotions when modifying generated videos, making it difficult to obtain results that satisfy the user.
[0644] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0645] In this invention, the server includes means for generating an initial image based on user input data and emotion data, moving image generation means for generating successive frames based on the initial image, and means for transmitting the generated moving image to the user. This makes it possible to generate a moving image that reflects the emotion data input by the user, and to regenerate the moving image while taking emotion into consideration even in subsequent modifications.
[0646] "User input data" refers to information about text, sketches, and video content entered by a user using a dedicated application or web interface.
[0647] "Emotion data" is information that represents the emotion or atmosphere that the user wants to impart to the content of the video.
[0648] An "initial image" is the first still image of a video generated based on the user's input data and emotion data.
[0649] "Video generation means" refers to the function or technology for generating a series of frames based on initial images and emotion data to create a video.
[0650] "Correction points" refer to requests for changes or corrections that a user specifies for a generated video.
[0651] "Regeneration" is the process of recreating a video based on user-specified modifications and emotion data.
[0652] An "emotion engine" is a software module or algorithm for analyzing and recognizing a user's emotional data.
[0653] A "deep learning model" is an artificial intelligence technology that uses large amounts of data to learn and generate sophisticated videos from the input data.
[0654] The present invention provides a system for generating videos that reflect a user's emotions by incorporating the user's emotional data. This system allows users to easily embody their own ideas and images as videos.
[0655] System Structure
[0656] The system consists of a terminal that receives user input data and a server that processes the data. Furthermore, by incorporating an emotion engine, it becomes possible to generate videos based on the user's emotions.
[0657] Hardware and software used
[0658] Device: The device on which the user provides input data (e.g., smartphone, tablet, computer)
[0659] Server: A server that processes input data and emotion data and generates videos.
[0660] Emotion engine: A software module that analyzes and recognizes user emotion data.
[0661] Generative AI model: A deep learning model that generates videos based on input data and emotion data
[0662] A concrete example of the video generation flow
[0663] User Input
[0664] The user inputs the desired video content using a dedicated application or web interface. For example, if the user wants to add the emotion "calm" to a video of a deer walking in a deep forest, the user inputs the information in text and selects "calm" as the emotion tag. This input is sent to the device.
[0665] Example prompt sentence:
[0666] Add the emotion "Happiness" to a video of a puppy running on a beach at sunset.
[0667] Sending input data and emotion data
[0668] The device transmits the user's input data and emotion data to the server. The transmitted data includes the text information and sketches entered by the user, as well as emotion tags.
[0669] Emotion Recognition and Early Image Generation
[0670] The server analyzes the received input data and emotion data and recognizes the emotion using the emotion engine. The recognized emotion is reflected in the initial image generation means. For example, if the user adds the emotion "calm," the server generates an initial image with a calm atmosphere.
[0671] Video generation
[0672] The server then uses a generative AI model to generate a series of frames based on the initial images, creating a video. This process also incorporates the user's emotional data, generating videos that are more in line with their emotions. For example, calm scenes will use gentle movements and soft lighting.
[0673] Submitting and reviewing videos
[0674] The generated video is sent to the device and can be viewed by the user, who can review the video and enter corrections or changes as necessary.
[0675] Reflecting and regenerating emotion-based modifications
[0676] When the user inputs corrections, the device sends the correction data to the server. The server then generates a new video based on the correction data and emotion data. The corrections are also optimized based on the emotion data.
[0677] Final check and output
[0678] The regenerated video is then sent back to the device, where the user can finally review the video to their satisfaction and save or share it. The device will provide the final video file and generate a sharing link if necessary.
[0679] As described above, the present invention provides a system that allows users to easily create and edit videos that reflect their own emotions and intentions. This system utilizes emotion data to provide more specific and emotionally rich videos.
[0680] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0681] Step 1:
[0682] Using a dedicated application or web interface, users input the content of the desired video and related emotional data. This input includes text information, sketches, and emotion tags. This input data is sent to the device. Specifically, the user inputs "a video of a deer walking in a deep forest" and the emotion "calm," and the data is sent to the device.
[0683] Input: Text information, sketch, emotion tag
[0684] Output: Input data sent to the terminal
[0685] Step 2:
[0686] The device combines the user's input data and emotion data into a single data packet and sends it to the server. Specifically, it combines the text information "A deer is walking in a deep forest" and the emotion tag "Calm" in JSON format. The device creates a data packet and sends it to the server.
[0687] Input: User input data and emotion data
[0688] Output: Data packet sent to the server
[0689] Step 3:
[0690] The server analyzes the received data packets and uses the emotion engine to recognize and analyze the user's emotions. Specifically, the server analyzes the received data and recognizes the emotion "calm." This emotion data is reflected in the initial image generation means, as it affects subsequent processing.
[0691] Input: Data packet (user input data and emotion data)
[0692] Output: Recognized emotion data
[0693] Step 4:
[0694] The server uses the initial image generation means to generate an initial image based on the user's input data and the recognized emotion data. Specifically, an initial image depicting a forest landscape in gentle colors is created. This initial image is used as the first step in video generation.
[0695] Input: User input data, recognized emotion data
[0696] Output: Initial image
[0697] Step 5:
[0698] The server uses a generative AI model to generate a series of frames based on the generated initial image, creating a video. Specifically, the video generation AI model receives the initial image and emotion data as input and generates a series of frames of a deer in a tranquil scene. This series of frames is then combined to form a complete video.
[0699] Input: Initial image, recognized emotion data
[0700] Output: Video with continuous frames
[0701] Step 6:
[0702] The generated video is sent from the server to the device, where the user can view it. The user plays the video through the application and checks the content. Specifically, the video data from the server arrives at the device, and the user clicks the "play" button to watch the video.
[0703] Input: Generated video
[0704] Output: Video sent to device, user confirmation
[0705] Step 7:
[0706] If a user wishes to make corrections to the video content, they input the corrections and send the correction data from their device to the server. Specifically, the user might input, "Make the deer move a little faster and add the sound of birds singing in the background." This correction data is then sent from the device to the server.
[0707] Input: User modifications
[0708] Output: Corrected data sent to the server
[0709] Step 8:
[0710] The server regenerates the initial image and video based on the received correction data and emotion data. Specifically, the server generates new initial images and consecutive frames, and regenerates a video that reflects the corrections and emotion data.
[0711] Input: Corrected data, recognized emotion data
[0712] Output: Regenerated video
[0713] Step 9:
[0714] The regenerated video is sent from the server to the device, and the user is finally confirmed to be satisfied. The user watches the video again, and if satisfied, can save or share it. Specifically, the server sends the regenerated video to the device, and if the user is satisfied, they click the "Save" or "Share" button.
[0715] Input: Regenerated video
[0716] Output: User final review, save or share
[0717] (Application example 2)
[0718] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0719] Current video generation systems make it difficult for users to easily create videos that reflect their desired emotions. They also lack the functionality to easily share generated videos on online platforms. This makes it difficult to quickly create and share emotionally rich, effective videos, especially in the advertising field.
[0720] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0721] In this invention, the server includes means for receiving user input data, means for generating an initial image based on the user input data, video generation means for generating successive frames based on the initial image, means for transmitting the generated video to the user, means for receiving corrections specified by the user, means for regenerating the video based on the corrections, means for the video generation means to generate a video reflecting the user's emotional data, an emotion engine for analyzing the user's emotional data, and means for generating a link for sharing the generated video on an online platform. This enables users to easily generate effective videos that reflect emotions and quickly share them online.
[0722] "User input data" refers to information such as text or sketches that a user provides to the system.
[0723] "Initial image" refers to the first still image generated based on user input data.
[0724] "Video generator" refers to a mechanism for generating a sequence of video frames based on an initial image.
[0725] "Correction points" refer to requests for changes or corrections that a user specifies for a video.
[0726] "Regeneration" refers to regenerating a video by reflecting the corrections.
[0727] "Emotional data" refers to information about emotions that a user provides to the system.
[0728] An "emotion engine" refers to software or algorithms that analyze the emotional data provided by the user and use that data to generate videos.
[0729] "Link for sharing on online platforms" refers to a URL or sharing link for sharing the generated video with other users on the Internet.
[0730] "Consecutive frames" refers to a number of still images that make up a moving image.
[0731] "Means for transmitting to the user" refers to a communication means for providing the generated video to the user.
[0732] The present invention relates to a system for enabling users to quickly and effectively create emotional videos and easily share them on an online platform, and a specific embodiment of the system will be described below.
[0733] System Configuration
[0734] The system of the present invention is composed of a terminal that receives user input data, a server that processes the data, and a communication means that provides the generated video to the user. It also includes an emotion engine for emotion recognition and an AI model for video generation.
[0735] Hardware and software used
[0736] Terminal: A device that allows users to input text and sketches necessary for video generation. This includes smartphones and PCs.
[0737] Server: Provides the computing resources to analyze user input data and emotion data and generate videos. Cloud-based servers are often used.
[0738] Emotion engine: Software for analyzing emotion data and generating initial images based on it.
[0739] Video generation AI model: A deep learning model that generates a series of frames based on initial images to create emotionally relevant videos.
[0740] Communication means: A network that transmits the generated video data to the terminal and allows the user to view the video.
[0741] Video generation process
[0742] 1. Getting User Input
[0743] Users can use a dedicated application or web interface to input the content of the video and the desired emotion using text or a sketch. For example, they can input "a video introducing the latest running shoes" and the emotion "excitement."
[0744] 2. Data transmission and processing
[0745] The device sends the user's input data and emotion data to the server, which receives the data, analyzes the emotion using an emotion engine, and generates an initial image.
[0746] 3. Generate initial images
[0747] The emotion engine analyzes the user's emotional data and generates an initial image that reflects that emotion. If the emotion is "excitement," vivid colors and dynamic compositions are used.
[0748] 4. Video Generation
[0749] The server uses a video generation AI model based on the generated initial images to generate successive frames, which then creates a video that reflects the emotion.
[0750] 5. Send and review your video
[0751] The generated video is sent from the server to the device, where the user can review it and make corrections if necessary.
[0752] 6. Processing and regenerating fixes
[0753] If the user inputs corrections, the data is sent to the server again, and the server generates a new video based on the corrections and emotion data.
[0754] 7. Video sharing
[0755] The final video is provided as a link that users can use to share the video on online platforms.
[0756] Examples and prompts
[0757] As a concrete example, consider a case where a user wants to generate an advertising video of a family picnic scene in a park with spring flowers blooming with the emotion of "happiness." The user inputs the following prompt sentence:
[0758] Example prompt sentence
[0759] Advertising video of a family picnic scene in a park with spring flowers
[0760] Emotion: euphoria
[0761] Based on this prompt, a video is generated that reflects the user's emotions, showing smiling faces and warm colors. The generated video can be easily shared on the user's preferred platform.
[0762] The above is a specific embodiment for carrying out the present invention. This system enables advertisers to quickly create emotionally rich and effective videos and use them in their marketing activities.
[0763] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0764] Step 1:
[0765] Users use a dedicated application or web interface to input text, sketches, and emotional information required for video generation. Input includes "a video introducing the latest running shoes" and the emotion "excitement." The device collects this input data and stores it as part of its data.
[0766] Step 2:
[0767] The device sends the user's input data (text, sketches, emotion information) to the server, which receives it and prepares each data for analysis. The input data is organized for further processing.
[0768] Step 3:
[0769] The server uses an emotion engine to analyze the user's emotional information. It generates an initial image based on the results of this analysis. If the input data is "excitement," an initial image with bright colors and a dynamic composition is generated. This initial image becomes the basis for the next video generation.
[0770] Step 4:
[0771] The server supplies the generated initial images to a video generation AI model, which generates successive frames. The generative AI model generates successive frames based on the initial images and emotion data received as input. The output is video data that reflects the emotion.
[0772] Step 5:
[0773] The server sends the generated video to the terminal. The terminal receives the video data and provides it to the user for confirmation. The user can view the video and input corrections as necessary.
[0774] Step 6:
[0775] If the user inputs corrections, the device sends the correction data back to the server, which then receives the corrections and reuses the emotion engine and video generation AI model to generate a new video that reflects the corrections.
[0776] Step 7:
[0777] The server then sends the new video with the modifications reflected back to the device, and the user repeats this process until they are finally satisfied, at which point they are asked to generate a link to save the video or share it on an online platform.
[0778] As described above, users, devices, and servers work together to efficiently generate videos that reflect emotions, and then modify and optimize them as needed, completing the entire process of sharing them online.
[0779] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0780] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0781] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0782] [Third embodiment]
[0783] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0784] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0785] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0786] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0787] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0788] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0789] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0790] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0791] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0792] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0793] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0794] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0795] The embodiment of the present invention will now be described in detail.
[0796] System Structure
[0797] The present invention provides a system that allows users to easily embody their own ideas and images as animations. The system consists of a terminal that receives input data from the user and a server that processes the data based on the input data.
[0798] Program processing flow
[0799] 1. User Input
[0800] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if a user wants a video of a deer walking through a deep forest, they can write that in text or input a simple sketch. This input data is sent to the device.
[0801] 2. Sending input data
[0802] The terminal sends the user's input data to the server, which receives it and prepares to generate an initial image based on the user's wishes.
[0803] 3. Generate initial images
[0804] The server uses an image generation AI model to generate an initial image based on the user's input data. This model uses the latest neural network technology, for example. For example, an initial image representing the user's input scene of "a deer walking in a deep forest" is generated.
[0805] 4. Video Generation
[0806] After the initial images are generated, the server then uses a deep learning-based video generation AI model to generate a series of frames to create a video. The server automatically generates frames to smoothly express the required movement between images.
[0807] 5. Send and review your video
[0808] The generated video is sent to the device and can be viewed by the user, who can then review the video and input any corrections or changes they wish to make.
[0809] 6. Reflecting the changes and regenerating
[0810] The device sends the user-specified corrections to the server, and the server reflects the corrections and regenerates the video. For example, if the user instructs the server to "make the deer walk faster," the correction is reflected.
[0811] 7. Final check and output
[0812] The regenerated video is sent back to the user. The user finally reviews the video to their satisfaction and proceeds to save or share. The device provides the final video file and generates a sharing link if necessary.
[0813] Specific examples
[0814] Specific examples are shown below.
[0815] Example: A user requests a video of a puppy running on a beach at sunset and enters that in the text.
[0816] Initial image: The server generates an initial image of a beach at sunset and a video of a puppy running.
[0817] Example fix: The user inputs the command "Make the puppy faster and add the sound of waves."
[0818] Final output: A video is generated that reflects the edits, and the user is happy to save and share the video.
[0819] As described above, the present invention enables users to intuitively and easily create high-quality moving images by automatically generating, modifying, and regenerating moving images based on user input.
[0820] The processing flow will be explained below.
[0821] Step 1:
[0822] The user opens a dedicated application or web interface, enters the desired video content in text or draws a rough sketch, and is ready to send their specific request to the device.
[0823] Step 2:
[0824] The device checks the user's input data and sends it to the server, including text information and sketches entered by the user, as well as optional settings.
[0825] Step 3:
[0826] The server receives the input data, analyzes it, and prepares to generate an initial image based on the user's request.
[0827] Step 4:
[0828] The server uses an image generation AI model to generate an initial image based on the user's input data. For example, if the user's text input is "A deer is walking in a deep forest," it generates an initial image that represents that scene.
[0829] Step 5:
[0830] The server prepares the video based on the generated initial images, which includes preprocessing to generate successive frames.
[0831] Step 6:
[0832] The server uses a video generation AI model to generate a series of frames that smoothly represent the required movement between the initial images to create the video. This model uses deep learning techniques, for example, to reproduce natural movements.
[0833] Step 7:
[0834] The server sends the generated video to the terminal, which receives the video and notifies the user to confirm it.
[0835] Step 8:
[0836] The user reviews the received video and decides whether they are satisfied. If the user wishes to make any corrections, they input specific corrections (e.g., color adjustments, changing the speed of the motion, etc.) and send them to the device.
[0837] Step 9:
[0838] The device sends the user's modified data to the server, which receives it and prepares to generate a new video that reflects the modifications.
[0839] Step 10:
[0840] The server regenerates the video based on the corrections, and then regenerates successive frames to complete the entire video.
[0841] Step 11:
[0842] The regenerated video is sent to the device again, and the device receives the video and notifies the user for final confirmation.
[0843] Step 12:
[0844] The user reviews the final video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[0845] Through these steps, users can embody their own images as videos and easily modify and regenerate them without requiring any specialized skills.
[0846] Example 1
[0847] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0848] Conventional video generation systems have struggled to easily and intuitively create high-quality videos based on user input data. In particular, it is technically complex to generate videos with natural, continuous movement based on input data in different formats, such as text and sketches. Furthermore, regenerating videos based on user correction requests requires cumbersome procedures and takes a long time.
[0849] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0850] In this invention, the server includes a terminal that receives user input data, a means for transmitting the user input data to the server, a means for utilizing an image generation AI model that generates an initial image based on the user input data, a means for utilizing a video generation AI model that generates successive frames based on the initial image, a means for transmitting the generated video to the terminal, a terminal that receives corrections specified by the user, a means for regenerating the video based on the corrections, and a means for transmitting the regenerated video to the terminal. This allows users to easily and intuitively create high-quality videos and quickly respond to correction requests.
[0851] "User" refers to an individual or organization that uses this system to generate moving images.
[0852] "Input Data" refers to information required for video generation that is provided by a user through a dedicated application or web interface, and may include formats such as text or sketches.
[0853] "Terminal" means a device used by a user to provide input data and view the generated video, including a smartphone, tablet, computer, etc.
[0854] The "server" is a central processing unit for generating moving images based on data input by the user, and performs the functions of data analysis, image generation, moving image generation, and correction reflection.
[0855] An "image generation AI model" is an artificial intelligence model that uses neural network technology to generate initial images based on user input data.
[0856] A "video generation AI model" is an artificial intelligence model that generates successive frames based on initial images to create videos that include natural movements.
[0857] The "corrections" are specific instructions that the user wishes to add or change to the generated video.
[0858] "Regeneration" is the process of regenerating a video by reflecting the corrections specified by the user.
[0859] MODE FOR CARRYING OUT THE INVENTION
[0860] This invention is a system that allows users to easily embody their own ideas and images as animations. This system consists of a terminal that receives input data from the user and a server that processes the data based on that input.
[0861] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if a user wants a video of a deer walking in a deep forest, they can write that in text or enter a simple sketch. The device receives this input data and sends it to the server.
[0862] The server receives the user's input data and generates an initial image using an image generation AI model. This model uses the latest neural network technology, such as Stable Diffusion. The generated initial image will represent the scene the user input, such as "a deer walking in a deep forest."
[0863] The server then uses a deep learning-based video generation AI model to generate a series of frames to create a video, using video generation technologies such as DALL-E 2. The deep learning model generates frames based on the initial images to create a natural-looking sequence of motion.
[0864] The generated video is sent to the device and can be viewed by the user. The user watches the video and inputs any desired corrections or changes in text. These corrections are then sent back to the server via the device.
[0865] The server receives the corrections specified by the user and regenerates the video using the image generation AI model and video generation AI model. For example, if the user requests that the deer walk faster, the server adjusts the video generation parameters according to the request and regenerates the video.
[0866] The regenerated video is sent back to the device, and this process is repeated until the user is satisfied. Finally, when a video that the user is satisfied with is generated, the device provides the video file and proceeds to the storage or sharing stage. For example, a sharing link for the video can be generated using a storage service or file sharing service.
[0867] Prompt Sentence Examples
[0868] Below is an example of a prompt sentence to input to the generative AI model.
[0869] Prompt: "Create a video of a deer walking through a deep forest."
[0870] Prompt: "Draw a scene of a puppy running on a beach at sunset."
[0871] In this way, the present invention allows users to intuitively and easily create high-quality videos by automatically generating, modifying, and regenerating videos based on user input.
[0872] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0873] Step 1:
[0874] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if they want a video of a deer walking through a deep forest, they can write a specific prompt in text or upload a simple sketch. This input data is sent to the device.
[0875] Input: User input data via text or sketch
[0876] Output: Input data stored on the device
[0877] Step 2:
[0878] The device sends the user's input data to the server using the secure HTTP / HTTPS protocol. The server analyzes the received data and prepares for initial image generation.
[0879] Input: Input data stored on the device
[0880] Output: The input data sent to the server
[0881] Step 3:
[0882] The server invokes an image generation AI model to generate an initial image based on the user's input data. This model uses the latest neural network technology, Stable Diffusion. The server inputs a prompt statement into the model and generates an initial image that embodies the specified scene.
[0883] Input: User input data, prompt text
[0884] Output: The initial image generated
[0885] Step 4:
[0886] The server generates a series of frames based on the initial image to create a video. This task uses DALL-E 2, a deep learning-based video generation AI model. The server inputs the initial image into the AI model and generates a series of frames that naturally express continuous movement to create a video.
[0887] Input: Initial image
[0888] Output: Video with consecutive frames
[0889] Step 5:
[0890] The generated video is sent to the device and can be viewed by the user, who can watch the video and enter any desired corrections or changes in text.
[0891] Input: Generated video
[0892] Output: User-supplied text of corrections and changes
[0893] Step 6:
[0894] The device sends the user-specified corrections to the server. The server analyzes the corrections and regenerates the video using the image generation AI model and video generation AI model. For example, if the user instructs the device to "make the deer walk faster," the server reflects that parameter.
[0895] Input: Text of user corrections or changes
[0896] Output: Regenerated video
[0897] Step 7:
[0898] The regenerated video is sent to the device, and the user repeats this process until they are satisfied. Finally, the device can also generate a sharing link in conjunction with a storage or file sharing service to save or share the final video.
[0899] Input: Regenerated video
[0900] Output: Final video file, share link (optional)
[0901] The above is the specific processing flow of the program of this system.
[0902] (Application example 1)
[0903] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0904] Currently, creating advertisements on the market requires a high level of expertise, time, and resources. Small businesses and individual marketers in particular face challenges in creating advertising videos quickly and effectively. There is a need for a way for users to easily realize their ideas and create, edit, and share videos suitable for advertising media.
[0905] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0906] In this invention, the server includes means for receiving user input data, means for generating an initial image based on the user input data, video generation means for generating successive frames based on the initial image, means for transmitting the generated video to the user, means for receiving corrections specified by the user, means for regenerating the video based on the corrections, and means for saving and sharing the user-generated video on an advertising medium, thereby enabling users to easily create high-quality advertising videos and quickly edit and share them.
[0907] "Means for receiving user input data" refers to a device or interface through which a user inputs their ideas or desired images in the form of text, sketches, or the like.
[0908] "Means for generating an initial image" refers to the algorithm or software that creates the initial still image based on user input data.
[0909] "Video generation means for generating successive frames" refers to a technology for creating frames that change continuously from a generated initial image and outputting them as a video.
[0910] "Means for transmitting the generated video to the user" refers to a protocol or system for transferring the generated video file to the user terminal via a network.
[0911] The "means for receiving corrections specified by the user" refers to an interface that allows the user to instruct changes or corrections to the generated video, and a system that receives the content of those changes or corrections.
[0912] "Means for regenerating a video based on the modifications" refers to algorithms or software for regenerating a video that reflects the modifications specified by the user.
[0913] "Means for storing and sharing on advertising media" refers to systems and technologies for storing the generated videos on advertising platforms such as websites, social media, and email, and sharing them with other users.
[0914] This invention provides a system that allows users to easily generate high-quality advertising videos and modify and share them as needed. This system includes means for receiving user input data, means for generating an initial image, means for generating a video that generates a series of frames, means for sending the generated video to the user, means for receiving modifications specified by the user, means for regenerating the video based on the modifications, and means for saving and sharing the generated video on an advertising medium.
[0915] Hardware and software used
[0916] Hardware: The device (smartphone, tablet, smart glasses, etc.) and server through which the user manipulates input data.
[0917] Software: Generative AI models (e.g., Stable Diffusion and DALL-E) and deep learning-based video generation models (e.g., DeepMind and OpenAI's Video GPT) are implemented on cloud servers.
[0918] Processing flow and data calculation
[0919] 1. User Input
[0920] Users input their ideas for the ad video they want to generate in text or sketch form. For example, they can input a prompt in text such as "A new smartphone flying through the air."
[0921] 2. Data transmission and initial image generation
[0922] The device sends input data to a cloud server, which uses a generative AI model to generate an initial image based on the input data, such as a scene of a new smartphone flying through the air.
[0923] 3. Video Generation
[0924] An initial image is input to a deep learning-based video generation model, which then generates successive frames based on the initial image, generating a continuous video from the initial image.
[0925] 4. Sending videos and receiving corrections
[0926] The generated video is sent to the device for the user to review. An interface is provided for the user to input desired corrections or changes, such as "increase the smartphone's flight speed and change the background to a sunset."
[0927] 5. Modify and Regenerate
[0928] The server then reflects the received corrections and regenerates the video. The corrections are made automatically by the AI model, significantly reducing the user's workload.
[0929] 6. Final confirmation, saving and sharing
[0930] The regenerated video is sent to the device for final confirmation by the user. If the user is satisfied, the video is saved and shared on advertising media (website, social media, email, etc.).
[0931] Specific examples
[0932] Prompt Sentence Examples
[0933] "I want to create a video of my new smartphone flying through the air."
[0934] Correction prompt: "Make my phone fly faster and change the background to a sunset."
[0935] This system allows users to easily realize their ideas and quickly create, edit, and share high-quality advertising videos.
[0936] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0937] Step 1:
[0938] Getting user input data
[0939] - Input: The user inputs the desired video content in the form of text or a sketch. For example, the user inputs "a video of a new smartphone flying in the air."
[0940] - Action: The user provides input data using a dedicated application or a web interface.
[0941] - Output: The input data is stored inside the application and is ready to be sent to the cloud server.
[0942] Step 2:
[0943] Generate the initial image
[0944] - Input: User input data sent to the cloud server.
[0945] - Operation: The server passes the received input data to a generative AI model, which generates an initial image from the text or sketch. The generative AI model uses, for example, Stable Diffusion or DALL-E.
[0946] - Output: The generative AI model generates an initial image, for example, an image of a new smartphone flying through the air.
[0947] Step 3:
[0948] Video frame generation
[0949] - Input: The generated initial image.
[0950] Operation: The server inputs the initial image into a deep learning-based video generation model to generate successive frames. The deep learning model used is, for example, Video GPT.
[0951] - Output: A series of frames are generated from the initial image and compiled into a video.
[0952] Step 4:
[0953] Submitting and reviewing videos
[0954] - Input: The generated video file.
[0955] - Operation: The server sends the generated video file to the user's device. The user checks the received video and specifies any corrections.
[0956] - Output: The video sent to the user's device and the modifications specified by the user (e.g., "increase the smartphone's flying speed and change the background to a sunset") are obtained.
[0957] Step 5:
[0958] Applying the changes and regenerating
[0959] - Input: User-specified corrections.
[0960] - Operation: The server uses the AI model and video generation model to modify the video based on the user's modifications. Specifically, it reflects instructions such as flight speed and background changes.
[0961] - Output: A new video file is generated with the modifications reflected.
[0962] Step 6:
[0963] Final confirmation, saving and sharing
[0964] - Input: The modified video file.
[0965] - Operation: The server resends the modified video file to the user's device. The user performs a final check and saves or shares the video on the advertising medium if satisfied.
[0966] - Output: The final video file is saved on the user's device, and a sharing link is generated to the specified advertising medium.
[0967] By following these steps, users can easily and quickly create, modify, and share high-quality advertising videos.
[0968] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0969] The embodiment of the present invention will now be described in detail.
[0970] System Structure
[0971] This invention provides a system that allows users to easily embody their ideas and images as videos. The system consists of a terminal that receives user input data and a server that processes the data. Furthermore, by incorporating an emotion engine that recognizes the user's emotions, it becomes possible to generate videos based on the user's emotions.
[0972] Program processing flow
[0973] 1. User Input
[0974] Using a dedicated application or web interface, users input the desired video content using text or sketches. For example, if a user wants to add the emotion "calm" to a video of a deer walking in a deep forest, they input that information. The input data is then sent to the device.
[0975] 2. Sending input data and emotion data
[0976] The device sends the user's input data and emotion data together to the server. The transmitted data includes the text information and sketches entered by the user, as well as emotion tags.
[0977] 3. Emotion Recognition and Initial Image Generation
[0978] The server receives the input data and emotion data. The server uses an emotion engine to analyze and identify the user's emotion. The recognized emotion is reflected in the initial image generation means. For example, if the user's emotion is "calm," an initial image with a calm atmosphere is generated to reflect this.
[0979] 4. Video Generation
[0980] After the initial image is generated, the server uses a video generation AI model to generate a series of frames to create a video. The model also incorporates the user's emotional data to generate videos that are more emotionally relevant. For example, a calm scene would use gentle movements and soft lighting.
[0981] 5. Send and review your video
[0982] The generated video is sent to the device and can be viewed by the user. The user can review the video and input corrections or changes as necessary.
[0983] 6. Reflecting and regenerating emotionally based modifications
[0984] When the user inputs corrections, the device sends the correction data to the server. The server then generates a new video based on the correction data and emotion data. The corrections are optimized based on the emotion data.
[0985] 7. Final check and output
[0986] The regenerated video is sent back to the device. The user finally reviews the video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[0987] Specific examples
[0988] Specific examples are shown below.
[0989] Example: A user inputs a video of a puppy running on a beach at sunset and adds the emotion "happiness."
[0990] Initial image: The server generates an initial image of a beach at sunset, and the emotion engine creates a scene with bright colors and a feeling of happiness.
[0991] Video Generation: A video of a puppy running on the beach is generated, with lively and happy movements.
[0992] Example fix: The user suggests making fixes like "Make the puppy a little faster and add the sound of the ocean."
[0993] Final output: A video is generated that reflects the edits and the user is happy to save and share the video.
[0994] As described above, the present invention is a system that incorporates a user's emotional data and generates videos that match the user's emotions, allowing the user to easily generate more specific and emotionally rich videos.
[0995] The processing flow will be explained below.
[0996] Step 1:
[0997] The user opens a dedicated application or web interface, enters the desired video content in text or draws a simple image using a sketch, and also enters the emotion they want the video to reflect (e.g., "calm" or "happy"). This prepares the device to send the specific request and emotional data.
[0998] Step 2:
[0999] The device checks the user's input data and emotion data and transmits them to the server. The transmitted data includes the user's input text information, sketches, and emotion tags.
[1000] Step 3:
[1001] The server receives the input data and emotion data, analyzes the received data, and prepares to generate an initial image based on the user's request and emotion.
[1002] Step 4:
[1003] The server uses an emotion engine to analyze the emotions contained in the user's text data and sketches. For example, if the user's emotion is analyzed as "calm," it will recognize it.
[1004] Step 5:
[1005] The server uses an image generation AI model to generate an initial image based on the user's input data and emotion data. For example, based on the user's input text "A deer walking in a deep forest" and the emotion "Calm," it generates an initial image with a calm atmosphere.
[1006] Step 6:
[1007] The server uses a video generation AI model based on the generated initial images to generate a series of frames to create a video. This model also reflects the user's emotional data, generating videos that are in line with their emotions. For example, gentle movements and soft lighting are used for calm scenes.
[1008] Step 7:
[1009] The server sends the generated video to the terminal, which receives the video and notifies the user to confirm it.
[1010] Step 8:
[1011] The user reviews the received video and decides whether they are satisfied. If the user wishes to make any corrections, they input specific corrections (e.g., changing the speed of the puppy's movements, adjusting the color) and send them to the device.
[1012] Step 9:
[1013] The device sends the user's modified data to the server, which receives it and prepares to generate a new video that reflects the modifications.
[1014] Step 10:
[1015] The server regenerates the video based on the modifications. Based on the modifications and emotion data, it regenerates consecutive frames to complete the entire video. For example, if a user requests the puppy to move a little faster, a video reflecting the speed change is generated.
[1016] Step 11:
[1017] The regenerated video is sent to the device again, and the device receives the video and notifies the user for final confirmation.
[1018] Step 12:
[1019] The user reviews the final video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[1020] Through these steps, users can embody their own images into emotive videos without requiring any specialized skills, and can easily modify and regenerate them.
[1021] Example 2
[1022] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1023] Currently, many video generation systems generate videos based on user input data, but lack the technology to adjust the content and atmosphere of the video based on the user's emotions. This makes it difficult to easily generate videos that embody the user's intended emotions and atmosphere. Furthermore, there is no way to reflect emotions when modifying generated videos, making it difficult to obtain results that satisfy the user.
[1024] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1025] In this invention, the server includes means for generating an initial image based on user input data and emotion data, moving image generation means for generating successive frames based on the initial image, and means for transmitting the generated moving image to the user. This makes it possible to generate a moving image that reflects the emotion data input by the user, and to regenerate the moving image while taking emotion into consideration even in subsequent modifications.
[1026] "User input data" refers to information about text, sketches, and video content entered by a user using a dedicated application or web interface.
[1027] "Emotion data" is information that represents the emotion or atmosphere that the user wants to impart to the content of the video.
[1028] An "initial image" is the first still image of a video generated based on the user's input data and emotion data.
[1029] "Video generation means" refers to the function or technology for generating a series of frames based on initial images and emotion data to create a video.
[1030] "Correction points" refer to requests for changes or corrections that a user specifies for a generated video.
[1031] "Regeneration" is the process of recreating a video based on user-specified modifications and emotion data.
[1032] An "emotion engine" is a software module or algorithm for analyzing and recognizing a user's emotional data.
[1033] A "deep learning model" is an artificial intelligence technology that uses large amounts of data to learn and generate sophisticated videos from the input data.
[1034] The present invention provides a system for generating videos that reflect a user's emotions by incorporating the user's emotional data. This system allows users to easily embody their own ideas and images as videos.
[1035] System Structure
[1036] The system consists of a terminal that receives user input data and a server that processes the data. Furthermore, by incorporating an emotion engine, it becomes possible to generate videos based on the user's emotions.
[1037] Hardware and software used
[1038] Device: The device on which the user provides input data (e.g., smartphone, tablet, computer)
[1039] Server: A server that processes input data and emotion data and generates videos.
[1040] Emotion engine: A software module that analyzes and recognizes user emotion data.
[1041] Generative AI model: A deep learning model that generates videos based on input data and emotion data
[1042] A concrete example of the video generation flow
[1043] User Input
[1044] The user inputs the desired video content using a dedicated application or web interface. For example, if the user wants to add the emotion "calm" to a video of a deer walking in a deep forest, the user inputs the information in text and selects "calm" as the emotion tag. This input is sent to the device.
[1045] Example prompt sentence:
[1046] Add the emotion "Happiness" to a video of a puppy running on a beach at sunset.
[1047] Sending input data and emotion data
[1048] The device transmits the user's input data and emotion data to the server. The transmitted data includes the text information and sketches entered by the user, as well as emotion tags.
[1049] Emotion Recognition and Early Image Generation
[1050] The server analyzes the received input data and emotion data and recognizes the emotion using the emotion engine. The recognized emotion is reflected in the initial image generation means. For example, if the user adds the emotion "calm," the server generates an initial image with a calm atmosphere.
[1051] Video generation
[1052] The server then uses a generative AI model to generate a series of frames based on the initial images, creating a video. This process also incorporates the user's emotional data, generating videos that are more in line with their emotions. For example, calm scenes will use gentle movements and soft lighting.
[1053] Submitting and reviewing videos
[1054] The generated video is sent to the device and can be viewed by the user, who can review the video and enter corrections or changes as necessary.
[1055] Reflecting and regenerating emotion-based modifications
[1056] When the user inputs corrections, the device sends the correction data to the server. The server then generates a new video based on the correction data and emotion data. The corrections are also optimized based on the emotion data.
[1057] Final check and output
[1058] The regenerated video is then sent back to the device, where the user can finally review the video to their satisfaction and save or share it. The device will provide the final video file and generate a sharing link if necessary.
[1059] As described above, the present invention provides a system that allows users to easily create and edit videos that reflect their own emotions and intentions. This system utilizes emotion data to provide more specific and emotionally rich videos.
[1060] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1061] Step 1:
[1062] Using a dedicated application or web interface, users input the content of the desired video and related emotional data. This input includes text information, sketches, and emotion tags. This input data is sent to the device. Specifically, the user inputs "a video of a deer walking in a deep forest" and the emotion "calm," and the data is sent to the device.
[1063] Input: Text information, sketch, emotion tag
[1064] Output: Input data sent to the terminal
[1065] Step 2:
[1066] The device combines the user's input data and emotion data into a single data packet and sends it to the server. Specifically, it combines the text information "A deer is walking in a deep forest" and the emotion tag "Calm" in JSON format. The device creates a data packet and sends it to the server.
[1067] Input: User input data and emotion data
[1068] Output: Data packet sent to the server
[1069] Step 3:
[1070] The server analyzes the received data packets and uses the emotion engine to recognize and analyze the user's emotions. Specifically, the server analyzes the received data and recognizes the emotion "calm." This emotion data is reflected in the initial image generation means, as it affects subsequent processing.
[1071] Input: Data packet (user input data and emotion data)
[1072] Output: Recognized emotion data
[1073] Step 4:
[1074] The server uses the initial image generation means to generate an initial image based on the user's input data and the recognized emotion data. Specifically, an initial image depicting a forest landscape in gentle colors is created. This initial image is used as the first step in video generation.
[1075] Input: User input data, recognized emotion data
[1076] Output: Initial image
[1077] Step 5:
[1078] The server uses a generative AI model to generate a series of frames based on the generated initial image, creating a video. Specifically, the video generation AI model receives the initial image and emotion data as input and generates a series of frames of a deer in a tranquil scene. This series of frames is then combined to form a complete video.
[1079] Input: Initial image, recognized emotion data
[1080] Output: Video with continuous frames
[1081] Step 6:
[1082] The generated video is sent from the server to the device, where the user can view it. The user plays the video through the application and checks the content. Specifically, the video data from the server arrives at the device, and the user clicks the "play" button to watch the video.
[1083] Input: Generated video
[1084] Output: Video sent to device, user confirmation
[1085] Step 7:
[1086] If a user wishes to make corrections to the video content, they input the corrections and send the correction data from their device to the server. Specifically, the user might input, "Make the deer move a little faster and add the sound of birds singing in the background." This correction data is then sent from the device to the server.
[1087] Input: User modifications
[1088] Output: Corrected data sent to the server
[1089] Step 8:
[1090] The server regenerates the initial image and video based on the received correction data and emotion data. Specifically, the server generates new initial images and consecutive frames, and regenerates a video that reflects the corrections and emotion data.
[1091] Input: Corrected data, recognized emotion data
[1092] Output: Regenerated video
[1093] Step 9:
[1094] The regenerated video is sent from the server to the device, and the user is finally confirmed to be satisfied. The user watches the video again, and if satisfied, can save or share it. Specifically, the server sends the regenerated video to the device, and if the user is satisfied, they click the "Save" or "Share" button.
[1095] Input: Regenerated video
[1096] Output: User final review, save or share
[1097] (Application example 2)
[1098] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1099] Current video generation systems make it difficult for users to easily create videos that reflect their desired emotions. They also lack the functionality to easily share generated videos on online platforms. This makes it difficult to quickly create and share emotionally rich, effective videos, especially in the advertising field.
[1100] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1101] In this invention, the server includes means for receiving user input data, means for generating an initial image based on the user input data, video generation means for generating successive frames based on the initial image, means for transmitting the generated video to the user, means for receiving corrections specified by the user, means for regenerating the video based on the corrections, means for the video generation means to generate a video reflecting the user's emotional data, an emotion engine for analyzing the user's emotional data, and means for generating a link for sharing the generated video on an online platform. This enables users to easily generate effective videos that reflect emotions and quickly share them online.
[1102] "User input data" refers to information such as text or sketches that a user provides to the system.
[1103] "Initial image" refers to the first still image generated based on user input data.
[1104] "Video generator" refers to a mechanism for generating a sequence of video frames based on an initial image.
[1105] "Correction points" refer to requests for changes or corrections that a user specifies for a video.
[1106] "Regeneration" refers to regenerating a video by reflecting the corrections.
[1107] "Emotional data" refers to information about emotions that a user provides to the system.
[1108] An "emotion engine" refers to software or algorithms that analyze the emotional data provided by the user and use that data to generate videos.
[1109] "Link for sharing on online platforms" refers to a URL or sharing link for sharing the generated video with other users on the Internet.
[1110] "Consecutive frames" refers to a number of still images that make up a moving image.
[1111] "Means for transmitting to the user" refers to a communication means for providing the generated video to the user.
[1112] The present invention relates to a system for enabling users to quickly and effectively create emotional videos and easily share them on an online platform, and a specific embodiment of the system will be described below.
[1113] System Configuration
[1114] The system of the present invention is composed of a terminal that receives user input data, a server that processes the data, and a communication means that provides the generated video to the user. It also includes an emotion engine for emotion recognition and an AI model for video generation.
[1115] Hardware and software used
[1116] Terminal: A device that allows users to input text and sketches necessary for video generation. This includes smartphones and PCs.
[1117] Server: Provides the computing resources to analyze user input data and emotion data and generate videos. Cloud-based servers are often used.
[1118] Emotion engine: Software for analyzing emotion data and generating initial images based on it.
[1119] Video generation AI model: A deep learning model that generates a series of frames based on initial images to create emotionally relevant videos.
[1120] Communication means: A network that transmits the generated video data to the terminal and allows the user to view the video.
[1121] Video generation process
[1122] 1. Getting User Input
[1123] Users can use a dedicated application or web interface to input the content of the video and the desired emotion using text or a sketch. For example, they can input "a video introducing the latest running shoes" and the emotion "excitement."
[1124] 2. Data transmission and processing
[1125] The device sends the user's input data and emotion data to the server, which receives the data, analyzes the emotion using an emotion engine, and generates an initial image.
[1126] 3. Generate initial images
[1127] The emotion engine analyzes the user's emotional data and generates an initial image that reflects that emotion. If the emotion is "excitement," vivid colors and dynamic compositions are used.
[1128] 4. Video Generation
[1129] The server uses a video generation AI model based on the generated initial images to generate successive frames, which then creates a video that reflects the emotion.
[1130] 5. Send and review your video
[1131] The generated video is sent from the server to the device, where the user can review it and make corrections if necessary.
[1132] 6. Processing and regenerating fixes
[1133] If the user inputs corrections, the data is sent to the server again, and the server generates a new video based on the corrections and emotion data.
[1134] 7. Video sharing
[1135] The final video is provided as a link that users can use to share the video on online platforms.
[1136] Examples and prompts
[1137] As a concrete example, consider a case where a user wants to generate an advertising video of a family picnic scene in a park with spring flowers blooming with the emotion of "happiness." The user inputs the following prompt sentence:
[1138] Example prompt sentence
[1139] Advertising video of a family picnic scene in a park with spring flowers
[1140] Emotion: euphoria
[1141] Based on this prompt, a video is generated that reflects the user's emotions, showing smiling faces and warm colors. The generated video can be easily shared on the user's preferred platform.
[1142] The above is a specific embodiment for carrying out the present invention. This system enables advertisers to quickly create emotionally rich and effective videos and use them in their marketing activities.
[1143] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1144] Step 1:
[1145] Users use a dedicated application or web interface to input text, sketches, and emotional information required for video generation. Input includes "a video introducing the latest running shoes" and the emotion "excitement." The device collects this input data and stores it as part of its data.
[1146] Step 2:
[1147] The device sends the user's input data (text, sketches, emotion information) to the server, which receives it and prepares each data for analysis. The input data is organized for further processing.
[1148] Step 3:
[1149] The server uses an emotion engine to analyze the user's emotional information. It generates an initial image based on the results of this analysis. If the input data is "excitement," an initial image with bright colors and a dynamic composition is generated. This initial image becomes the basis for the next video generation.
[1150] Step 4:
[1151] The server supplies the generated initial images to a video generation AI model, which generates successive frames. The generative AI model generates successive frames based on the initial images and emotion data received as input. The output is video data that reflects the emotion.
[1152] Step 5:
[1153] The server sends the generated video to the terminal. The terminal receives the video data and provides it to the user for confirmation. The user can view the video and input corrections as necessary.
[1154] Step 6:
[1155] If the user inputs corrections, the device sends the correction data back to the server, which then receives the corrections and reuses the emotion engine and video generation AI model to generate a new video that reflects the corrections.
[1156] Step 7:
[1157] The server then sends the new video with the modifications reflected back to the device, and the user repeats this process until they are finally satisfied, at which point they are asked to generate a link to save the video or share it on an online platform.
[1158] As described above, users, devices, and servers work together to efficiently generate videos that reflect emotions, and then modify and optimize them as needed, completing the entire process of sharing them online.
[1159] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1160] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1161] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1162] [Fourth embodiment]
[1163] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1164] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1165] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1166] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1167] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1168] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1169] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1170] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1171] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1172] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1173] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1174] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1175] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1176] The embodiment of the present invention will now be described in detail.
[1177] System Structure
[1178] The present invention provides a system that allows users to easily embody their own ideas and images as animations. The system consists of a terminal that receives input data from the user and a server that processes the data based on the input data.
[1179] Program processing flow
[1180] 1. User Input
[1181] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if a user wants a video of a deer walking through a deep forest, they can write that in text or input a simple sketch. This input data is sent to the device.
[1182] 2. Sending input data
[1183] The terminal sends the user's input data to the server, which receives it and prepares to generate an initial image based on the user's wishes.
[1184] 3. Generate initial images
[1185] The server uses an image generation AI model to generate an initial image based on the user's input data. This model uses the latest neural network technology, for example. For example, an initial image representing the user's input scene of "a deer walking in a deep forest" is generated.
[1186] 4. Video Generation
[1187] After the initial images are generated, the server then uses a deep learning-based video generation AI model to generate a series of frames to create a video. The server automatically generates frames to smoothly express the required movement between images.
[1188] 5. Send and review your video
[1189] The generated video is sent to the device and can be viewed by the user, who can then review the video and input any corrections or changes they wish to make.
[1190] 6. Reflecting the changes and regenerating
[1191] The device sends the user-specified corrections to the server, and the server reflects the corrections and regenerates the video. For example, if the user instructs the server to "make the deer walk faster," the correction is reflected.
[1192] 7. Final check and output
[1193] The regenerated video is sent back to the user. The user finally reviews the video to their satisfaction and proceeds to save or share. The device provides the final video file and generates a sharing link if necessary.
[1194] Specific examples
[1195] Specific examples are shown below.
[1196] Example: A user requests a video of a puppy running on a beach at sunset and enters that in the text.
[1197] Initial image: The server generates an initial image of a beach at sunset and a video of a puppy running.
[1198] Example fix: The user inputs the command "Make the puppy faster and add the sound of waves."
[1199] Final output: A video is generated that reflects the edits, and the user is happy to save and share the video.
[1200] As described above, the present invention enables users to intuitively and easily create high-quality moving images by automatically generating, modifying, and regenerating moving images based on user input.
[1201] The processing flow will be explained below.
[1202] Step 1:
[1203] The user opens a dedicated application or web interface, enters the desired video content in text or draws a rough sketch, and is ready to send their specific request to the device.
[1204] Step 2:
[1205] The device checks the user's input data and sends it to the server, including text information and sketches entered by the user, as well as optional settings.
[1206] Step 3:
[1207] The server receives the input data, analyzes it, and prepares to generate an initial image based on the user's request.
[1208] Step 4:
[1209] The server uses an image generation AI model to generate an initial image based on the user's input data. For example, if the user's text input is "A deer is walking in a deep forest," it generates an initial image that represents that scene.
[1210] Step 5:
[1211] The server prepares the video based on the generated initial images, which includes preprocessing to generate successive frames.
[1212] Step 6:
[1213] The server uses a video generation AI model to generate a series of frames that smoothly represent the required movement between the initial images to create the video. This model uses deep learning techniques, for example, to reproduce natural movements.
[1214] Step 7:
[1215] The server sends the generated video to the terminal, which receives the video and notifies the user to confirm it.
[1216] Step 8:
[1217] The user reviews the received video and decides whether they are satisfied. If the user wishes to make any corrections, they input specific corrections (e.g., color adjustments, changing the speed of the motion, etc.) and send them to the device.
[1218] Step 9:
[1219] The device sends the user's modified data to the server, which receives it and prepares to generate a new video that reflects the modifications.
[1220] Step 10:
[1221] The server regenerates the video based on the corrections, and then regenerates successive frames to complete the entire video.
[1222] Step 11:
[1223] The regenerated video is sent to the device again, and the device receives the video and notifies the user for final confirmation.
[1224] Step 12:
[1225] The user reviews the final video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[1226] Through these steps, users can embody their own images as videos and easily modify and regenerate them without requiring any specialized skills.
[1227] Example 1
[1228] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1229] Conventional video generation systems have struggled to easily and intuitively create high-quality videos based on user input data. In particular, it is technically complex to generate videos with natural, continuous movement based on input data in different formats, such as text and sketches. Furthermore, regenerating videos based on user correction requests requires cumbersome procedures and takes a long time.
[1230] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1231] In this invention, the server includes a terminal that receives user input data, a means for transmitting the user input data to the server, a means for utilizing an image generation AI model that generates an initial image based on the user input data, a means for utilizing a video generation AI model that generates successive frames based on the initial image, a means for transmitting the generated video to the terminal, a terminal that receives corrections specified by the user, a means for regenerating the video based on the corrections, and a means for transmitting the regenerated video to the terminal. This allows users to easily and intuitively create high-quality videos and quickly respond to correction requests.
[1232] "User" refers to an individual or organization that uses this system to generate moving images.
[1233] "Input Data" refers to information required for video generation that is provided by a user through a dedicated application or web interface, and may include formats such as text or sketches.
[1234] "Terminal" means a device used by a user to provide input data and view the generated video, including a smartphone, tablet, computer, etc.
[1235] The "server" is a central processing unit for generating moving images based on data input by the user, and performs the functions of data analysis, image generation, moving image generation, and correction reflection.
[1236] An "image generation AI model" is an artificial intelligence model that uses neural network technology to generate initial images based on user input data.
[1237] A "video generation AI model" is an artificial intelligence model that generates successive frames based on initial images to create videos that include natural movements.
[1238] The "corrections" are specific instructions that the user wishes to add or change to the generated video.
[1239] "Regeneration" is the process of regenerating a video by reflecting the corrections specified by the user.
[1240] MODE FOR CARRYING OUT THE INVENTION
[1241] This invention is a system that allows users to easily embody their own ideas and images as animations. This system consists of a terminal that receives input data from the user and a server that processes the data based on that input.
[1242] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if a user wants a video of a deer walking in a deep forest, they can write that in text or enter a simple sketch. The device receives this input data and sends it to the server.
[1243] The server receives the user's input data and generates an initial image using an image generation AI model. This model uses the latest neural network technology, such as Stable Diffusion. The generated initial image will represent the scene the user input, such as "a deer walking in a deep forest."
[1244] The server then uses a deep learning-based video generation AI model to generate a series of frames to create a video, using video generation technologies such as DALL-E 2. The deep learning model generates frames based on the initial images to create a natural-looking sequence of motion.
[1245] The generated video is sent to the device and can be viewed by the user. The user watches the video and inputs any desired corrections or changes in text. These corrections are then sent back to the server via the device.
[1246] The server receives the corrections specified by the user and regenerates the video using the image generation AI model and video generation AI model. For example, if the user requests that the deer walk faster, the server adjusts the video generation parameters according to the request and regenerates the video.
[1247] The regenerated video is sent back to the device, and this process is repeated until the user is satisfied. Finally, when a video that the user is satisfied with is generated, the device provides the video file and proceeds to the storage or sharing stage. For example, a sharing link for the video can be generated using a storage service or file sharing service.
[1248] Prompt Sentence Examples
[1249] Below is an example of a prompt sentence to input to the generative AI model.
[1250] Prompt: "Create a video of a deer walking through a deep forest."
[1251] Prompt: "Draw a scene of a puppy running on a beach at sunset."
[1252] In this way, the present invention allows users to intuitively and easily create high-quality videos by automatically generating, modifying, and regenerating videos based on user input.
[1253] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1254] Step 1:
[1255] Users use a dedicated application or web interface to input the desired image using text or a sketch. For example, if they want a video of a deer walking through a deep forest, they can write a specific prompt in text or upload a simple sketch. This input data is sent to the device.
[1256] Input: User input data via text or sketch
[1257] Output: Input data stored on the device
[1258] Step 2:
[1259] The device sends the user's input data to the server using the secure HTTP / HTTPS protocol. The server analyzes the received data and prepares for initial image generation.
[1260] Input: Input data stored on the device
[1261] Output: The input data sent to the server
[1262] Step 3:
[1263] The server invokes an image generation AI model to generate an initial image based on the user's input data. This model uses the latest neural network technology, Stable Diffusion. The server inputs a prompt statement into the model and generates an initial image that embodies the specified scene.
[1264] Input: User input data, prompt text
[1265] Output: The initial image generated
[1266] Step 4:
[1267] The server generates a series of frames based on the initial image to create a video. This task uses DALL-E 2, a deep learning-based video generation AI model. The server inputs the initial image into the AI model and generates a series of frames that naturally express continuous movement to create a video.
[1268] Input: Initial image
[1269] Output: Video with consecutive frames
[1270] Step 5:
[1271] The generated video is sent to the device and can be viewed by the user, who can watch the video and enter any desired corrections or changes in text.
[1272] Input: Generated video
[1273] Output: User-supplied text of corrections and changes
[1274] Step 6:
[1275] The device sends the user-specified corrections to the server. The server analyzes the corrections and regenerates the video using the image generation AI model and video generation AI model. For example, if the user instructs the device to "make the deer walk faster," the server reflects that parameter.
[1276] Input: Text of user corrections or changes
[1277] Output: Regenerated video
[1278] Step 7:
[1279] The regenerated video is sent to the device, and the user repeats this process until they are satisfied. Finally, the device can also generate a sharing link in conjunction with a storage or file sharing service to save or share the final video.
[1280] Input: Regenerated video
[1281] Output: Final video file, share link (optional)
[1282] The above is the specific processing flow of the program of this system.
[1283] (Application example 1)
[1284] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1285] Currently, creating advertisements on the market requires a high level of expertise, time, and resources. Small businesses and individual marketers in particular face challenges in creating advertising videos quickly and effectively. There is a need for a way for users to easily realize their ideas and create, edit, and share videos suitable for advertising media.
[1286] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1287] In this invention, the server includes means for receiving user input data, means for generating an initial image based on the user input data, video generation means for generating successive frames based on the initial image, means for transmitting the generated video to the user, means for receiving corrections specified by the user, means for regenerating the video based on the corrections, and means for saving and sharing the user-generated video on an advertising medium, thereby enabling users to easily create high-quality advertising videos and quickly edit and share them.
[1288] "Means for receiving user input data" refers to a device or interface through which a user inputs their ideas or desired images in the form of text, sketches, or the like.
[1289] "Means for generating an initial image" refers to the algorithm or software that creates the initial still image based on user input data.
[1290] "Video generation means for generating successive frames" refers to a technology for creating frames that change continuously from a generated initial image and outputting them as a video.
[1291] "Means for transmitting the generated video to the user" refers to a protocol or system for transferring the generated video file to the user terminal via a network.
[1292] The "means for receiving corrections specified by the user" refers to an interface that allows the user to instruct changes or corrections to the generated video, and a system that receives the content of those changes or corrections.
[1293] "Means for regenerating a video based on the modifications" refers to algorithms or software for regenerating a video that reflects the modifications specified by the user.
[1294] "Means for storing and sharing on advertising media" refers to systems and technologies for storing the generated videos on advertising platforms such as websites, social media, and email, and sharing them with other users.
[1295] This invention provides a system that allows users to easily generate high-quality advertising videos and modify and share them as needed. This system includes means for receiving user input data, means for generating an initial image, means for generating a video that generates a series of frames, means for sending the generated video to the user, means for receiving modifications specified by the user, means for regenerating the video based on the modifications, and means for saving and sharing the generated video on an advertising medium.
[1296] Hardware and software used
[1297] Hardware: The device (smartphone, tablet, smart glasses, etc.) and server through which the user manipulates input data.
[1298] Software: Generative AI models (e.g., Stable Diffusion and DALL-E) and deep learning-based video generation models (e.g., DeepMind and OpenAI's Video GPT) are implemented on cloud servers.
[1299] Processing flow and data calculation
[1300] 1. User Input
[1301] Users input their ideas for the ad video they want to generate in text or sketch form. For example, they can input a prompt in text such as "A new smartphone flying through the air."
[1302] 2. Data transmission and initial image generation
[1303] The device sends input data to a cloud server, which uses a generative AI model to generate an initial image based on the input data, such as a scene of a new smartphone flying through the air.
[1304] 3. Video Generation
[1305] An initial image is input to a deep learning-based video generation model, which then generates successive frames based on the initial image, generating a continuous video from the initial image.
[1306] 4. Sending videos and receiving corrections
[1307] The generated video is sent to the device for the user to review. An interface is provided for the user to input desired corrections or changes, such as "increase the smartphone's flight speed and change the background to a sunset."
[1308] 5. Modify and Regenerate
[1309] The server then reflects the received corrections and regenerates the video. The corrections are made automatically by the AI model, significantly reducing the user's workload.
[1310] 6. Final confirmation, saving and sharing
[1311] The regenerated video is sent to the device for final confirmation by the user. If the user is satisfied, the video is saved and shared on advertising media (website, social media, email, etc.).
[1312] Specific examples
[1313] Prompt Sentence Examples
[1314] "I want to create a video of my new smartphone flying through the air."
[1315] Correction prompt: "Make my phone fly faster and change the background to a sunset."
[1316] This system allows users to easily realize their ideas and quickly create, edit, and share high-quality advertising videos.
[1317] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1318] Step 1:
[1319] Getting user input data
[1320] - Input: The user inputs the desired video content in the form of text or a sketch. For example, the user inputs "a video of a new smartphone flying in the air."
[1321] - Action: The user provides input data using a dedicated application or a web interface.
[1322] - Output: The input data is stored inside the application and is ready to be sent to the cloud server.
[1323] Step 2:
[1324] Generate the initial image
[1325] - Input: User input data sent to the cloud server.
[1326] - Operation: The server passes the received input data to a generative AI model, which generates an initial image from the text or sketch. The generative AI model uses, for example, Stable Diffusion or DALL-E.
[1327] - Output: The generative AI model generates an initial image, for example, an image of a new smartphone flying through the air.
[1328] Step 3:
[1329] Video frame generation
[1330] - Input: The generated initial image.
[1331] Operation: The server inputs the initial image into a deep learning-based video generation model to generate successive frames. The deep learning model used is, for example, Video GPT.
[1332] - Output: A series of frames are generated from the initial image and compiled into a video.
[1333] Step 4:
[1334] Submitting and reviewing videos
[1335] - Input: The generated video file.
[1336] - Operation: The server sends the generated video file to the user's device. The user checks the received video and specifies any corrections.
[1337] - Output: The video sent to the user's device and the modifications specified by the user (e.g., "increase the smartphone's flying speed and change the background to a sunset") are obtained.
[1338] Step 5:
[1339] Applying the changes and regenerating
[1340] - Input: User-specified corrections.
[1341] - Operation: The server uses the AI model and video generation model to modify the video based on the user's modifications. Specifically, it reflects instructions such as flight speed and background changes.
[1342] - Output: A new video file is generated with the modifications reflected.
[1343] Step 6:
[1344] Final confirmation, saving and sharing
[1345] - Input: The modified video file.
[1346] - Operation: The server resends the modified video file to the user's device. The user performs a final check and saves or shares the video on the advertising medium if satisfied.
[1347] - Output: The final video file is saved on the user's device, and a sharing link is generated to the specified advertising medium.
[1348] By following these steps, users can easily and quickly create, modify, and share high-quality advertising videos.
[1349] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1350] The embodiment of the present invention will now be described in detail.
[1351] System Structure
[1352] This invention provides a system that allows users to easily embody their ideas and images as videos. The system consists of a terminal that receives user input data and a server that processes the data. Furthermore, by incorporating an emotion engine that recognizes the user's emotions, it becomes possible to generate videos based on the user's emotions.
[1353] Program processing flow
[1354] 1. User Input
[1355] Using a dedicated application or web interface, users input the desired video content using text or sketches. For example, if a user wants to add the emotion "calm" to a video of a deer walking in a deep forest, they input that information. The input data is then sent to the device.
[1356] 2. Sending input data and emotion data
[1357] The device sends the user's input data and emotion data together to the server. The transmitted data includes the text information and sketches entered by the user, as well as emotion tags.
[1358] 3. Emotion Recognition and Initial Image Generation
[1359] The server receives the input data and emotion data. The server uses an emotion engine to analyze and identify the user's emotion. The recognized emotion is reflected in the initial image generation means. For example, if the user's emotion is "calm," an initial image with a calm atmosphere is generated to reflect this.
[1360] 4. Video Generation
[1361] After the initial image is generated, the server uses a video generation AI model to generate a series of frames to create a video. The model also incorporates the user's emotional data to generate videos that are more emotionally relevant. For example, a calm scene would use gentle movements and soft lighting.
[1362] 5. Send and review your video
[1363] The generated video is sent to the device and can be viewed by the user. The user can review the video and input corrections or changes as necessary.
[1364] 6. Reflecting and regenerating emotionally based modifications
[1365] When the user inputs corrections, the device sends the correction data to the server. The server then generates a new video based on the correction data and emotion data. The corrections are optimized based on the emotion data.
[1366] 7. Final check and output
[1367] The regenerated video is sent back to the device. The user finally reviews the video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[1368] Specific examples
[1369] Specific examples are shown below.
[1370] Example: A user inputs a video of a puppy running on a beach at sunset and adds the emotion "happiness."
[1371] Initial image: The server generates an initial image of a beach at sunset, and the emotion engine creates a scene with bright colors and a feeling of happiness.
[1372] Video Generation: A video of a puppy running on the beach is generated, with lively and happy movements.
[1373] Example fix: The user suggests making fixes like "Make the puppy a little faster and add the sound of the ocean."
[1374] Final output: A video is generated that reflects the edits and the user is happy to save and share the video.
[1375] As described above, the present invention is a system that incorporates a user's emotional data and generates videos that match the user's emotions, allowing the user to easily generate more specific and emotionally rich videos.
[1376] The processing flow will be explained below.
[1377] Step 1:
[1378] The user opens a dedicated application or web interface, enters the desired video content in text or draws a simple image using a sketch, and also enters the emotion they want the video to reflect (e.g., "calm" or "happy"). This prepares the device to send the specific request and emotional data.
[1379] Step 2:
[1380] The device checks the user's input data and emotion data and transmits them to the server. The transmitted data includes the user's input text information, sketches, and emotion tags.
[1381] Step 3:
[1382] The server receives the input data and emotion data, analyzes the received data, and prepares to generate an initial image based on the user's request and emotion.
[1383] Step 4:
[1384] The server uses an emotion engine to analyze the emotions contained in the user's text data and sketches. For example, if the user's emotion is analyzed as "calm," it will recognize it.
[1385] Step 5:
[1386] The server uses an image generation AI model to generate an initial image based on the user's input data and emotion data. For example, based on the user's input text "A deer walking in a deep forest" and the emotion "Calm," it generates an initial image with a calm atmosphere.
[1387] Step 6:
[1388] The server uses a video generation AI model based on the generated initial images to generate a series of frames to create a video. This model also reflects the user's emotional data, generating videos that are in line with their emotions. For example, gentle movements and soft lighting are used for calm scenes.
[1389] Step 7:
[1390] The server sends the generated video to the terminal, which receives the video and notifies the user to confirm it.
[1391] Step 8:
[1392] The user reviews the received video and decides whether they are satisfied. If the user wishes to make any corrections, they input specific corrections (e.g., changing the speed of the puppy's movements, adjusting the color) and send them to the device.
[1393] Step 9:
[1394] The device sends the user's modified data to the server, which receives it and prepares to generate a new video that reflects the modifications.
[1395] Step 10:
[1396] The server regenerates the video based on the modifications. Based on the modifications and emotion data, it regenerates consecutive frames to complete the entire video. For example, if a user requests the puppy to move a little faster, a video reflecting the speed change is generated.
[1397] Step 11:
[1398] The regenerated video is sent to the device again, and the device receives the video and notifies the user for final confirmation.
[1399] Step 12:
[1400] The user reviews the final video to their satisfaction and chooses to save or share it. The device provides the final video file and generates a sharing link if necessary.
[1401] Through these steps, users can embody their own images into emotive videos without requiring any specialized skills, and can easily modify and regenerate them.
[1402] Example 2
[1403] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1404] Currently, many video generation systems generate videos based on user input data, but lack the technology to adjust the content and atmosphere of the video based on the user's emotions. This makes it difficult to easily generate videos that embody the user's intended emotions and atmosphere. Furthermore, there is no way to reflect emotions when modifying generated videos, making it difficult to obtain results that satisfy the user.
[1405] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1406] In this invention, the server includes means for generating an initial image based on user input data and emotion data, moving image generation means for generating successive frames based on the initial image, and means for transmitting the generated moving image to the user. This makes it possible to generate a moving image that reflects the emotion data input by the user, and to regenerate the moving image while taking emotion into consideration even in subsequent modifications.
[1407] "User input data" refers to information about text, sketches, and video content entered by a user using a dedicated application or web interface.
[1408] "Emotion data" is information that represents the emotion or atmosphere that the user wants to impart to the content of the video.
[1409] An "initial image" is the first still image of a video generated based on the user's input data and emotion data.
[1410] "Video generation means" refers to the function or technology for generating a series of frames based on initial images and emotion data to create a video.
[1411] "Correction points" refer to requests for changes or corrections that a user specifies for a generated video.
[1412] "Regeneration" is the process of recreating a video based on user-specified modifications and emotion data.
[1413] An "emotion engine" is a software module or algorithm for analyzing and recognizing a user's emotional data.
[1414] A "deep learning model" is an artificial intelligence technology that uses large amounts of data to learn and generate sophisticated videos from the input data.
[1415] The present invention provides a system for generating videos that reflect a user's emotions by incorporating the user's emotional data. This system allows users to easily embody their own ideas and images as videos.
[1416] System Structure
[1417] The system consists of a terminal that receives user input data and a server that processes the data. Furthermore, by incorporating an emotion engine, it becomes possible to generate videos based on the user's emotions.
[1418] Hardware and software used
[1419] Device: The device on which the user provides input data (e.g., smartphone, tablet, computer)
[1420] Server: A server that processes input data and emotion data and generates videos.
[1421] Emotion engine: A software module that analyzes and recognizes user emotion data.
[1422] Generative AI model: A deep learning model that generates videos based on input data and emotion data
[1423] A concrete example of the video generation flow
[1424] User Input
[1425] The user inputs the desired video content using a dedicated application or web interface. For example, if the user wants to add the emotion "calm" to a video of a deer walking in a deep forest, the user inputs the information in text and selects "calm" as the emotion tag. This input is sent to the device.
[1426] Example prompt sentence:
[1427] Add the emotion "Happiness" to a video of a puppy running on a beach at sunset.
[1428] Sending input data and emotion data
[1429] The device transmits the user's input data and emotion data to the server. The transmitted data includes the text information and sketches entered by the user, as well as emotion tags.
[1430] Emotion Recognition and Early Image Generation
[1431] The server analyzes the received input data and emotion data and recognizes the emotion using the emotion engine. The recognized emotion is reflected in the initial image generation means. For example, if the user adds the emotion "calm," the server generates an initial image with a calm atmosphere.
[1432] Video generation
[1433] The server then uses a generative AI model to generate a series of frames based on the initial images, creating a video. This process also incorporates the user's emotional data, generating videos that are more in line with their emotions. For example, calm scenes will use gentle movements and soft lighting.
[1434] Submitting and reviewing videos
[1435] The generated video is sent to the device and can be viewed by the user, who can review the video and enter corrections or changes as necessary.
[1436] Reflecting and regenerating emotion-based modifications
[1437] When the user inputs corrections, the device sends the correction data to the server. The server then generates a new video based on the correction data and emotion data. The corrections are also optimized based on the emotion data.
[1438] Final check and output
[1439] The regenerated video is then sent back to the device, where the user can finally review the video to their satisfaction and save or share it. The device will provide the final video file and generate a sharing link if necessary.
[1440] As described above, the present invention provides a system that allows users to easily create and edit videos that reflect their own emotions and intentions. This system utilizes emotion data to provide more specific and emotionally rich videos.
[1441] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1442] Step 1:
[1443] Using a dedicated application or web interface, users input the content of the desired video and related emotional data. This input includes text information, sketches, and emotion tags. This input data is sent to the device. Specifically, the user inputs "a video of a deer walking in a deep forest" and the emotion "calm," and the data is sent to the device.
[1444] Input: Text information, sketch, emotion tag
[1445] Output: Input data sent to the terminal
[1446] Step 2:
[1447] The device combines the user's input data and emotion data into a single data packet and sends it to the server. Specifically, it combines the text information "A deer is walking in a deep forest" and the emotion tag "Calm" in JSON format. The device creates a data packet and sends it to the server.
[1448] Input: User input data and emotion data
[1449] Output: Data packet sent to the server
[1450] Step 3:
[1451] The server analyzes the received data packets and uses the emotion engine to recognize and analyze the user's emotions. Specifically, the server analyzes the received data and recognizes the emotion "calm." This emotion data is reflected in the initial image generation means, as it affects subsequent processing.
[1452] Input: Data packet (user input data and emotion data)
[1453] Output: Recognized emotion data
[1454] Step 4:
[1455] The server uses the initial image generation means to generate an initial image based on the user's input data and the recognized emotion data. Specifically, an initial image depicting a forest landscape in gentle colors is created. This initial image is used as the first step in video generation.
[1456] Input: User input data, recognized emotion data
[1457] Output: Initial image
[1458] Step 5:
[1459] The server uses a generative AI model to generate a series of frames based on the generated initial image, creating a video. Specifically, the video generation AI model receives the initial image and emotion data as input and generates a series of frames of a deer in a tranquil scene. This series of frames is then combined to form a complete video.
[1460] Input: Initial image, recognized emotion data
[1461] Output: Video with continuous frames
[1462] Step 6:
[1463] The generated video is sent from the server to the device, where the user can view it. The user plays the video through the application and checks the content. Specifically, the video data from the server arrives at the device, and the user clicks the "play" button to watch the video.
[1464] Input: Generated video
[1465] Output: Video sent to device, user confirmation
[1466] Step 7:
[1467] If a user wishes to make corrections to the video content, they input the corrections and send the correction data from their device to the server. Specifically, the user might input, "Make the deer move a little faster and add the sound of birds singing in the background." This correction data is then sent from the device to the server.
[1468] Input: User modifications
[1469] Output: Corrected data sent to the server
[1470] Step 8:
[1471] The server regenerates the initial image and video based on the received correction data and emotion data. Specifically, the server generates new initial images and consecutive frames, and regenerates a video that reflects the corrections and emotion data.
[1472] Input: Corrected data, recognized emotion data
[1473] Output: Regenerated video
[1474] Step 9:
[1475] The regenerated video is sent from the server to the device, and the user is finally confirmed to be satisfied. The user watches the video again, and if satisfied, can save or share it. Specifically, the server sends the regenerated video to the device, and if the user is satisfied, they click the "Save" or "Share" button.
[1476] Input: Regenerated video
[1477] Output: User final review, save or share
[1478] (Application example 2)
[1479] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1480] Current video generation systems make it difficult for users to easily create videos that reflect their desired emotions. They also lack the functionality to easily share generated videos on online platforms. This makes it difficult to quickly create and share emotionally rich, effective videos, especially in the advertising field.
[1481] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1482] In this invention, the server includes means for receiving user input data, means for generating an initial image based on the user input data, video generation means for generating successive frames based on the initial image, means for transmitting the generated video to the user, means for receiving corrections specified by the user, means for regenerating the video based on the corrections, means for the video generation means to generate a video reflecting the user's emotional data, an emotion engine for analyzing the user's emotional data, and means for generating a link for sharing the generated video on an online platform. This enables users to easily generate effective videos that reflect emotions and quickly share them online.
[1483] "User input data" refers to information such as text or sketches that a user provides to the system.
[1484] "Initial image" refers to the first still image generated based on user input data.
[1485] "Video generator" refers to a mechanism for generating a sequence of video frames based on an initial image.
[1486] "Correction points" refer to requests for changes or corrections that a user specifies for a video.
[1487] "Regeneration" refers to regenerating a video by reflecting the corrections.
[1488] "Emotional data" refers to information about emotions that a user provides to the system.
[1489] An "emotion engine" refers to software or algorithms that analyze the emotional data provided by the user and use that data to generate videos.
[1490] "Link for sharing on online platforms" refers to a URL or sharing link for sharing the generated video with other users on the Internet.
[1491] "Consecutive frames" refers to a number of still images that make up a moving image.
[1492] "Means for transmitting to the user" refers to a communication means for providing the generated video to the user.
[1493] The present invention relates to a system for enabling users to quickly and effectively create emotional videos and easily share them on an online platform, and a specific embodiment of the system will be described below.
[1494] System Configuration
[1495] The system of the present invention is composed of a terminal that receives user input data, a server that processes the data, and a communication means that provides the generated video to the user. It also includes an emotion engine for emotion recognition and an AI model for video generation.
[1496] Hardware and software used
[1497] Terminal: A device that allows users to input text and sketches necessary for video generation. This includes smartphones and PCs.
[1498] Server: Provides the computing resources to analyze user input data and emotion data and generate videos. Cloud-based servers are often used.
[1499] Emotion engine: Software for analyzing emotion data and generating initial images based on it.
[1500] Video generation AI model: A deep learning model that generates a series of frames based on initial images to create emotionally relevant videos.
[1501] Communication means: A network that transmits the generated video data to the terminal and allows the user to view the video.
[1502] Video generation process
[1503] 1. Getting User Input
[1504] Users can use a dedicated application or web interface to input the content of the video and the desired emotion using text or a sketch. For example, they can input "a video introducing the latest running shoes" and the emotion "excitement."
[1505] 2. Data transmission and processing
[1506] The device sends the user's input data and emotion data to the server, which receives the data, analyzes the emotion using an emotion engine, and generates an initial image.
[1507] 3. Generate initial images
[1508] The emotion engine analyzes the user's emotional data and generates an initial image that reflects that emotion. If the emotion is "excitement," vivid colors and dynamic compositions are used.
[1509] 4. Video Generation
[1510] The server uses a video generation AI model based on the generated initial images to generate successive frames, which then creates a video that reflects the emotion.
[1511] 5. Send and review your video
[1512] The generated video is sent from the server to the device, where the user can review it and make corrections if necessary.
[1513] 6. Processing and regenerating fixes
[1514] If the user inputs corrections, the data is sent to the server again, and the server generates a new video based on the corrections and emotion data.
[1515] 7. Video sharing
[1516] The final video is provided as a link that users can use to share the video on online platforms.
[1517] Examples and prompts
[1518] As a concrete example, consider a case where a user wants to generate an advertising video of a family picnic scene in a park with spring flowers blooming with the emotion of "happiness." The user inputs the following prompt sentence:
[1519] Example prompt sentence
[1520] Advertising video of a family picnic scene in a park with spring flowers
[1521] Emotion: euphoria
[1522] Based on this prompt, a video is generated that reflects the user's emotions, showing smiling faces and warm colors. The generated video can be easily shared on the user's preferred platform.
[1523] The above is a specific embodiment for carrying out the present invention. This system enables advertisers to quickly create emotionally rich and effective videos and use them in their marketing activities.
[1524] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1525] Step 1:
[1526] Users use a dedicated application or web interface to input text, sketches, and emotional information required for video generation. Input includes "a video introducing the latest running shoes" and the emotion "excitement." The device collects this input data and stores it as part of its data.
[1527] Step 2:
[1528] The device sends the user's input data (text, sketches, emotion information) to the server, which receives it and prepares each data for analysis. The input data is organized for further processing.
[1529] Step 3:
[1530] The server uses an emotion engine to analyze the user's emotional information. It generates an initial image based on the results of this analysis. If the input data is "excitement," an initial image with bright colors and a dynamic composition is generated. This initial image becomes the basis for the next video generation.
[1531] Step 4:
[1532] The server supplies the generated initial images to a video generation AI model, which generates successive frames. The generative AI model generates successive frames based on the initial images and emotion data received as input. The output is video data that reflects the emotion.
[1533] Step 5:
[1534] The server sends the generated video to the terminal. The terminal receives the video data and provides it to the user for confirmation. The user can view the video and input corrections as necessary.
[1535] Step 6:
[1536] If the user inputs corrections, the device sends the correction data back to the server, which then receives the corrections and reuses the emotion engine and video generation AI model to generate a new video that reflects the corrections.
[1537] Step 7:
[1538] The server then sends the new video with the modifications reflected back to the device, and the user repeats this process until they are finally satisfied, at which point they are asked to generate a link to save the video or share it on an online platform.
[1539] As described above, users, devices, and servers work together to efficiently generate videos that reflect emotions, and then modify and optimize them as needed, completing the entire process of sharing them online.
[1540] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1541] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1542] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1543] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1544] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1545] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1546] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1547] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1548] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1549] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1550] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1551] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1552] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1553] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1554] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1555] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1556] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1557] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1558] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1559] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1560] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1561] The following is further disclosed regarding the above embodiment.
[1562] (Claim 1)
[1563] means for receiving user input data;
[1564] means for generating an initial image based on the user input data;
[1565] a video generation means for generating successive frames based on the initial image;
[1566] means for transmitting the generated video to a user;
[1567] means for receiving said user-specified corrections;
[1568] means for regenerating the video based on the modifications;
[1569] A system including:
[1570] (Claim 2)
[1571] The system of claim 1 , wherein the user provides input data by text or a sketch.
[1572] (Claim 3)
[1573] 2. The system of claim 1, wherein the video generation means utilizes a deep learning model to generate videos including continuous motion.
[1574] "Example 1"
[1575] (Claim 1)
[1576] a terminal for receiving user input data;
[1577] means for transmitting the user input data to a server;
[1578] A means for utilizing an image generation AI model to generate an initial image based on the user's input data;
[1579] a means for utilizing an AI video generation model to generate successive frames based on the initial images;
[1580] means for transmitting the generated video to a terminal;
[1581] a terminal that receives the corrections specified by the user;
[1582] means for regenerating the video based on the modifications;
[1583] means for transmitting the regenerated video to the terminal;
[1584] A system including:
[1585] (Claim 2)
[1586] The system of claim 1 , wherein the user provides input data by text or a sketch.
[1587] (Claim 3)
[1588] 2. The system of claim 1, wherein the video generation means utilizes a deep learning model to generate videos including continuous motion.
[1589] "Application Example 1"
[1590] (Claim 1)
[1591] means for receiving user input data;
[1592] means for generating an initial image based on the user input data;
[1593] a video generation means for generating successive frames based on the initial image;
[1594] means for transmitting the generated video to a user;
[1595] means for receiving said user-specified corrections;
[1596] means for regenerating the video based on the modifications;
[1597] a means for storing and sharing user-generated videos on advertising media;
[1598] A system including:
[1599] (Claim 2)
[1600] The system of claim 1 , wherein the user provides input data by text or a sketch.
[1601] (Claim 3)
[1602] 2. The system of claim 1, wherein the video generation means utilizes a deep learning model to generate videos including continuous motion.
[1603] "Example 2: Combining Emotion Engines"
[1604] (Claim 1)
[1605] means for receiving user input data;
[1606] means for generating an initial image based on the user's input data and emotion data;
[1607] a video generation means for generating successive frames based on the initial image;
[1608] means for transmitting the generated video to a user;
[1609] means for receiving said user-specified corrections;
[1610] means for regenerating a video based on the corrections and emotion data;
[1611] means for analyzing and recognizing the emotions of the user;
[1612] A system including:
[1613] (Claim 2)
[1614] The system of claim 1 , wherein the user provides input data by text or a sketch.
[1615] (Claim 3)
[1616] 2. The system of claim 1, wherein the video generation means utilizes a generative AI model to generate videos including continuous motion.
[1617] "Application example 2 when combining emotion engines"
[1618] (Claim 1)
[1619] means for receiving user input data;
[1620] means for generating an initial image based on the user input data;
[1621] a video generation means for generating successive frames based on the initial image;
[1622] means for transmitting the generated video to a user;
[1623] means for receiving said user-specified corrections;
[1624] means for regenerating the video based on the modifications;
[1625] means for generating a video by the video generating means, the video generating means generating a video by reflecting the emotion data of the user;
[1626] an emotion engine that analyzes emotion data of the user;
[1627] means for generating a link for sharing the generated video on an online platform;
[1628] A system including:
[1629] (Claim 2)
[1630] The system of claim 1 , wherein the user provides input data by text or a sketch.
[1631] (Claim 3)
[1632] 2. The system of claim 1, wherein the video generation means utilizes a deep learning model to generate videos including continuous motion. [Explanation of symbols]
[1633] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving user input data; means for generating an initial image based on the user input data; a video generation means for generating successive frames based on the initial image; means for transmitting the generated video to a user; means for receiving said user-specified corrections; means for regenerating the video based on the modifications; A system including:
2. The system of claim 1 , wherein the user provides input data by text or a sketch.
3. The system of claim 1 , wherein the video generation means utilizes a deep learning model to generate videos including continuous motion.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A