System

The system uses a generative AI model and video generation engine to create high-quality videos from user inputs, addressing the challenges of skill and hardware requirements in conventional video production.

JP2026021083APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122765
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Conventional video production methods require specialized skills, long hours of work, and expensive hardware, making it difficult for ordinary users to easily visualize their ideas and create high-quality video content.

Method used

A system utilizing a generative AI model to generate a video script based on user linguistic instructions, a video generation engine to create the video file, and a database to store and manage the video files with unique identifiers, allowing users to easily create and access high-quality videos without advanced knowledge or expensive hardware.

Benefits of technology

Enables users to quickly visualize their ideas and generate high-quality video content efficiently, without the need for specialized skills or expensive equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021083000001_ABST
    Figure 2026021083000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for generating a video script based on a verbal command of a user using a generative artificial intelligence model; video generation engine means for creating a video file based on the generated video script; and database means for storing the generated video file and issuing an identifier of a storage location.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional video production methods require specialized skills, long hours of work, and a large amount of resources. This makes it difficult for ordinary users to easily visualize their ideas. Furthermore, with content requiring expensive hardware like the Metaverse, it is difficult to provide video content that viewers can easily consume. These issues need to be resolved. [Means for solving the problem]

[0005] The present invention is a system that includes a means for generating a video script based on a user's linguistic instructions using a generative AI model, a video generation engine means for creating a video file based on the generated video script, and a database means for saving the generated video file and issuing an identifier for the storage destination. This allows users to quickly visualize their ideas without requiring advanced specialized knowledge. Furthermore, users can easily enjoy high-quality video content without the need for expensive hardware.

[0006] A "generative artificial intelligence model" refers to an algorithm or program that analyzes a user's linguistic instructions and preferences and generates a video script based on them.

[0007] "User's verbal instructions" refer to the user's provided ideas, preferences, and detailed instructions regarding the content of the video content.

[0008] "Video script" refers to text data that describes detailed instructions and structure for the scenes, narration, music, etc. of the generated video.

[0009] "Video generation engine" refers to a combination of software and hardware for creating video files based on a generated video script.

[0010] "Database" refers to a system for storing generated video files and managing and issuing identifiers for the storage destinations.

[0011] "Identifier" refers to a unique string or code used to identify the location and data of a stored video file. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0020] [First embodiment]

[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0033] An embodiment of the present invention is described below: The system uses a generative AI model to generate a video script based on a user's linguistic instructions, creates and saves a video file based on the generated video script, and provides the user with an identifier for the saved file.

[0034] 1. Creating and Sending a Request

[0035] The user inputs their ideas into the interface. For example, they specify an idea for an "adventure with a space travel image" or a preference for "upbeat, up-tempo music." This information is converted into a request object by the terminal and sent to the server.

[0036] 2. Receiving and parsing the request

[0037] The server receives requests from users and extracts information such as their ideas and preferences, which are then prepared as input data for a generative AI model.

[0038] 3. Generate the script

[0039] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences, detailing the video's scenes, narration, music, and more.

[0040] 4. Image Generation

[0041] The server's video generation engine creates the video based on the generated script, including rendering the scene and integrating music. The generation engine creates high-quality video files according to the instructions in the script.

[0042] 5. Storage of footage and issuance of identifiers

[0043] The server saves the created video file in a database. Once saving is complete, the server generates an identifier for the storage location and returns it to the user. This identifier allows the user to access and view the created video.

[0044] Specific examples

[0045] For example, consider a user who wants to generate a video with the theme of "an adventure with the image of space travel." The user inputs the theme and preferences into the interface and submits a request. The server receives this request and uses a generative AI model to generate a script based on the theme of "an adventure with the image of space travel." The server's video generation engine then creates a video based on the script and stores it in a database. Finally, the user can watch their video using the identifier returned by the server.

[0046] The processing flow will be explained below.

[0047] Step 1:

[0048] The user inputs their ideas and preferences into the interface. For example, they specify the idea of ​​"an adventure with the image of space travel" and the preference of "upbeat, up-tempo music."

[0049] Step 2:

[0050] The device generates a request object (user_request) based on the user's input, which contains information about the user's ideas and preferences.

[0051] Step 3:

[0052] The terminal sends the generated request object to the server, which contains data based on the user's input.

[0053] Step 4:

[0054] The server receives the user's request and analyzes it, which includes extracting information about the user's ideas and preferences.

[0055] Step 5:

[0056] Based on the extracted information, the server prepares input data (ai_input) for the generative AI model, which contains detailed information about the user's ideas and preferences.

[0057] Step 6:

[0058] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences. The AI ​​model analyzes the input data and outputs the script.

[0059] Step 7:

[0060] The server's video generation engine creates the video based on the generated script, a process that includes rendering the scene and integrating music.

[0061] Step 8:

[0062] The server saves the completed video file in a database. During this saving process, an identifier is issued to ensure that the video file is properly managed.

[0063] Step 9:

[0064] The server returns the identifier of the video stored in the database to the user, allowing the user to access the generated video.

[0065] Step 10:

[0066] The user uses the returned identifier to view the video on their device, which allows them to stream or download their video content.

[0067] Example 1

[0068] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0069] Conventional video generation systems require users to have advanced skills and specialized knowledge to create specific video scripts, making them difficult for average users to use. Furthermore, the process of manually creating a script and then generating a video file is time-consuming, labor-intensive, and inefficient.

[0070] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0071] In this invention, the server includes: means for a user to input ideas and preferences as linguistic instructions into an interface; means for converting the user input into a request object and sending the request to the server; means for the server to receive the request and extract information on the ideas and preferences; means for generating a video script based on the user's ideas and preferences using a generative AI model; means including a video generation engine that creates a video file based on the generated video script; and means for saving the generated video file in a database and issuing an identifier for the storage destination. This allows users without advanced skills or specialized knowledge to easily generate video files based on linguistic instructions and efficiently use the videos.

[0072] A "user" is a person who uses the system to input verbal instructions into the interface to generate a video file.

[0073] An "interface" is an input device or software that allows a user to input their ideas and preferences.

[0074] A "request object" is a collection of information that is converted from user input into a data format and sent to the server.

[0075] "Server" means a computer system that receives and analyzes a request object from a user, generates a video script using a generative AI model, and then generates and stores the video file.

[0076] A "generative AI model" is an artificial intelligence model for generating video scripts based on a user's ideas and preferences.

[0077] A "video script" is a document that describes detailed instructions for scenes, narration, music, etc. for generating a video file.

[0078] A "video generation engine" is a device or software that creates a video file based on a video script.

[0079] A "database" is a system that stores generated video files and manages information for later access.

[0080] An "identifier" is a character string or number that uniquely identifies the generated video file and returns information to the user.

[0081] This embodiment of the present invention is a system that generates a video script based on a user's verbal instructions, and then creates and saves a video file based on the script. This system consists of three main components: a user, a terminal, and a server.

[0082] First, the user inputs their ideas and preferences into an interface, which can be implemented as a computer or smartphone application. For example, specific instructions such as "an adventure with a space travel image" or "upbeat, upbeat music" can be entered in text format.

[0083] The device then converts the user's input into a request object and sends it to the server, which includes the user's specified themes and preferences, allowing the server to convey the user's request in the appropriate data format.

[0084] The server receives the request sent from the device and analyzes the ideas and preferences contained therein. The analyzed information is prepared as input data for a generative AI model, such as OpenAI GPT-4.

[0085] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences. The script contains detailed instructions for the video's scenes, narration, music, etc. After the script is generated, the server's video generation engine creates a video file based on the script.

[0086] The video generation engine uses software such as Adobe After Effects to render the scenes and integrate the music. The generated video files are of high quality and are based on the detailed instructions in the script.

[0087] Finally, the server stores the generated video file in a database. The database uses a high-speed storage system to efficiently manage the video files. Once the storage is complete, the server generates a unique identifier for the video file and returns this identifier to the user. The user can use this identifier to access and watch the generated video file.

[0088] As a concrete example, consider a case where a user wants to generate a video with the theme of "adventure with a space travel image" and an upbeat, up-tempo music. The user inputs the theme and preferences into the interface and sends the request to the server. The server analyzes the request and generates a video script using a generative AI model. The server's video generation engine then creates a video based on the script and stores it in a database. Finally, the user can watch their video using the identifier returned by the server.

[0089] Examples of prompt sentences include the following:

[0090] Please create a video with the theme of "An adventure with the image of space travel." I would like to use upbeat music with a bright atmosphere.

[0091] In this way, by using this system, users can easily generate and manage high-quality video based on their own ideas and preferences, without needing advanced specialized knowledge.

[0092] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0093] Step 1: User Input

[0094] The user inputs their ideas and preferences for the video into the interface. For example, they can input a theme such as "an adventure with a space travel image" or preferences such as "a bright atmosphere and up-tempo music." The input data is sent to the terminal as text data.

[0095] Step 2: Create a request object

[0096] The terminal converts textual input data from the user into a request object, which contains themes and preferences as key-value pairs. For example, a request object might contain the following data:

[0097] json

[0098] {

[0099] "theme": "Space travel-inspired adventure",

[0100] "mood": "upbeat, upbeat music"

[0101] }

[0102] The created request object is sent to the server.

[0103] Step 3: Receiving and Parsing the Request

[0104] The server receives the request object sent from the device, extracts idea and preference information from the request object, and prepares the extracted data as input data for the generative AI model.

[0105] Step 4: Input to the generative AI model

[0106] The server inputs the extracted ideas and preferences into a generative AI model, which generates a prompt like this:

[0107] Create a video with the theme of "An adventure with the image of space travel." Use upbeat, up-tempo music.

[0108] Input data is transformed into a generative AI model.

[0109] Step 5: Generate a video script

[0110] The server invokes the generative AI model to generate a video script based on the input data. The generated script contains details such as video scenes, narration, and music. For example, the following script might be generated:

[0111] json

[0112] {

[0113] "scenes": [

[0114] {"description": "A spaceship traveling through the galaxy", "duration": "30 seconds"},

[0115] {"description": "Spaceship landing scene", "duration": "20 seconds"}

[0116] ],

[0117] "narration": "The space travel adventure begins...",

[0118] "music": "Uptempo background music"

[0119] }

[0120] Step 6: Rendering the footage

[0121] The server's video generation engine creates the video based on the generated script. First, the video generation engine renders the scene. This process includes generating CG and editing existing footage. Next, music and narration are integrated, and the final video file is generated.

[0122] Step 7: Save the video file

[0123] The server stores the generated video files in a database, a process that uses a high-speed storage system.

[0124] Step 8: Generate and return identifiers

[0125] The server generates a unique identifier for each video file that has completed the storage process. This identifier will have the following format:

[0126] json

[0127] {

[0128] "identifier": "abcd1234"

[0129] }

[0130] The identifier is returned to the user, who can then use it to access and view the generated video file.

[0131] The above is the flow of processing of the program of this system.

[0132] (Application example 1)

[0133] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0134] Conventional video generation systems lack the ability to quickly generate and instantly distribute high-quality videos based on user ideas. In particular, there has been no system that allows users to easily input ideas using a smartphone and then view and share the videos generated on the spot. This has made it difficult to meet users' demands for creative content generation and instant distribution.

[0135] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0136] In this invention, the server includes: means for generating a video script based on a user's linguistic instructions using a generative artificial intelligence model; a video generation engine; a database; means for a user to input ideas and preferences using a smartphone; means for formatting the input data as input data for the generative AI model; and means for providing the user with a preview and sharing of the generated video file, thereby enabling the user to generate high-quality videos in real time using a smartphone and instantly share the videos on various platforms.

[0137] A "generative artificial intelligence model" is an algorithm that analyzes a user's linguistic instructions and automatically generates a video script based on that information.

[0138] A "video script" is a document that details the video's scenes, narration, music, etc.

[0139] A "video generation engine" is software or hardware that generates the actual video file based on a video script.

[0140] A "database" is an information management system that stores the generated video files and allows users to access them.

[0141] An "identifier" is an ID or code that uniquely identifies a specific video file within a database.

[0142] A "smartphone" is a multifunctional mobile device that combines mobile communication and computer functions.

[0143] "Formatting" is the process of converting input data into a specific format that is easy for a generative AI model to understand.

[0144] "Preview" is a function that displays a part or all of a generated video file in advance before the user finally views it.

[0145] "Sharing" is the process of sharing the generated video file with other users and platforms.

[0146] The system embodying this invention allows users to easily input their ideas and preferences using a smartphone, and then generate, view, and share high-quality video on the spot.

[0147] First, the user uses a dedicated smartphone application to input their ideas and preferences for the video they want to create, including a specific prompt such as "An adventure story exploring the mysteries of space in a light-hearted atmosphere." The application then formats the data entered by the user into JSON format and sends it to the server.

[0148] The server generates a video script using a generative AI model based on the received request. The generative AI model analyzes the user's linguistic instructions and creates a video script that details the development of scenes, narration content, music selection, etc. This script is then converted into an actual video file by the video generation engine on the server.

[0149] The video generation engine renders multiple scenes based on the generated video script and integrates appropriate music. The server stores the rendered and music-integrated video files in a database and issues an identifier for the storage location. This identifier is used by users to access and watch the generated video files.

[0150] Finally, users receive the issued identifier through a smartphone application and can preview the video generated based on their input idea. This video can be instantly shared within the application or on platforms such as social media.

[0151] The primary hardware used includes smartphones (e.g., iPhone, Android smartphone) and servers (e.g., AWS EC2 instances).The software used includes smartphone applications (e.g., developed with React Native), server-side generative AI models (e.g., GPT-4), a Flask-based backend, and a video generation engine (e.g., FFmpeg).

[0152] In this way, users can easily generate high-quality videos and share them on a variety of platforms. An example of a specific prompt is "An adventure story exploring the mysteries of space in a cheerful atmosphere." Based on this prompt, the generative AI model generates a video script that meets the user's needs, and high-quality videos are automatically generated based on that script.

[0153] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0154] Step 1:

[0155] Users launch a dedicated smartphone application and input their ideas and preferences for the video they want to create into the interface, typically entering prompts in text format such as "An adventure story exploring the mysteries of space in a cheerful atmosphere." The input data is then formatted in JSON format within the application.

[0156] Input: User's ideas and preferences (prompt text)

[0157] Output: Request data in JSON format

[0158] Step 2:

[0159] The device sends formatted JSON data to the server, and the request contains the user's ideas and preferences.

[0160] Input: Request data in JSON format

[0161] Output: Request sent to server

[0162] Step 3:

[0163] The server analyzes the received request and prepares the prompt sentence as input data for the generative AI model. The server then invokes the generative AI model to generate a video script based on the user's linguistic instructions.

[0164] Input: Request data in JSON format

[0165] Output: Generated video script

[0166] Step 4:

[0167] The server's video generation engine creates video files based on the generated video script, rendering the scenes and integrating music.

[0168] Input: Generated video script

[0169] Output: Video file

[0170] Step 5:

[0171] The server stores the generated video file in a database and issues an identifier for the storage location, which is used by users to access the generated video.

[0172] Input: Video file

[0173] Output: Destination identifier

[0174] Step 6:

[0175] The server returns the storage location identifier to the smartphone application, allowing the user to check the video generated based on their own idea.

[0176] Input: Destination identifier

[0177] Output: Send identifier to smartphone

[0178] Step 7:

[0179] Users can use a smartphone application to preview and watch the generated video, and can also share the video within the application or on platforms such as social media.

[0180] Input: Destination identifier

[0181] Output: Preview and share your footage

[0182] This will enable high-quality videos to be quickly generated, viewed, and shared based on user-input ideas and preferences.

[0183] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0184] An embodiment of the present invention is described below: The system uses a generative AI model and an emotion engine to generate a video script based on a user's linguistic instructions and recognized emotions, and creates and saves a video file based on the video script.

[0185] 1. Creating and Sending a Request

[0186] The user inputs their ideas and preferences into the interface. For example, they specify the idea of ​​"an adventure with a space travel image" and their preference of "upbeat, upbeat music." At this point, the emotion engine is activated to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to obtain emotional data. This information is converted into a request object by the terminal and sent to the server.

[0187] 2. Receiving and parsing the request

[0188] The server receives user requests and emotional data, extracts information on ideas, preferences, and recognized emotions, and prepares this information as input to a generative AI model.

[0189] 3. Generate the script

[0190] The server invokes a generative AI model to generate a video script based on the user's ideas, preferences, and emotions. The generative AI model analyzes these inputs and outputs a script that reflects the video's scenes, narration, musical details, and emotional tone.

[0191] 4. Image Generation

[0192] The server's video generation engine creates the video based on the generated script. This process includes rendering the scene and integrating music. The video generation engine creates high-quality video files according to the instructions in the script.

[0193] 5. Storage of footage and issuance of identifiers

[0194] The server stores the completed video file in a database. During this storage process, an identifier is issued to ensure the video file is properly managed.

[0195] Specific examples

[0196] For example, consider a user who wants to create a video with the theme of "an adventure with the image of space travel." The user inputs the theme and preferences into the interface, while the emotion engine recognizes the user's excited facial expression and tone of voice. The server receives this request and emotion data and generates a script using a generative AI model. The script reflects the theme of "an adventure with the image of space travel" and the excited tone. The video generation engine then generates a video based on the script and stores it in a database. Finally, the user can view their video using an identifier returned by the server. This identifier allows the user to access the generated video through a web browser or application.

[0197] The processing flow will be explained below.

[0198] Step 1:

[0199] The user inputs their ideas and preferences into the user interface. For example, they might specify a theme such as "an adventure with a space travel image" and preferences such as "a cheerful atmosphere and up-tempo music." The emotion engine also analyzes the user's facial expressions and voice via a camera and microphone built into the user interface.

[0200] Step 2:

[0201] The device receives the user's input data and the emotion data from the emotion engine, and combines them to generate a request object (user_request), which contains all the themes, preferences, and emotion data.

[0202] Step 3:

[0203] The terminal sends the generated request object to the server, which includes the user's ideas, preferences, and emotion data.

[0204] Step 4:

[0205] The server receives the request object from the user and analyzes its contents, which includes extracting the user's themes, preferences, and sentiment data.

[0206] Step 5:

[0207] Based on the extracted information, the server prepares input data (ai_input) for the generative AI model, which includes the user's themes, preferences, and emotional data.

[0208] Step 6:

[0209] The server invokes a generative AI model to generate a video script based on the user's themes, preferences, and emotions. The generative AI model analyzes these input data and outputs a video script, which includes video scenes, narration, music, and a tone that reflects the user's emotions.

[0210] Step 7:

[0211] The server's video generation engine creates the video based on the generated script. This process includes rendering the scenes described in the script and integrating music. The video generation engine creates high-quality video files according to the instructions in the script.

[0212] Step 8:

[0213] The server stores the completed video file in a database. During this storage process, an identifier is issued to ensure the video file is properly managed.

[0214] Step 9:

[0215] The server returns to the user an identifier for the video stored in the database, allowing the user to access the generated video.

[0216] Step 10:

[0217] The user uses the returned identifier to watch the video on their device, which allows them to stream or download their video content through a web browser or dedicated app.

[0218] Example 2

[0219] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0220] Conventional video generation systems generate scripts based solely on the user's verbal instructions, making it difficult to reflect the user's emotions and detailed preferences. Furthermore, when the management of generated video files and the issuance of identifiers are done manually, this increases the administrative workload and increases the likelihood of errors.

[0221] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for generating a video script based on a user's linguistic instructions and recognized emotions using a generative AI model, a video generation engine means for creating a video file based on the generated video script, a database means for saving the generated video file and issuing an identifier for the saving destination, and means for analyzing the user's input data and emotion data, generating it as a request object, and transmitting it to the server. This enables video generation that reflects the user's emotions and detailed preferences, and further enables efficient automatic management of the generated video files and the issuance of identifiers.

[0222] A "generative artificial intelligence model" is a model that uses artificial intelligence technology to generate video scripts and other content based on user input data and emotional data.

[0223] A "video script" is a set of instructions that describes the scene structure, narration, musical details, and emotional tone for creating a video.

[0224] A "video generation engine" is a software or hardware system for creating the actual video based on the generated video script.

[0225] "Database" refers to a storage device and associated management system for storing and managing generated video files.

[0226] A "request object" is a data structure that includes user input data and emotion data and is generated to be sent to a server.

[0227] The "emotion engine" is a system that analyzes the user's facial expressions and tone of voice, recognizes emotions, and generates and provides that data.

[0228] An "identifier" is a unique code or number that uniquely identifies the generated video file.

[0229] "Input data" is data, including ideas and preferences, that a user provides to the system.

[0230] An embodiment of the present invention is a system for generating a video script based on a user's linguistic instructions and recognized emotions using a generative AI model and an emotion engine, and creating and saving a high-quality video file based on the script. The system includes a server, a terminal, and various user interfaces.

[0231] The user interface (e.g., a PC or smartphone application) provides a means for the user to input their ideas and preferences. The emotion engine is used to analyze the user's facial expressions and tone of voice to obtain emotion data. This emotion data, along with the user's input data, is packaged into a request object and sent from the device to the server.

[0232] The server receives the user's request object and analyzes the user's ideas, preferences, and emotional data using a video script generation method based on a generative AI model. The generative AI model is often operated using cloud computing resources. Based on the input data, the model generates a specific video script including scene composition, narration, musical details, and emotional tone.

[0233] The generated video script is then converted into a video file by a server-based video generation engine, which consists of advanced software and hardware for scene rendering and music integration. The video generation engine performs the necessary data calculations and processing according to the instructions to create a high-quality video file.

[0234] The generated video file is stored in a database on the server. During the storage process, an identifier is generated to uniquely identify the video file. This identifier allows users to access the completed video file. The identifier is used in web browsers and applications, providing users with an easy way to view the video.

[0235] Specific examples

[0236] For example, consider a case where a user wants to generate a video with the theme of "adventure with the image of space travel." The user inputs the theme and preferences (e.g., "upbeat atmosphere, up-tempo music") into the interface, and the emotion engine recognizes the user's excited facial expression and tone of voice. Next, the user's input data and emotion data are packaged into a request object and sent to the server.

[0237] The server receives this request and generates a script using a generative AI model. The generative AI model outputs a script that reflects an exciting tone in line with the theme of "an adventure with the image of space travel." Based on this script, a video generation engine renders the scene and integrates music and narration to create a video file. Finally, the server stores the video file in a database and returns an identifier to the user. The user can use this identifier to watch the video they generated.

[0238] Prompt Sentence Examples

[0239] "The theme specified by the user is 'an adventure with the image of space travel,' and their preference is 'a cheerful atmosphere, up-tempo music.' We also recognized an excited tone as the user's emotion. Please generate a video script based on this."

[0240] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0241] Step 1: User Input and Emotion Recognition

[0242] The user uses the interface to input their ideas and preferences. For example, the user may specify a theme of "adventure with space travel imagery" and a preference for "upbeat, upbeat music."

[0243] The emotion engine analyzes the user's facial expressions and tone of voice in real time to obtain emotion data. For example, if the user has an excited expression, that data is obtained by the emotion engine.

[0244] Input: Your ideas, preferences, facial expressions and tone of voice.

[0245] Output: User input data and emotion data.

[0246] Step 2: Generate and send a request

[0247] The terminal collects the acquired user input data and emotion data and organizes them into a request object, which includes the user's ideas, preferences, and emotion data.

[0248] The terminal sends a request object to the server.

[0249] Input: User input data and emotion data.

[0250] Output: The request object that is sent to the server.

[0251] Step 3: Receiving and parsing the request

[0252] The server receives the request object.

[0253] The server extracts the user's ideas, preferences, and sentiment data from the request object, checking the data for consistency and preprocessing the data if necessary.

[0254] Input: A request object.

[0255] Output: Parsed ideas, preferences, and sentiment data.

[0256] Step 4: Generate the script

[0257] The server invokes the generative AI model and generates a script using the parsed data as input.

[0258] The generative AI model outputs a specific video script containing detailed scenes, narration, musical details, and emotional tone based on user input and emotional data.

[0259] Input: Parsed idea, preference, and sentiment data.

[0260] Output: The generated video script.

[0261] Step 5: Generate the footage

[0262] The server's video generation engine creates video files based on the generated script, and performs processes such as rendering scenes and integrating music and narration.

[0263] The image generation engine follows the instructions of the script and performs the necessary data calculations and processing to output high-quality images.

[0264] Input: The generated video script.

[0265] Output: The generated video file.

[0266] Step 6: Storing the footage and issuing an identifier

[0267] The server stores the completed video files in a database, a process that may also involve compression and format conversion of the video files.

[0268] The server issues a unique identifier for the stored video file and returns this identifier to the user, allowing them to access the video at a later time.

[0269] Input: The generated video file.

[0270] Output: The file stored in the database and the issued identifier.

[0271] (Application example 2)

[0272] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0273] Modern virtual stores face the challenge of generating and displaying promotional videos in real time that reflect users' interests and emotions. Furthermore, there is no established method for reflecting user emotions in the generation of such videos, making it difficult to customize the user experience with conventional content generation systems. Furthermore, while video file management and high-speed display are required, achieving these goals requires advanced technology.

[0274] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0275] In this invention, the server includes means for generating a video script based on the user's linguistic instructions and recognized emotions, a video generation engine for creating a video file based on the generated video script, and a database for saving the generated video file and issuing an identifier for the saving destination, thereby enabling the real-time generation and display of promotional videos according to the user's interests and emotions.

[0276] A "generative artificial intelligence model" is an artificial intelligence model for generating a video script based on a user's instructions and emotions.

[0277] The "video generation engine" is an engine that creates a video file based on the generated video script.

[0278] The "database" is a system that stores the generated video files and issues identifiers for the storage locations.

[0279] The "emotion engine" is a system that analyzes the user's facial expressions and tone of voice to obtain emotional data.

[0280] The "video display means" is a means for generating promotional video relating to a designated product in real time and showing it on the spot.

[0281] A "user interface" is an interface through which a user inputs instructions and preferences into a system.

[0282] "Scene rendering" is the computational process for generating multiple video scenes.

[0283] "Integrated music" means selecting music that is appropriate for the video scene and playing it in an integrated manner as a whole.

[0284] This invention is a system that combines a generative AI model and an emotion engine to generate and display videos in real time based on user instructions and emotions. This system enables users to automatically generate promotional videos and display them on the spot while browsing products in a virtual store.

[0285] Hardware and Software Configuration

[0286] The system includes the following main hardware and software:

[0287] Hardware:

[0288] 1. Camera: Used to capture the user's facial expressions.

[0289] 2. Microphone: Used to capture the tone of the user's voice.

[0290] 3. Smartphone or head-mounted display (HMD): Used to run applications and play videos.

[0291] software:

[0292] 1. OpenCV: Analyzes camera footage and extracts facial expression data.

[0293] 2. TensorFlow: Used to build and run emotion engines and generative AI models.

[0294] 3. JSON: Used for data exchange and storage.

[0295] 4. FFmpeg: A library for creating footage based on the generated script.

[0296] Overall system processing flow

[0297] 1. User Interface: The user selects products in the virtual store and inputs the promotion theme and preferences. For example, assume that the user prefers "up-tempo music" with a theme of "adventure with space travel images."

[0298] 2. Emotion engine: The camera and microphone are used to capture the user's facial expressions and tone of voice, which are then analyzed to obtain emotional data. For example, if the user is excited, this data is input into the system.

[0299] 3. Generative AI model: Generate a video script based on the user's input of themes, preferences, and emotional data. Generate the script using a generative AI model (e.g., GPT-3).

[0300] 4. Video Generation Engine: Generates high-quality video based on the generated script. This process includes scene rendering and music integration.

[0301] 5. Database: Stores the generated video files and issues identifiers that users can use to access the videos.

[0302] 6. Video display means: The generated video is played in real time on a smartphone or HMD. For example, a promotional video of an "adventure inspired by space travel" is played in a virtual store, which is expected to arouse users' interest in the product.

[0303] Specific examples

[0304] For example, suppose a user requests an "adventure"-themed video with an excited expression while browsing products in a virtual store. In this case, the emotion engine analyzes the user's excited facial expression and tone of voice, and provides the prompt "an adventure with the image of space travel" as input data to the generative AI model (GPT-3). The video generation engine generates a video based on the script, saves it in a database, and issues an identifier. Finally, the video is played in real time on a smartphone or HMD.

[0305] Prompt Sentence Examples

[0306] "I'd like some exciting, up-tempo music that evokes the image of space travel and adventure."

[0307] This system makes it possible to provide promotional videos in real time based on the user's emotions and individual preferences.

[0308] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0309] Step 1:

[0310] Users access the virtual store via their smartphone or HMD and select products. At this time, the user inputs the theme and preferences of the promotional video. For example, if a user inputs the instructions "adventure with a space travel image" and "up-tempo music," this data is sent to the system.

[0311] Step 2:

[0312] The device uses a camera and microphone to capture the user's facial expressions and tone of voice in real time, which then generates emotion data. For example, if the user is excited, their facial expressions and tone of voice are analyzed and the emotion data for "excitement" is generated.

[0313] Step 3:

[0314] The server receives user-entered themes, preferences, and emotional data. This data is organized in JSON format and prepared as input for the generative AI model. Examples of input data include "adventure with space travel images," "up-tempo music," and "excitement."

[0315] Step 4:

[0316] The server generates a video script using a generative AI model. Specifically, a generative AI model using TensorFlow (e.g., GPT-3) analyzes the input data and outputs a video script based on the theme of "an adventure with the image of space travel" and the emotions of "up-tempo music" and "excitement."

[0317] Step 5:

[0318] The server's video generation engine creates video files based on the generated script, using software such as FFmpeg to render the scene and integrate music. A high-quality video file is generated as the output.

[0319] Step 6:

[0320] The server stores the generated video file in a database and issues a file identifier, which is required for users to access the video later and ensures that the video file is properly managed in the database.

[0321] Step 7:

[0322] The device receives the identifier issued by the server and plays the video generated in real time. While the user browses products in the virtual store, a promotional video for an "adventure inspired by space travel" is played, providing an experience based on the user's interests and emotions.

[0323] The above steps will realize a system that generates and displays promotional videos in real time that reflect the user's preferences and emotions.

[0324] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0325] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0326] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0327] [Second embodiment]

[0328] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0329] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0330] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0331] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0332] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0333] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0334] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0335] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0336] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0337] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0338] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0339] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0340] An embodiment of the present invention is described below: The system uses a generative AI model to generate a video script based on a user's linguistic instructions, creates and saves a video file based on the generated video script, and provides the user with an identifier for the saved file.

[0341] 1. Creating and Sending a Request

[0342] The user inputs their ideas into the interface. For example, they specify an idea for an "adventure with a space travel image" or a preference for "upbeat, up-tempo music." This information is converted into a request object by the terminal and sent to the server.

[0343] 2. Receiving and parsing the request

[0344] The server receives requests from users and extracts information such as their ideas and preferences, which are then prepared as input data for a generative AI model.

[0345] 3. Generate the script

[0346] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences, detailing the video's scenes, narration, music, and more.

[0347] 4. Image Generation

[0348] The server's video generation engine creates the video based on the generated script, including rendering the scene and integrating music. The generation engine creates high-quality video files according to the instructions in the script.

[0349] 5. Storage of footage and issuance of identifiers

[0350] The server saves the created video file in a database. Once saving is complete, the server generates an identifier for the storage location and returns it to the user. This identifier allows the user to access and view the created video.

[0351] Specific examples

[0352] For example, consider a user who wants to generate a video with the theme of "an adventure with the image of space travel." The user inputs the theme and preferences into the interface and submits a request. The server receives this request and uses a generative AI model to generate a script based on the theme of "an adventure with the image of space travel." The server's video generation engine then creates a video based on the script and stores it in a database. Finally, the user can watch their video using the identifier returned by the server.

[0353] The processing flow will be explained below.

[0354] Step 1:

[0355] The user inputs their ideas and preferences into the interface. For example, they specify the idea of ​​"an adventure with the image of space travel" and the preference of "upbeat, up-tempo music."

[0356] Step 2:

[0357] The device generates a request object (user_request) based on the user's input, which contains information about the user's ideas and preferences.

[0358] Step 3:

[0359] The terminal sends the generated request object to the server, which contains data based on the user's input.

[0360] Step 4:

[0361] The server receives the user's request and analyzes it, which includes extracting information about the user's ideas and preferences.

[0362] Step 5:

[0363] Based on the extracted information, the server prepares input data (ai_input) for the generative AI model, which contains detailed information about the user's ideas and preferences.

[0364] Step 6:

[0365] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences. The AI ​​model analyzes the input data and outputs the script.

[0366] Step 7:

[0367] The server's video generation engine creates the video based on the generated script, a process that includes rendering the scene and integrating music.

[0368] Step 8:

[0369] The server saves the completed video file in a database. During this saving process, an identifier is issued to ensure that the video file is properly managed.

[0370] Step 9:

[0371] The server returns the identifier of the video stored in the database to the user, allowing the user to access the generated video.

[0372] Step 10:

[0373] The user uses the returned identifier to view the video on their device, which allows them to stream or download their video content.

[0374] Example 1

[0375] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0376] Conventional video generation systems require users to have advanced skills and specialized knowledge to create specific video scripts, making them difficult for average users to use. Furthermore, the process of manually creating a script and then generating a video file is time-consuming, labor-intensive, and inefficient.

[0377] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0378] In this invention, the server includes: means for a user to input ideas and preferences as linguistic instructions into an interface; means for converting the user input into a request object and sending the request to the server; means for the server to receive the request and extract information on the ideas and preferences; means for generating a video script based on the user's ideas and preferences using a generative AI model; means including a video generation engine that creates a video file based on the generated video script; and means for saving the generated video file in a database and issuing an identifier for the storage destination. This allows users without advanced skills or specialized knowledge to easily generate video files based on linguistic instructions and efficiently use the videos.

[0379] A "user" is a person who uses the system to input verbal instructions into the interface to generate a video file.

[0380] An "interface" is an input device or software that allows a user to input their ideas and preferences.

[0381] A "request object" is a collection of information that is converted from user input into a data format and sent to the server.

[0382] "Server" means a computer system that receives and analyzes a request object from a user, generates a video script using a generative AI model, and then generates and stores the video file.

[0383] A "generative AI model" is an artificial intelligence model for generating video scripts based on a user's ideas and preferences.

[0384] A "video script" is a document that describes detailed instructions for scenes, narration, music, etc. for generating a video file.

[0385] A "video generation engine" is a device or software that creates a video file based on a video script.

[0386] A "database" is a system that stores generated video files and manages information for later access.

[0387] An "identifier" is a character string or number that uniquely identifies the generated video file and returns information to the user.

[0388] This embodiment of the present invention is a system that generates a video script based on a user's verbal instructions, and then creates and saves a video file based on the script. This system consists of three main components: a user, a terminal, and a server.

[0389] First, the user inputs their ideas and preferences into an interface, which can be implemented as a computer or smartphone application. For example, specific instructions such as "an adventure with a space travel image" or "upbeat, upbeat music" can be entered in text format.

[0390] The device then converts the user's input into a request object and sends it to the server, which includes the user's specified themes and preferences, allowing the server to convey the user's request in the appropriate data format.

[0391] The server receives the request sent from the device and analyzes the ideas and preferences contained therein. The analyzed information is prepared as input data for a generative AI model, such as OpenAI GPT-4.

[0392] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences. The script contains detailed instructions for the video's scenes, narration, music, etc. After the script is generated, the server's video generation engine creates a video file based on the script.

[0393] The video generation engine uses software such as Adobe After Effects to render the scenes and integrate the music. The generated video files are of high quality and are based on the detailed instructions in the script.

[0394] Finally, the server stores the generated video file in a database. The database uses a high-speed storage system to efficiently manage the video files. Once the storage is complete, the server generates a unique identifier for the video file and returns this identifier to the user. The user can use this identifier to access and watch the generated video file.

[0395] As a concrete example, consider a case where a user wants to generate a video with the theme of "adventure with a space travel image" and an upbeat, up-tempo music. The user inputs the theme and preferences into the interface and sends the request to the server. The server analyzes the request and generates a video script using a generative AI model. The server's video generation engine then creates a video based on the script and stores it in a database. Finally, the user can watch their video using the identifier returned by the server.

[0396] Examples of prompt sentences include the following:

[0397] Please create a video with the theme of "An adventure with the image of space travel." I would like to use upbeat music with a bright atmosphere.

[0398] In this way, by using this system, users can easily generate and manage high-quality video based on their own ideas and preferences, without needing advanced specialized knowledge.

[0399] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0400] Step 1: User Input

[0401] The user inputs their ideas and preferences for the video into the interface. For example, they can input a theme such as "an adventure with a space travel image" or preferences such as "a bright atmosphere and up-tempo music." The input data is sent to the terminal as text data.

[0402] Step 2: Create a request object

[0403] The terminal converts textual input data from the user into a request object, which contains themes and preferences as key-value pairs. For example, a request object might contain the following data:

[0404] json

[0405] {

[0406] "theme": "Space travel-inspired adventure",

[0407] "mood": "upbeat, upbeat music"

[0408] }

[0409] The created request object is sent to the server.

[0410] Step 3: Receiving and Parsing the Request

[0411] The server receives the request object sent from the device, extracts idea and preference information from the request object, and prepares the extracted data as input data for the generative AI model.

[0412] Step 4: Input to the generative AI model

[0413] The server inputs the extracted ideas and preferences into a generative AI model, which generates a prompt like this:

[0414] Create a video with the theme of "An adventure with the image of space travel." Use upbeat, up-tempo music.

[0415] Input data is transformed into a generative AI model.

[0416] Step 5: Generate a video script

[0417] The server invokes the generative AI model to generate a video script based on the input data. The generated script contains details such as video scenes, narration, and music. For example, the following script might be generated:

[0418] json

[0419] {

[0420] "scenes": [

[0421] {"description": "A spaceship traveling through the galaxy", "duration": "30 seconds"},

[0422] {"description": "Spaceship landing scene", "duration": "20 seconds"}

[0423] ],

[0424] "narration": "The space travel adventure begins...",

[0425] "music": "Uptempo background music"

[0426] }

[0427] Step 6: Rendering the footage

[0428] The server's video generation engine creates the video based on the generated script. First, the video generation engine renders the scene. This process includes generating CG and editing existing footage. Next, music and narration are integrated, and the final video file is generated.

[0429] Step 7: Save the video file

[0430] The server stores the generated video files in a database, a process that uses a high-speed storage system.

[0431] Step 8: Generate and return identifiers

[0432] The server generates a unique identifier for each video file that has completed the storage process. This identifier will have the following format:

[0433] json

[0434] {

[0435] "identifier": "abcd1234"

[0436] }

[0437] The identifier is returned to the user, who can then use it to access and view the generated video file.

[0438] The above is the flow of processing of the program of this system.

[0439] (Application example 1)

[0440] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0441] Conventional video generation systems lack the ability to quickly generate and instantly distribute high-quality videos based on user ideas. In particular, there has been no system that allows users to easily input ideas using a smartphone and then view and share the videos generated on the spot. This has made it difficult to meet users' demands for creative content generation and instant distribution.

[0442] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0443] In this invention, the server includes: means for generating a video script based on a user's linguistic instructions using a generative artificial intelligence model; a video generation engine; a database; means for a user to input ideas and preferences using a smartphone; means for formatting the input data as input data for the generative AI model; and means for providing the user with a preview and sharing of the generated video file, thereby enabling the user to generate high-quality videos in real time using a smartphone and instantly share the videos on various platforms.

[0444] A "generative artificial intelligence model" is an algorithm that analyzes a user's linguistic instructions and automatically generates a video script based on that information.

[0445] A "video script" is a document that details the video's scenes, narration, music, etc.

[0446] A "video generation engine" is software or hardware that generates the actual video file based on a video script.

[0447] A "database" is an information management system that stores the generated video files and allows users to access them.

[0448] An "identifier" is an ID or code that uniquely identifies a specific video file within a database.

[0449] A "smartphone" is a multifunctional mobile device that combines mobile communication and computer functions.

[0450] "Formatting" is the process of converting input data into a specific format that is easy for a generative AI model to understand.

[0451] "Preview" is a function that displays a part or all of a generated video file in advance before the user finally views it.

[0452] "Sharing" is the process of sharing the generated video file with other users and platforms.

[0453] The system embodying this invention allows users to easily input their ideas and preferences using a smartphone, and then generate, view, and share high-quality video on the spot.

[0454] First, the user uses a dedicated smartphone application to input their ideas and preferences for the video they want to create, including a specific prompt such as "An adventure story exploring the mysteries of space in a light-hearted atmosphere." The application then formats the data entered by the user into JSON format and sends it to the server.

[0455] The server generates a video script using a generative AI model based on the received request. The generative AI model analyzes the user's linguistic instructions and creates a video script that details the development of scenes, narration content, music selection, etc. This script is then converted into an actual video file by the video generation engine on the server.

[0456] The video generation engine renders multiple scenes based on the generated video script and integrates appropriate music. The server stores the rendered and music-integrated video files in a database and issues an identifier for the storage location. This identifier is used by users to access and watch the generated video files.

[0457] Finally, users receive the issued identifier through a smartphone application and can preview the video generated based on their input idea. This video can be instantly shared within the application or on platforms such as social media.

[0458] The primary hardware used includes smartphones (e.g., iPhone, Android smartphone) and servers (e.g., AWS EC2 instances).The software used includes smartphone applications (e.g., developed with React Native), server-side generative AI models (e.g., GPT-4), a Flask-based backend, and a video generation engine (e.g., FFmpeg).

[0459] In this way, users can easily generate high-quality videos and share them on a variety of platforms. An example of a specific prompt is "An adventure story exploring the mysteries of space in a cheerful atmosphere." Based on this prompt, the generative AI model generates a video script that meets the user's needs, and high-quality videos are automatically generated based on that script.

[0460] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0461] Step 1:

[0462] Users launch a dedicated smartphone application and input their ideas and preferences for the video they want to create into the interface, typically entering prompts in text format such as "An adventure story exploring the mysteries of space in a cheerful atmosphere." The input data is then formatted in JSON format within the application.

[0463] Input: User's ideas and preferences (prompt text)

[0464] Output: Request data in JSON format

[0465] Step 2:

[0466] The device sends formatted JSON data to the server, and the request contains the user's ideas and preferences.

[0467] Input: Request data in JSON format

[0468] Output: Request sent to server

[0469] Step 3:

[0470] The server analyzes the received request and prepares the prompt sentence as input data for the generative AI model. The server then invokes the generative AI model to generate a video script based on the user's linguistic instructions.

[0471] Input: Request data in JSON format

[0472] Output: Generated video script

[0473] Step 4:

[0474] The server's video generation engine creates video files based on the generated video script, rendering the scenes and integrating music.

[0475] Input: Generated video script

[0476] Output: Video file

[0477] Step 5:

[0478] The server stores the generated video file in a database and issues an identifier for the storage location, which is used by users to access the generated video.

[0479] Input: Video file

[0480] Output: Destination identifier

[0481] Step 6:

[0482] The server returns the storage location identifier to the smartphone application, allowing the user to check the video generated based on their own idea.

[0483] Input: Destination identifier

[0484] Output: Send identifier to smartphone

[0485] Step 7:

[0486] Users can use a smartphone application to preview and watch the generated video, and can also share the video within the application or on platforms such as social media.

[0487] Input: Destination identifier

[0488] Output: Preview and share your footage

[0489] This will enable high-quality videos to be quickly generated, viewed, and shared based on user-input ideas and preferences.

[0490] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0491] An embodiment of the present invention is described below: The system uses a generative AI model and an emotion engine to generate a video script based on a user's linguistic instructions and recognized emotions, and creates and saves a video file based on the video script.

[0492] 1. Creating and Sending a Request

[0493] The user inputs their ideas and preferences into the interface. For example, they specify the idea of ​​"an adventure with a space travel image" and their preference of "upbeat, upbeat music." At this point, the emotion engine is activated to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to obtain emotional data. This information is converted into a request object by the terminal and sent to the server.

[0494] 2. Receiving and parsing the request

[0495] The server receives user requests and emotional data, extracts information on ideas, preferences, and recognized emotions, and prepares this information as input to a generative AI model.

[0496] 3. Generate the script

[0497] The server invokes a generative AI model to generate a video script based on the user's ideas, preferences, and emotions. The generative AI model analyzes these inputs and outputs a script that reflects the video's scenes, narration, musical details, and emotional tone.

[0498] 4. Image Generation

[0499] The server's video generation engine creates the video based on the generated script. This process includes rendering the scene and integrating music. The video generation engine creates high-quality video files according to the instructions in the script.

[0500] 5. Storage of footage and issuance of identifiers

[0501] The server stores the completed video file in a database. During this storage process, an identifier is issued to ensure the video file is properly managed.

[0502] Specific examples

[0503] For example, consider a user who wants to create a video with the theme of "an adventure with the image of space travel." The user inputs the theme and preferences into the interface, while the emotion engine recognizes the user's excited facial expression and tone of voice. The server receives this request and emotion data and generates a script using a generative AI model. The script reflects the theme of "an adventure with the image of space travel" and the excited tone. The video generation engine then generates a video based on the script and stores it in a database. Finally, the user can view their video using an identifier returned by the server. This identifier allows the user to access the generated video through a web browser or application.

[0504] The processing flow will be explained below.

[0505] Step 1:

[0506] The user inputs their ideas and preferences into the user interface. For example, they might specify a theme such as "an adventure with a space travel image" and preferences such as "a cheerful atmosphere and up-tempo music." The emotion engine also analyzes the user's facial expressions and voice via a camera and microphone built into the user interface.

[0507] Step 2:

[0508] The device receives the user's input data and the emotion data from the emotion engine, and combines them to generate a request object (user_request), which contains all the themes, preferences, and emotion data.

[0509] Step 3:

[0510] The terminal sends the generated request object to the server, which includes the user's ideas, preferences, and emotion data.

[0511] Step 4:

[0512] The server receives the request object from the user and analyzes its contents, which includes extracting the user's themes, preferences, and sentiment data.

[0513] Step 5:

[0514] Based on the extracted information, the server prepares input data (ai_input) for the generative AI model, which includes the user's themes, preferences, and emotional data.

[0515] Step 6:

[0516] The server invokes a generative AI model to generate a video script based on the user's themes, preferences, and emotions. The generative AI model analyzes these input data and outputs a video script, which includes video scenes, narration, music, and a tone that reflects the user's emotions.

[0517] Step 7:

[0518] The server's video generation engine creates the video based on the generated script. This process includes rendering the scenes described in the script and integrating music. The video generation engine creates high-quality video files according to the instructions in the script.

[0519] Step 8:

[0520] The server stores the completed video file in a database. During this storage process, an identifier is issued to ensure the video file is properly managed.

[0521] Step 9:

[0522] The server returns to the user an identifier for the video stored in the database, allowing the user to access the generated video.

[0523] Step 10:

[0524] The user uses the returned identifier to watch the video on their device, which allows them to stream or download their video content through a web browser or dedicated app.

[0525] Example 2

[0526] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0527] Conventional video generation systems generate scripts based solely on the user's verbal instructions, making it difficult to reflect the user's emotions and detailed preferences. Furthermore, when the management of generated video files and the issuance of identifiers are done manually, this increases the administrative workload and increases the likelihood of errors.

[0528] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for generating a video script based on a user's linguistic instructions and recognized emotions using a generative AI model, a video generation engine means for creating a video file based on the generated video script, a database means for saving the generated video file and issuing an identifier for the saving destination, and means for analyzing the user's input data and emotion data, generating it as a request object, and transmitting it to the server. This enables video generation that reflects the user's emotions and detailed preferences, and further enables efficient automatic management of the generated video files and the issuance of identifiers.

[0529] A "generative artificial intelligence model" is a model that uses artificial intelligence technology to generate video scripts and other content based on user input data and emotional data.

[0530] A "video script" is a set of instructions that describes the scene structure, narration, musical details, and emotional tone for creating a video.

[0531] A "video generation engine" is a software or hardware system for creating the actual video based on the generated video script.

[0532] "Database" refers to a storage device and associated management system for storing and managing generated video files.

[0533] A "request object" is a data structure that includes user input data and emotion data and is generated to be sent to a server.

[0534] The "emotion engine" is a system that analyzes the user's facial expressions and tone of voice, recognizes emotions, and generates and provides that data.

[0535] An "identifier" is a unique code or number that uniquely identifies the generated video file.

[0536] "Input data" is data, including ideas and preferences, that a user provides to the system.

[0537] An embodiment of the present invention is a system for generating a video script based on a user's linguistic instructions and recognized emotions using a generative AI model and an emotion engine, and creating and saving a high-quality video file based on the script. The system includes a server, a terminal, and various user interfaces.

[0538] The user interface (e.g., a PC or smartphone application) provides a means for the user to input their ideas and preferences. The emotion engine is used to analyze the user's facial expressions and tone of voice to obtain emotion data. This emotion data, along with the user's input data, is packaged into a request object and sent from the device to the server.

[0539] The server receives the user's request object and analyzes the user's ideas, preferences, and emotional data using a video script generation method based on a generative AI model. The generative AI model is often operated using cloud computing resources. Based on the input data, the model generates a specific video script including scene composition, narration, musical details, and emotional tone.

[0540] The generated video script is then converted into a video file by a server-based video generation engine, which consists of advanced software and hardware for scene rendering and music integration. The video generation engine performs the necessary data calculations and processing according to the instructions to create a high-quality video file.

[0541] The generated video file is stored in a database on the server. During the storage process, an identifier is generated to uniquely identify the video file. This identifier allows users to access the completed video file. The identifier is used in web browsers and applications, providing users with an easy way to view the video.

[0542] Specific examples

[0543] For example, consider a case where a user wants to generate a video with the theme of "adventure with the image of space travel." The user inputs the theme and preferences (e.g., "upbeat atmosphere, up-tempo music") into the interface, and the emotion engine recognizes the user's excited facial expression and tone of voice. Next, the user's input data and emotion data are packaged into a request object and sent to the server.

[0544] The server receives this request and generates a script using a generative AI model. The generative AI model outputs a script that reflects an exciting tone in line with the theme of "an adventure with the image of space travel." Based on this script, a video generation engine renders the scene and integrates music and narration to create a video file. Finally, the server stores the video file in a database and returns an identifier to the user. The user can use this identifier to watch the video they generated.

[0545] Prompt Sentence Examples

[0546] "The theme specified by the user is 'an adventure with the image of space travel,' and their preference is 'a cheerful atmosphere, up-tempo music.' We also recognized an excited tone as the user's emotion. Please generate a video script based on this."

[0547] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0548] Step 1: User Input and Emotion Recognition

[0549] The user uses the interface to input their ideas and preferences. For example, the user may specify a theme of "adventure with space travel imagery" and a preference for "upbeat, upbeat music."

[0550] The emotion engine analyzes the user's facial expressions and tone of voice in real time to obtain emotion data. For example, if the user has an excited expression, that data is obtained by the emotion engine.

[0551] Input: Your ideas, preferences, facial expressions and tone of voice.

[0552] Output: User input data and emotion data.

[0553] Step 2: Generate and send a request

[0554] The terminal collects the acquired user input data and emotion data and organizes them into a request object, which includes the user's ideas, preferences, and emotion data.

[0555] The terminal sends a request object to the server.

[0556] Input: User input data and emotion data.

[0557] Output: The request object that is sent to the server.

[0558] Step 3: Receiving and parsing the request

[0559] The server receives the request object.

[0560] The server extracts the user's ideas, preferences, and sentiment data from the request object, checking the data for consistency and preprocessing the data if necessary.

[0561] Input: A request object.

[0562] Output: Parsed ideas, preferences, and sentiment data.

[0563] Step 4: Generate the script

[0564] The server invokes the generative AI model and generates a script using the parsed data as input.

[0565] The generative AI model outputs a specific video script containing detailed scenes, narration, musical details, and emotional tone based on user input and emotional data.

[0566] Input: Parsed idea, preference, and sentiment data.

[0567] Output: The generated video script.

[0568] Step 5: Generate the footage

[0569] The server's video generation engine creates video files based on the generated script, and performs processes such as rendering scenes and integrating music and narration.

[0570] The image generation engine follows the instructions of the script and performs the necessary data calculations and processing to output high-quality images.

[0571] Input: The generated video script.

[0572] Output: The generated video file.

[0573] Step 6: Storing the footage and issuing an identifier

[0574] The server stores the completed video files in a database, a process that may also involve compression and format conversion of the video files.

[0575] The server issues a unique identifier for the stored video file and returns this identifier to the user, allowing them to access the video at a later time.

[0576] Input: The generated video file.

[0577] Output: The file stored in the database and the issued identifier.

[0578] (Application example 2)

[0579] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0580] Modern virtual stores face the challenge of generating and displaying promotional videos in real time that reflect users' interests and emotions. Furthermore, there is no established method for reflecting user emotions in the generation of such videos, making it difficult to customize the user experience with conventional content generation systems. Furthermore, while video file management and high-speed display are required, achieving these goals requires advanced technology.

[0581] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0582] In this invention, the server includes means for generating a video script based on the user's linguistic instructions and recognized emotions, a video generation engine for creating a video file based on the generated video script, and a database for saving the generated video file and issuing an identifier for the saving destination, thereby enabling the real-time generation and display of promotional videos according to the user's interests and emotions.

[0583] A "generative artificial intelligence model" is an artificial intelligence model for generating a video script based on a user's instructions and emotions.

[0584] The "video generation engine" is an engine that creates a video file based on the generated video script.

[0585] The "database" is a system that stores the generated video files and issues identifiers for the storage locations.

[0586] The "emotion engine" is a system that analyzes the user's facial expressions and tone of voice to obtain emotional data.

[0587] The "video display means" is a means for generating promotional video relating to a designated product in real time and showing it on the spot.

[0588] A "user interface" is an interface through which a user inputs instructions and preferences into a system.

[0589] "Scene rendering" is the computational process for generating multiple video scenes.

[0590] "Integrated music" means selecting music that is appropriate for the video scene and playing it in an integrated manner as a whole.

[0591] This invention is a system that combines a generative AI model and an emotion engine to generate and display videos in real time based on user instructions and emotions. This system enables users to automatically generate promotional videos and display them on the spot while browsing products in a virtual store.

[0592] Hardware and Software Configuration

[0593] The system includes the following main hardware and software:

[0594] Hardware:

[0595] 1. Camera: Used to capture the user's facial expressions.

[0596] 2. Microphone: Used to capture the tone of the user's voice.

[0597] 3. Smartphone or head-mounted display (HMD): Used to run applications and play videos.

[0598] software:

[0599] 1. OpenCV: Analyzes camera footage and extracts facial expression data.

[0600] 2. TensorFlow: Used to build and run emotion engines and generative AI models.

[0601] 3. JSON: Used for data exchange and storage.

[0602] 4. FFmpeg: A library for creating footage based on the generated script.

[0603] Overall system processing flow

[0604] 1. User Interface: The user selects products in the virtual store and inputs the promotion theme and preferences. For example, assume that the user prefers "up-tempo music" with a theme of "adventure with space travel images."

[0605] 2. Emotion engine: The camera and microphone are used to capture the user's facial expressions and tone of voice, which are then analyzed to obtain emotional data. For example, if the user is excited, this data is input into the system.

[0606] 3. Generative AI model: Generate a video script based on the user's input of themes, preferences, and emotional data. Generate the script using a generative AI model (e.g., GPT-3).

[0607] 4. Video Generation Engine: Generates high-quality video based on the generated script. This process includes scene rendering and music integration.

[0608] 5. Database: Stores the generated video files and issues identifiers that users can use to access the videos.

[0609] 6. Video display means: The generated video is played in real time on a smartphone or HMD. For example, a promotional video of an "adventure inspired by space travel" is played in a virtual store, which is expected to arouse users' interest in the product.

[0610] Specific examples

[0611] For example, suppose a user requests an "adventure"-themed video with an excited expression while browsing products in a virtual store. In this case, the emotion engine analyzes the user's excited facial expression and tone of voice, and provides the prompt "an adventure with the image of space travel" as input data to the generative AI model (GPT-3). The video generation engine generates a video based on the script, saves it in a database, and issues an identifier. Finally, the video is played in real time on a smartphone or HMD.

[0612] Prompt Sentence Examples

[0613] "I'd like some exciting, up-tempo music that evokes the image of space travel and adventure."

[0614] This system makes it possible to provide promotional videos in real time based on the user's emotions and individual preferences.

[0615] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0616] Step 1:

[0617] Users access the virtual store via their smartphone or HMD and select products. At this time, the user inputs the theme and preferences of the promotional video. For example, if a user inputs the instructions "adventure with a space travel image" and "up-tempo music," this data is sent to the system.

[0618] Step 2:

[0619] The device uses a camera and microphone to capture the user's facial expressions and tone of voice in real time, which then generates emotion data. For example, if the user is excited, their facial expressions and tone of voice are analyzed and the emotion data for "excitement" is generated.

[0620] Step 3:

[0621] The server receives user-entered themes, preferences, and emotional data. This data is organized in JSON format and prepared as input for the generative AI model. Examples of input data include "adventure with space travel images," "up-tempo music," and "excitement."

[0622] Step 4:

[0623] The server generates a video script using a generative AI model. Specifically, a generative AI model using TensorFlow (e.g., GPT-3) analyzes the input data and outputs a video script based on the theme of "an adventure with the image of space travel" and the emotions of "up-tempo music" and "excitement."

[0624] Step 5:

[0625] The server's video generation engine creates video files based on the generated script, using software such as FFmpeg to render the scene and integrate music. A high-quality video file is generated as the output.

[0626] Step 6:

[0627] The server stores the generated video file in a database and issues a file identifier, which is required for users to access the video later and ensures that the video file is properly managed in the database.

[0628] Step 7:

[0629] The device receives the identifier issued by the server and plays the video generated in real time. While the user browses products in the virtual store, a promotional video for an "adventure inspired by space travel" is played, providing an experience based on the user's interests and emotions.

[0630] The above steps will realize a system that generates and displays promotional videos in real time that reflect the user's preferences and emotions.

[0631] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0632] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0633] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0634] [Third embodiment]

[0635] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0636] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0637] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0638] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0639] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0640] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0641] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0642] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0643] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0644] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0645] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0646] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0647] An embodiment of the present invention is described below: The system uses a generative AI model to generate a video script based on a user's linguistic instructions, creates and saves a video file based on the generated video script, and provides the user with an identifier for the saved file.

[0648] 1. Creating and Sending a Request

[0649] The user inputs their ideas into the interface. For example, they specify an idea for an "adventure with a space travel image" or a preference for "upbeat, up-tempo music." This information is converted into a request object by the terminal and sent to the server.

[0650] 2. Receiving and parsing the request

[0651] The server receives requests from users and extracts information such as their ideas and preferences, which are then prepared as input data for a generative AI model.

[0652] 3. Generate the script

[0653] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences, detailing the video's scenes, narration, music, and more.

[0654] 4. Image Generation

[0655] The server's video generation engine creates the video based on the generated script, including rendering the scene and integrating music. The generation engine creates high-quality video files according to the instructions in the script.

[0656] 5. Storage of footage and issuance of identifiers

[0657] The server saves the created video file in a database. Once saving is complete, the server generates an identifier for the storage location and returns it to the user. This identifier allows the user to access and view the created video.

[0658] Specific examples

[0659] For example, consider a user who wants to generate a video with the theme of "an adventure with the image of space travel." The user inputs the theme and preferences into the interface and submits a request. The server receives this request and uses a generative AI model to generate a script based on the theme of "an adventure with the image of space travel." The server's video generation engine then creates a video based on the script and stores it in a database. Finally, the user can watch their video using the identifier returned by the server.

[0660] The processing flow will be explained below.

[0661] Step 1:

[0662] The user inputs their ideas and preferences into the interface. For example, they specify the idea of ​​"an adventure with the image of space travel" and the preference of "upbeat, up-tempo music."

[0663] Step 2:

[0664] The device generates a request object (user_request) based on the user's input, which contains information about the user's ideas and preferences.

[0665] Step 3:

[0666] The terminal sends the generated request object to the server, which contains data based on the user's input.

[0667] Step 4:

[0668] The server receives the user's request and analyzes it, which includes extracting information about the user's ideas and preferences.

[0669] Step 5:

[0670] Based on the extracted information, the server prepares input data (ai_input) for the generative AI model, which contains detailed information about the user's ideas and preferences.

[0671] Step 6:

[0672] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences. The AI ​​model analyzes the input data and outputs the script.

[0673] Step 7:

[0674] The server's video generation engine creates the video based on the generated script, a process that includes rendering the scene and integrating music.

[0675] Step 8:

[0676] The server saves the completed video file in a database. During this saving process, an identifier is issued to ensure that the video file is properly managed.

[0677] Step 9:

[0678] The server returns the identifier of the video stored in the database to the user, allowing the user to access the generated video.

[0679] Step 10:

[0680] The user uses the returned identifier to view the video on their device, which allows them to stream or download their video content.

[0681] Example 1

[0682] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0683] Conventional video generation systems require users to have advanced skills and specialized knowledge to create specific video scripts, making them difficult for average users to use. Furthermore, the process of manually creating a script and then generating a video file is time-consuming, labor-intensive, and inefficient.

[0684] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0685] In this invention, the server includes: means for a user to input ideas and preferences as linguistic instructions into an interface; means for converting the user input into a request object and sending the request to the server; means for the server to receive the request and extract information on the ideas and preferences; means for generating a video script based on the user's ideas and preferences using a generative AI model; means including a video generation engine that creates a video file based on the generated video script; and means for saving the generated video file in a database and issuing an identifier for the storage destination. This allows users without advanced skills or specialized knowledge to easily generate video files based on linguistic instructions and efficiently use the videos.

[0686] A "user" is a person who uses the system to input verbal instructions into the interface to generate a video file.

[0687] An "interface" is an input device or software that allows a user to input their ideas and preferences.

[0688] A "request object" is a collection of information that is converted from user input into a data format and sent to the server.

[0689] "Server" means a computer system that receives and analyzes a request object from a user, generates a video script using a generative AI model, and then generates and stores the video file.

[0690] A "generative AI model" is an artificial intelligence model for generating video scripts based on a user's ideas and preferences.

[0691] A "video script" is a document that describes detailed instructions for scenes, narration, music, etc. for generating a video file.

[0692] A "video generation engine" is a device or software that creates a video file based on a video script.

[0693] A "database" is a system that stores generated video files and manages information for later access.

[0694] An "identifier" is a character string or number that uniquely identifies the generated video file and returns information to the user.

[0695] This embodiment of the present invention is a system that generates a video script based on a user's verbal instructions, and then creates and saves a video file based on the script. This system consists of three main components: a user, a terminal, and a server.

[0696] First, the user inputs their ideas and preferences into an interface, which can be implemented as a computer or smartphone application. For example, specific instructions such as "an adventure with a space travel image" or "upbeat, upbeat music" can be entered in text format.

[0697] The device then converts the user's input into a request object and sends it to the server, which includes the user's specified themes and preferences, allowing the server to convey the user's request in the appropriate data format.

[0698] The server receives the request sent from the device and analyzes the ideas and preferences contained therein. The analyzed information is prepared as input data for a generative AI model, such as OpenAI GPT-4.

[0699] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences. The script contains detailed instructions for the video's scenes, narration, music, etc. After the script is generated, the server's video generation engine creates a video file based on the script.

[0700] The video generation engine uses software such as Adobe After Effects to render the scenes and integrate the music. The generated video files are of high quality and are based on the detailed instructions in the script.

[0701] Finally, the server stores the generated video file in a database. The database uses a high-speed storage system to efficiently manage the video files. Once the storage is complete, the server generates a unique identifier for the video file and returns this identifier to the user. The user can use this identifier to access and watch the generated video file.

[0702] As a concrete example, consider a case where a user wants to generate a video with the theme of "adventure with a space travel image" and an upbeat, up-tempo music. The user inputs the theme and preferences into the interface and sends the request to the server. The server analyzes the request and generates a video script using a generative AI model. The server's video generation engine then creates a video based on the script and stores it in a database. Finally, the user can watch their video using the identifier returned by the server.

[0703] Examples of prompt sentences include the following:

[0704] Please create a video with the theme of "An adventure with the image of space travel." I would like to use upbeat music with a bright atmosphere.

[0705] In this way, by using this system, users can easily generate and manage high-quality video based on their own ideas and preferences, without needing advanced specialized knowledge.

[0706] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0707] Step 1: User Input

[0708] The user inputs their ideas and preferences for the video into the interface. For example, they can input a theme such as "an adventure with a space travel image" or preferences such as "a bright atmosphere and up-tempo music." The input data is sent to the terminal as text data.

[0709] Step 2: Create a request object

[0710] The terminal converts textual input data from the user into a request object, which contains themes and preferences as key-value pairs. For example, a request object might contain the following data:

[0711] json

[0712] {

[0713] "theme": "Space travel-inspired adventure",

[0714] "mood": "upbeat, upbeat music"

[0715] }

[0716] The created request object is sent to the server.

[0717] Step 3: Receiving and Parsing the Request

[0718] The server receives the request object sent from the device, extracts idea and preference information from the request object, and prepares the extracted data as input data for the generative AI model.

[0719] Step 4: Input to the generative AI model

[0720] The server inputs the extracted ideas and preferences into a generative AI model, which generates a prompt like this:

[0721] Create a video with the theme of "An adventure with the image of space travel." Use upbeat, up-tempo music.

[0722] Input data is transformed into a generative AI model.

[0723] Step 5: Generate a video script

[0724] The server invokes the generative AI model to generate a video script based on the input data. The generated script contains details such as video scenes, narration, and music. For example, the following script might be generated:

[0725] json

[0726] {

[0727] "scenes": [

[0728] {"description": "A spaceship traveling through the galaxy", "duration": "30 seconds"},

[0729] {"description": "Spaceship landing scene", "duration": "20 seconds"}

[0730] ],

[0731] "narration": "The space travel adventure begins...",

[0732] "music": "Uptempo background music"

[0733] }

[0734] Step 6: Rendering the footage

[0735] The server's video generation engine creates the video based on the generated script. First, the video generation engine renders the scene. This process includes generating CG and editing existing footage. Next, music and narration are integrated, and the final video file is generated.

[0736] Step 7: Save the video file

[0737] The server stores the generated video files in a database, a process that uses a high-speed storage system.

[0738] Step 8: Generate and return identifiers

[0739] The server generates a unique identifier for each video file that has completed the storage process. This identifier will have the following format:

[0740] json

[0741] {

[0742] "identifier": "abcd1234"

[0743] }

[0744] The identifier is returned to the user, who can then use it to access and view the generated video file.

[0745] The above is the flow of processing of the program of this system.

[0746] (Application example 1)

[0747] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0748] Conventional video generation systems lack the ability to quickly generate and instantly distribute high-quality videos based on user ideas. In particular, there has been no system that allows users to easily input ideas using a smartphone and then view and share the videos generated on the spot. This has made it difficult to meet users' demands for creative content generation and instant distribution.

[0749] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0750] In this invention, the server includes: means for generating a video script based on a user's linguistic instructions using a generative artificial intelligence model; a video generation engine; a database; means for a user to input ideas and preferences using a smartphone; means for formatting the input data as input data for the generative AI model; and means for providing the user with a preview and sharing of the generated video file, thereby enabling the user to generate high-quality videos in real time using a smartphone and instantly share the videos on various platforms.

[0751] A "generative artificial intelligence model" is an algorithm that analyzes a user's linguistic instructions and automatically generates a video script based on that information.

[0752] A "video script" is a document that details the video's scenes, narration, music, etc.

[0753] A "video generation engine" is software or hardware that generates the actual video file based on a video script.

[0754] A "database" is an information management system that stores the generated video files and allows users to access them.

[0755] An "identifier" is an ID or code that uniquely identifies a specific video file within a database.

[0756] A "smartphone" is a multifunctional mobile device that combines mobile communication and computer functions.

[0757] "Formatting" is the process of converting input data into a specific format that is easy for a generative AI model to understand.

[0758] "Preview" is a function that displays a part or all of a generated video file in advance before the user finally views it.

[0759] "Sharing" is the process of sharing the generated video file with other users and platforms.

[0760] The system embodying this invention allows users to easily input their ideas and preferences using a smartphone, and then generate, view, and share high-quality video on the spot.

[0761] First, the user uses a dedicated smartphone application to input their ideas and preferences for the video they want to create, including a specific prompt such as "An adventure story exploring the mysteries of space in a light-hearted atmosphere." The application then formats the data entered by the user into JSON format and sends it to the server.

[0762] The server generates a video script using a generative AI model based on the received request. The generative AI model analyzes the user's linguistic instructions and creates a video script that details the development of scenes, narration content, music selection, etc. This script is then converted into an actual video file by the video generation engine on the server.

[0763] The video generation engine renders multiple scenes based on the generated video script and integrates appropriate music. The server stores the rendered and music-integrated video files in a database and issues an identifier for the storage location. This identifier is used by users to access and watch the generated video files.

[0764] Finally, users receive the issued identifier through a smartphone application and can preview the video generated based on their input idea. This video can be instantly shared within the application or on platforms such as social media.

[0765] The primary hardware used includes smartphones (e.g., iPhone, Android smartphone) and servers (e.g., AWS EC2 instances).The software used includes smartphone applications (e.g., developed with React Native), server-side generative AI models (e.g., GPT-4), a Flask-based backend, and a video generation engine (e.g., FFmpeg).

[0766] In this way, users can easily generate high-quality videos and share them on a variety of platforms. An example of a specific prompt is "An adventure story exploring the mysteries of space in a cheerful atmosphere." Based on this prompt, the generative AI model generates a video script that meets the user's needs, and high-quality videos are automatically generated based on that script.

[0767] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0768] Step 1:

[0769] Users launch a dedicated smartphone application and input their ideas and preferences for the video they want to create into the interface, typically entering prompts in text format such as "An adventure story exploring the mysteries of space in a cheerful atmosphere." The input data is then formatted in JSON format within the application.

[0770] Input: User's ideas and preferences (prompt text)

[0771] Output: Request data in JSON format

[0772] Step 2:

[0773] The device sends formatted JSON data to the server, and the request contains the user's ideas and preferences.

[0774] Input: Request data in JSON format

[0775] Output: Request sent to server

[0776] Step 3:

[0777] The server analyzes the received request and prepares the prompt sentence as input data for the generative AI model. The server then invokes the generative AI model to generate a video script based on the user's linguistic instructions.

[0778] Input: Request data in JSON format

[0779] Output: Generated video script

[0780] Step 4:

[0781] The server's video generation engine creates video files based on the generated video script, rendering the scenes and integrating music.

[0782] Input: Generated video script

[0783] Output: Video file

[0784] Step 5:

[0785] The server stores the generated video file in a database and issues an identifier for the storage location, which is used by users to access the generated video.

[0786] Input: Video file

[0787] Output: Destination identifier

[0788] Step 6:

[0789] The server returns the storage location identifier to the smartphone application, allowing the user to check the video generated based on their own idea.

[0790] Input: Destination identifier

[0791] Output: Send identifier to smartphone

[0792] Step 7:

[0793] Users can use a smartphone application to preview and watch the generated video, and can also share the video within the application or on platforms such as social media.

[0794] Input: Destination identifier

[0795] Output: Preview and share your footage

[0796] This will enable high-quality videos to be quickly generated, viewed, and shared based on user-input ideas and preferences.

[0797] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0798] An embodiment of the present invention is described below: The system uses a generative AI model and an emotion engine to generate a video script based on a user's linguistic instructions and recognized emotions, and creates and saves a video file based on the video script.

[0799] 1. Creating and Sending a Request

[0800] The user inputs their ideas and preferences into the interface. For example, they specify the idea of ​​"an adventure with a space travel image" and their preference of "upbeat, upbeat music." At this point, the emotion engine is activated to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to obtain emotional data. This information is converted into a request object by the terminal and sent to the server.

[0801] 2. Receiving and parsing the request

[0802] The server receives user requests and emotional data, extracts information on ideas, preferences, and recognized emotions, and prepares this information as input to a generative AI model.

[0803] 3. Generate the script

[0804] The server invokes a generative AI model to generate a video script based on the user's ideas, preferences, and emotions. The generative AI model analyzes these inputs and outputs a script that reflects the video's scenes, narration, musical details, and emotional tone.

[0805] 4. Image Generation

[0806] The server's video generation engine creates the video based on the generated script. This process includes rendering the scene and integrating music. The video generation engine creates high-quality video files according to the instructions in the script.

[0807] 5. Storage of footage and issuance of identifiers

[0808] The server stores the completed video file in a database. During this storage process, an identifier is issued to ensure the video file is properly managed.

[0809] Specific examples

[0810] For example, consider a user who wants to create a video with the theme of "an adventure with the image of space travel." The user inputs the theme and preferences into the interface, while the emotion engine recognizes the user's excited facial expression and tone of voice. The server receives this request and emotion data and generates a script using a generative AI model. The script reflects the theme of "an adventure with the image of space travel" and the excited tone. The video generation engine then generates a video based on the script and stores it in a database. Finally, the user can view their video using an identifier returned by the server. This identifier allows the user to access the generated video through a web browser or application.

[0811] The processing flow will be explained below.

[0812] Step 1:

[0813] The user inputs their ideas and preferences into the user interface. For example, they might specify a theme such as "an adventure with a space travel image" and preferences such as "a cheerful atmosphere and up-tempo music." The emotion engine also analyzes the user's facial expressions and voice via a camera and microphone built into the user interface.

[0814] Step 2:

[0815] The device receives the user's input data and the emotion data from the emotion engine, and combines them to generate a request object (user_request), which contains all the themes, preferences, and emotion data.

[0816] Step 3:

[0817] The terminal sends the generated request object to the server, which includes the user's ideas, preferences, and emotion data.

[0818] Step 4:

[0819] The server receives the request object from the user and analyzes its contents, which includes extracting the user's themes, preferences, and sentiment data.

[0820] Step 5:

[0821] Based on the extracted information, the server prepares input data (ai_input) for the generative AI model, which includes the user's themes, preferences, and emotional data.

[0822] Step 6:

[0823] The server invokes a generative AI model to generate a video script based on the user's themes, preferences, and emotions. The generative AI model analyzes these input data and outputs a video script, which includes video scenes, narration, music, and a tone that reflects the user's emotions.

[0824] Step 7:

[0825] The server's video generation engine creates the video based on the generated script. This process includes rendering the scenes described in the script and integrating music. The video generation engine creates high-quality video files according to the instructions in the script.

[0826] Step 8:

[0827] The server stores the completed video file in a database. During this storage process, an identifier is issued to ensure the video file is properly managed.

[0828] Step 9:

[0829] The server returns to the user an identifier for the video stored in the database, allowing the user to access the generated video.

[0830] Step 10:

[0831] The user uses the returned identifier to watch the video on their device, which allows them to stream or download their video content through a web browser or dedicated app.

[0832] Example 2

[0833] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0834] Conventional video generation systems generate scripts based solely on the user's verbal instructions, making it difficult to reflect the user's emotions and detailed preferences. Furthermore, when the management of generated video files and the issuance of identifiers are done manually, this increases the administrative workload and increases the likelihood of errors.

[0835] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for generating a video script based on a user's linguistic instructions and recognized emotions using a generative AI model, a video generation engine means for creating a video file based on the generated video script, a database means for saving the generated video file and issuing an identifier for the saving destination, and means for analyzing the user's input data and emotion data, generating it as a request object, and transmitting it to the server. This enables video generation that reflects the user's emotions and detailed preferences, and further enables efficient automatic management of the generated video files and the issuance of identifiers.

[0836] A "generative artificial intelligence model" is a model that uses artificial intelligence technology to generate video scripts and other content based on user input data and emotional data.

[0837] A "video script" is a set of instructions that describes the scene structure, narration, musical details, and emotional tone for creating a video.

[0838] A "video generation engine" is a software or hardware system for creating the actual video based on the generated video script.

[0839] "Database" refers to a storage device and associated management system for storing and managing generated video files.

[0840] A "request object" is a data structure that includes user input data and emotion data and is generated to be sent to a server.

[0841] The "emotion engine" is a system that analyzes the user's facial expressions and tone of voice, recognizes emotions, and generates and provides that data.

[0842] An "identifier" is a unique code or number that uniquely identifies the generated video file.

[0843] "Input data" is data, including ideas and preferences, that a user provides to the system.

[0844] An embodiment of the present invention is a system for generating a video script based on a user's linguistic instructions and recognized emotions using a generative AI model and an emotion engine, and creating and saving a high-quality video file based on the script. The system includes a server, a terminal, and various user interfaces.

[0845] The user interface (e.g., a PC or smartphone application) provides a means for the user to input their ideas and preferences. The emotion engine is used to analyze the user's facial expressions and tone of voice to obtain emotion data. This emotion data, along with the user's input data, is packaged into a request object and sent from the device to the server.

[0846] The server receives the user's request object and analyzes the user's ideas, preferences, and emotional data using a video script generation method based on a generative AI model. The generative AI model is often operated using cloud computing resources. Based on the input data, the model generates a specific video script including scene composition, narration, musical details, and emotional tone.

[0847] The generated video script is then converted into a video file by a server-based video generation engine, which consists of advanced software and hardware for scene rendering and music integration. The video generation engine performs the necessary data calculations and processing according to the instructions to create a high-quality video file.

[0848] The generated video file is stored in a database on the server. During the storage process, an identifier is generated to uniquely identify the video file. This identifier allows users to access the completed video file. The identifier is used in web browsers and applications, providing users with an easy way to view the video.

[0849] Specific examples

[0850] For example, consider a case where a user wants to generate a video with the theme of "adventure with the image of space travel." The user inputs the theme and preferences (e.g., "upbeat atmosphere, up-tempo music") into the interface, and the emotion engine recognizes the user's excited facial expression and tone of voice. Next, the user's input data and emotion data are packaged into a request object and sent to the server.

[0851] The server receives this request and generates a script using a generative AI model. The generative AI model outputs a script that reflects an exciting tone in line with the theme of "an adventure with the image of space travel." Based on this script, a video generation engine renders the scene and integrates music and narration to create a video file. Finally, the server stores the video file in a database and returns an identifier to the user. The user can use this identifier to watch the video they generated.

[0852] Prompt Sentence Examples

[0853] "The theme specified by the user is 'an adventure with the image of space travel,' and their preference is 'a cheerful atmosphere, up-tempo music.' We also recognized an excited tone as the user's emotion. Please generate a video script based on this."

[0854] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0855] Step 1: User Input and Emotion Recognition

[0856] The user uses the interface to input their ideas and preferences. For example, the user may specify a theme of "adventure with space travel imagery" and a preference for "upbeat, upbeat music."

[0857] The emotion engine analyzes the user's facial expressions and tone of voice in real time to obtain emotion data. For example, if the user has an excited expression, that data is obtained by the emotion engine.

[0858] Input: Your ideas, preferences, facial expressions and tone of voice.

[0859] Output: User input data and emotion data.

[0860] Step 2: Generate and send a request

[0861] The terminal collects the acquired user input data and emotion data and organizes them into a request object, which includes the user's ideas, preferences, and emotion data.

[0862] The terminal sends a request object to the server.

[0863] Input: User input data and emotion data.

[0864] Output: The request object that is sent to the server.

[0865] Step 3: Receiving and parsing the request

[0866] The server receives the request object.

[0867] The server extracts the user's ideas, preferences, and sentiment data from the request object, checking the data for consistency and preprocessing the data if necessary.

[0868] Input: A request object.

[0869] Output: Parsed ideas, preferences, and sentiment data.

[0870] Step 4: Generate the script

[0871] The server invokes the generative AI model and generates a script using the parsed data as input.

[0872] The generative AI model outputs a specific video script containing detailed scenes, narration, musical details, and emotional tone based on user input and emotional data.

[0873] Input: Parsed idea, preference, and sentiment data.

[0874] Output: The generated video script.

[0875] Step 5: Generate the footage

[0876] The server's video generation engine creates video files based on the generated script, and performs processes such as rendering scenes and integrating music and narration.

[0877] The image generation engine follows the instructions of the script and performs the necessary data calculations and processing to output high-quality images.

[0878] Input: The generated video script.

[0879] Output: The generated video file.

[0880] Step 6: Storing the footage and issuing an identifier

[0881] The server stores the completed video files in a database, a process that may also involve compression and format conversion of the video files.

[0882] The server issues a unique identifier for the stored video file and returns this identifier to the user, allowing them to access the video at a later time.

[0883] Input: The generated video file.

[0884] Output: The file stored in the database and the issued identifier.

[0885] (Application example 2)

[0886] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0887] Modern virtual stores face the challenge of generating and displaying promotional videos in real time that reflect users' interests and emotions. Furthermore, there is no established method for reflecting user emotions in the generation of such videos, making it difficult to customize the user experience with conventional content generation systems. Furthermore, while video file management and high-speed display are required, achieving these goals requires advanced technology.

[0888] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0889] In this invention, the server includes means for generating a video script based on the user's linguistic instructions and recognized emotions, a video generation engine for creating a video file based on the generated video script, and a database for saving the generated video file and issuing an identifier for the saving destination, thereby enabling the real-time generation and display of promotional videos according to the user's interests and emotions.

[0890] A "generative artificial intelligence model" is an artificial intelligence model for generating a video script based on a user's instructions and emotions.

[0891] The "video generation engine" is an engine that creates a video file based on the generated video script.

[0892] The "database" is a system that stores the generated video files and issues identifiers for the storage locations.

[0893] The "emotion engine" is a system that analyzes the user's facial expressions and tone of voice to obtain emotional data.

[0894] The "video display means" is a means for generating promotional video relating to a designated product in real time and showing it on the spot.

[0895] A "user interface" is an interface through which a user inputs instructions and preferences into a system.

[0896] "Scene rendering" is the computational process for generating multiple video scenes.

[0897] "Integrated music" means selecting music that is appropriate for the video scene and playing it in an integrated manner as a whole.

[0898] This invention is a system that combines a generative AI model and an emotion engine to generate and display videos in real time based on user instructions and emotions. This system enables users to automatically generate promotional videos and display them on the spot while browsing products in a virtual store.

[0899] Hardware and Software Configuration

[0900] The system includes the following main hardware and software:

[0901] Hardware:

[0902] 1. Camera: Used to capture the user's facial expressions.

[0903] 2. Microphone: Used to capture the tone of the user's voice.

[0904] 3. Smartphone or head-mounted display (HMD): Used to run applications and play videos.

[0905] software:

[0906] 1. OpenCV: Analyzes camera footage and extracts facial expression data.

[0907] 2. TensorFlow: Used to build and run emotion engines and generative AI models.

[0908] 3. JSON: Used for data exchange and storage.

[0909] 4. FFmpeg: A library for creating footage based on the generated script.

[0910] Overall system processing flow

[0911] 1. User Interface: The user selects products in the virtual store and inputs the promotion theme and preferences. For example, assume that the user prefers "up-tempo music" with a theme of "adventure with space travel images."

[0912] 2. Emotion engine: The camera and microphone are used to capture the user's facial expressions and tone of voice, which are then analyzed to obtain emotional data. For example, if the user is excited, this data is input into the system.

[0913] 3. Generative AI model: Generate a video script based on the user's input of themes, preferences, and emotional data. Generate the script using a generative AI model (e.g., GPT-3).

[0914] 4. Video Generation Engine: Generates high-quality video based on the generated script. This process includes scene rendering and music integration.

[0915] 5. Database: Stores the generated video files and issues identifiers that users can use to access the videos.

[0916] 6. Video display means: The generated video is played in real time on a smartphone or HMD. For example, a promotional video of an "adventure inspired by space travel" is played in a virtual store, which is expected to arouse users' interest in the product.

[0917] Specific examples

[0918] For example, suppose a user requests an "adventure"-themed video with an excited expression while browsing products in a virtual store. In this case, the emotion engine analyzes the user's excited facial expression and tone of voice, and provides the prompt "an adventure with the image of space travel" as input data to the generative AI model (GPT-3). The video generation engine generates a video based on the script, saves it in a database, and issues an identifier. Finally, the video is played in real time on a smartphone or HMD.

[0919] Prompt Sentence Examples

[0920] "I'd like some exciting, up-tempo music that evokes the image of space travel and adventure."

[0921] This system makes it possible to provide promotional videos in real time based on the user's emotions and individual preferences.

[0922] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0923] Step 1:

[0924] Users access the virtual store via their smartphone or HMD and select products. At this time, the user inputs the theme and preferences of the promotional video. For example, if a user inputs the instructions "adventure with a space travel image" and "up-tempo music," this data is sent to the system.

[0925] Step 2:

[0926] The device uses a camera and microphone to capture the user's facial expressions and tone of voice in real time, which then generates emotion data. For example, if the user is excited, their facial expressions and tone of voice are analyzed and the emotion data for "excitement" is generated.

[0927] Step 3:

[0928] The server receives user-entered themes, preferences, and emotional data. This data is organized in JSON format and prepared as input for the generative AI model. Examples of input data include "adventure with space travel images," "up-tempo music," and "excitement."

[0929] Step 4:

[0930] The server generates a video script using a generative AI model. Specifically, a generative AI model using TensorFlow (e.g., GPT-3) analyzes the input data and outputs a video script based on the theme of "an adventure with the image of space travel" and the emotions of "up-tempo music" and "excitement."

[0931] Step 5:

[0932] The server's video generation engine creates video files based on the generated script, using software such as FFmpeg to render the scene and integrate music. A high-quality video file is generated as the output.

[0933] Step 6:

[0934] The server stores the generated video file in a database and issues a file identifier, which is required for users to access the video later and ensures that the video file is properly managed in the database.

[0935] Step 7:

[0936] The device receives the identifier issued by the server and plays the video generated in real time. While the user browses products in the virtual store, a promotional video for an "adventure inspired by space travel" is played, providing an experience based on the user's interests and emotions.

[0937] The above steps will realize a system that generates and displays promotional videos in real time that reflect the user's preferences and emotions.

[0938] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0939] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0940] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0941] [Fourth embodiment]

[0942] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0943] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0944] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0945] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0946] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0947] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0948] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0949] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0950] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0951] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0952] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0953] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0954] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0955] An embodiment of the present invention is described below: The system uses a generative AI model to generate a video script based on a user's linguistic instructions, creates and saves a video file based on the generated video script, and provides the user with an identifier for the saved file.

[0956] 1. Creating and Sending a Request

[0957] The user inputs their ideas into the interface. For example, they specify an idea for an "adventure with a space travel image" or a preference for "upbeat, up-tempo music." This information is converted into a request object by the terminal and sent to the server.

[0958] 2. Receiving and parsing the request

[0959] The server receives requests from users and extracts information such as their ideas and preferences, which are then prepared as input data for a generative AI model.

[0960] 3. Generate the script

[0961] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences, detailing the video's scenes, narration, music, and other details.

[0962] 4. Image Generation

[0963] The server's video generation engine creates the video based on the generated script, including rendering the scene and integrating music. The generation engine creates high-quality video files according to the instructions in the script.

[0964] 5. Storage of footage and issuance of identifiers

[0965] The server saves the created video file in a database. Once saving is complete, the server generates an identifier for the storage location and returns it to the user. This identifier allows the user to access and view the created video.

[0966] Specific examples

[0967] For example, consider a user who wants to generate a video with the theme of "an adventure with the image of space travel." The user inputs the theme and preferences into the interface and submits a request. The server receives this request and uses a generative AI model to generate a script based on the theme of "an adventure with the image of space travel." The server's video generation engine then creates a video based on the script and stores it in a database. Finally, the user can watch their video using the identifier returned by the server.

[0968] The processing flow will be explained below.

[0969] Step 1:

[0970] The user inputs their ideas and preferences into the interface. For example, they specify the idea of ​​"an adventure with the image of space travel" and the preference of "upbeat, up-tempo music."

[0971] Step 2:

[0972] The device generates a request object (user_request) based on the user's input, which contains information about the user's ideas and preferences.

[0973] Step 3:

[0974] The terminal sends the generated request object to the server, which contains data based on the user's input.

[0975] Step 4:

[0976] The server receives the user's request and analyzes it, which includes extracting information about the user's ideas and preferences.

[0977] Step 5:

[0978] Based on the extracted information, the server prepares input data (ai_input) for the generative AI model, which contains detailed information about the user's ideas and preferences.

[0979] Step 6:

[0980] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences. The AI ​​model analyzes the input data and outputs the script.

[0981] Step 7:

[0982] The server's video generation engine creates the video based on the generated script, a process that includes rendering the scene and integrating music.

[0983] Step 8:

[0984] The server saves the completed video file in a database. During this saving process, an identifier is issued to ensure that the video file is properly managed.

[0985] Step 9:

[0986] The server returns the identifier of the video stored in the database to the user, allowing the user to access the generated video.

[0987] Step 10:

[0988] The user uses the returned identifier to view the video on their device, which allows them to stream or download their video content.

[0989] Example 1

[0990] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0991] Conventional video generation systems require users to have advanced skills and specialized knowledge to create specific video scripts, making them difficult for average users to use. Furthermore, the process of manually creating a script and then generating a video file is time-consuming, labor-intensive, and inefficient.

[0992] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0993] In this invention, the server includes: means for a user to input ideas and preferences as linguistic instructions into an interface; means for converting the user input into a request object and sending the request to the server; means for the server to receive the request and extract information on the ideas and preferences; means for generating a video script based on the user's ideas and preferences using a generative AI model; means including a video generation engine that creates a video file based on the generated video script; and means for saving the generated video file in a database and issuing an identifier for the storage destination. This allows users without advanced skills or specialized knowledge to easily generate video files based on linguistic instructions and efficiently use the videos.

[0994] A "user" is a person who uses the system to input verbal instructions into the interface to generate a video file.

[0995] An "interface" is an input device or software that allows a user to input their ideas and preferences.

[0996] A "request object" is a collection of information that is converted from user input into a data format and sent to the server.

[0997] "Server" means a computer system that receives and analyzes a request object from a user, generates a video script using a generative AI model, and then generates and stores the video file.

[0998] A "generative AI model" is an artificial intelligence model for generating video scripts based on a user's ideas and preferences.

[0999] A "video script" is a document that describes detailed instructions for scenes, narration, music, etc. for generating a video file.

[1000] A "video generation engine" is a device or software that creates a video file based on a video script.

[1001] A "database" is a system that stores generated video files and manages information for later access.

[1002] An "identifier" is a character string or number that uniquely identifies the generated video file and returns information to the user.

[1003] The embodiment of this invention is a system that generates a video script based on the user's verbal instructions, and creates and saves a video file based on the script. This system consists of three main components: a user, a terminal, and a server.

[1004] First, the user inputs their ideas and preferences into an interface, which can be implemented as a computer or smartphone application. For example, specific instructions such as "an adventure with a space travel image" or "upbeat, upbeat music" can be entered in text format.

[1005] The device then converts the user's input into a request object and sends it to the server, which includes the user's specified themes and preferences, allowing the server to convey the user's request in the appropriate data format.

[1006] The server receives the request sent from the device and analyzes the ideas and preferences contained therein. The analyzed information is prepared as input data for a generative AI model, such as OpenAI GPT-4.

[1007] The server invokes a generative AI model to generate a video script based on the user's ideas and preferences. The script contains detailed instructions for the video's scenes, narration, music, etc. After the script is generated, the server's video generation engine creates a video file based on the script.

[1008] The video generation engine uses software such as Adobe After Effects to render the scenes and integrate the music. The generated video files are of high quality and are based on the detailed instructions in the script.

[1009] Finally, the server stores the generated video file in a database. The database uses a high-speed storage system to efficiently manage the video files. Once the storage is complete, the server generates a unique identifier for the video file and returns this identifier to the user. The user can use this identifier to access and watch the generated video file.

[1010] As a concrete example, consider a case where a user wants to generate a video with the theme of "adventure with a space travel image" and an upbeat, up-tempo music. The user inputs the theme and preferences into the interface and sends the request to the server. The server analyzes the request and generates a video script using a generative AI model. The server's video generation engine then creates a video based on the script and stores it in a database. Finally, the user can watch their video using the identifier returned by the server.

[1011] Examples of prompt sentences include the following:

[1012] Please create a video with the theme of "An adventure with the image of space travel." I would like to use upbeat music with a bright atmosphere.

[1013] In this way, by using this system, users can easily generate and manage high-quality video based on their own ideas and preferences, without needing advanced specialized knowledge.

[1014] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1015] Step 1: User Input

[1016] The user inputs their ideas and preferences for the video into the interface. For example, they can input a theme such as "an adventure with a space travel image" or preferences such as "a bright atmosphere and up-tempo music." The input data is sent to the terminal as text data.

[1017] Step 2: Create a request object

[1018] The terminal converts textual input data from the user into a request object, which contains themes and preferences as key-value pairs. For example, a request object might contain the following data:

[1019] json

[1020] {

[1021] "theme": "Space travel-inspired adventure",

[1022] "mood": "upbeat, upbeat music"

[1023] }

[1024] The created request object is sent to the server.

[1025] Step 3: Receiving and Parsing the Request

[1026] The server receives the request object sent from the device, extracts idea and preference information from the request object, and prepares the extracted data as input data for the generative AI model.

[1027] Step 4: Input to the generative AI model

[1028] The server inputs the extracted ideas and preferences into a generative AI model, which generates a prompt like this:

[1029] Create a video with the theme of "An adventure with the image of space travel." Use upbeat, up-tempo music.

[1030] Input data is transformed into a generative AI model.

[1031] Step 5: Generate a video script

[1032] The server invokes the generative AI model to generate a video script based on the input data. The generated script contains details such as video scenes, narration, and music. For example, the following script might be generated:

[1033] json

[1034] {

[1035] "scenes": [

[1036] {"description": "A spaceship traveling through the galaxy", "duration": "30 seconds"},

[1037] {"description": "Spaceship landing scene", "duration": "20 seconds"}

[1038] ],

[1039] "narration": "The space travel adventure begins...",

[1040] "music": "Uptempo background music"

[1041] }

[1042] Step 6: Rendering the footage

[1043] The server's video generation engine creates the video based on the generated script. First, the video generation engine renders the scene. This process includes generating CG and editing existing footage. Next, music and narration are integrated, and the final video file is generated.

[1044] Step 7: Save the video file

[1045] The server stores the generated video files in a database, a process that uses a high-speed storage system.

[1046] Step 8: Generate and return identifiers

[1047] The server generates a unique identifier for each video file that has completed the storage process. This identifier will have the following format:

[1048] json

[1049] {

[1050] "identifier": "abcd1234"

[1051] }

[1052] The identifier is returned to the user, who can then use it to access and view the generated video file.

[1053] The above is the flow of processing of the program of this system.

[1054] (Application example 1)

[1055] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1056] Conventional video generation systems lack the ability to quickly generate and instantly distribute high-quality videos based on user ideas. In particular, there has been no system that allows users to easily input ideas using a smartphone and then view and share the videos generated on the spot. This has made it difficult to meet users' demands for creative content generation and instant distribution.

[1057] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1058] In this invention, the server includes: means for generating a video script based on a user's linguistic instructions using a generative artificial intelligence model; a video generation engine; a database; means for a user to input ideas and preferences using a smartphone; means for formatting the input data as input data for the generative AI model; and means for providing the user with a preview and sharing of the generated video file, thereby enabling the user to generate high-quality videos in real time using a smartphone and instantly share the videos on various platforms.

[1059] A "generative artificial intelligence model" is an algorithm that analyzes a user's linguistic instructions and automatically generates a video script based on that information.

[1060] A "video script" is a document that details the video's scenes, narration, music, etc.

[1061] A "video generation engine" is software or hardware that generates the actual video file based on a video script.

[1062] The "database" is an information management system that stores the generated video files and allows users to access them.

[1063] An "identifier" is an ID or code that uniquely identifies a specific video file within a database.

[1064] A "smartphone" is a multifunctional mobile device that combines mobile communication and computer functions.

[1065] "Formatting" is the process of converting input data into a specific format that is easy for a generative AI model to understand.

[1066] "Preview" is a function that displays a part or all of a generated video file in advance before the user finally views it.

[1067] "Sharing" is the process of sharing the generated video file with other users and platforms.

[1068] The system embodying this invention allows users to easily input their ideas and preferences using a smartphone, and then generate, view, and share high-quality video on the spot.

[1069] First, the user uses a dedicated smartphone application to input their ideas and preferences for the video they want to create, including a specific prompt such as "An adventure story exploring the mysteries of space in a light-hearted atmosphere." The application then formats the data entered by the user into JSON format and sends it to the server.

[1070] The server generates a video script using a generative AI model based on the received request. The generative AI model analyzes the user's linguistic instructions and creates a video script that details the development of scenes, narration content, music selection, etc. This script is then converted into an actual video file by the video generation engine on the server.

[1071] The video generation engine renders multiple scenes based on the generated video script and integrates appropriate music. The server stores the rendered and music-integrated video files in a database and issues an identifier for the storage location. This identifier is used by users to access and watch the generated video files.

[1072] Finally, users receive the issued identifier through a smartphone application and can preview the video generated based on their input idea. This video can be instantly shared within the application or on platforms such as social media.

[1073] The primary hardware used includes smartphones (e.g., iPhone, Android smartphone) and servers (e.g., AWS EC2 instances).The software used includes smartphone applications (e.g., developed with React Native), server-side generative AI models (e.g., GPT-4), a Flask-based backend, and a video generation engine (e.g., FFmpeg).

[1074] In this way, users can easily generate high-quality videos and share them on a variety of platforms. An example of a specific prompt is "An adventure story exploring the mysteries of space in a cheerful atmosphere." Based on this prompt, the generative AI model generates a video script that meets the user's needs, and high-quality videos are automatically generated based on that script.

[1075] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1076] Step 1:

[1077] Users launch a dedicated smartphone application and input their ideas and preferences for the video they want to create into the interface, typically entering prompts in text format such as "An adventure story exploring the mysteries of space in a cheerful atmosphere." The input data is then formatted in JSON format within the application.

[1078] Input: User's ideas and preferences (prompt text)

[1079] Output: Request data in JSON format

[1080] Step 2:

[1081] The device sends formatted JSON data to the server, and the request contains the user's ideas and preferences.

[1082] Input: Request data in JSON format

[1083] Output: Request sent to server

[1084] Step 3:

[1085] The server analyzes the received request and prepares the prompt sentence as input data for the generative AI model. The server then invokes the generative AI model to generate a video script based on the user's linguistic instructions.

[1086] Input: Request data in JSON format

[1087] Output: Generated video script

[1088] Step 4:

[1089] The server's video generation engine creates video files based on the generated video script, rendering the scenes and integrating music.

[1090] Input: Generated video script

[1091] Output: Video file

[1092] Step 5:

[1093] The server stores the generated video file in a database and issues an identifier for the storage location, which is used by users to access the generated video.

[1094] Input: Video file

[1095] Output: Destination identifier

[1096] Step 6:

[1097] The server returns the storage location identifier to the smartphone application, allowing the user to check the video generated based on their own idea.

[1098] Input: Destination identifier

[1099] Output: Send identifier to smartphone

[1100] Step 7:

[1101] Users can use a smartphone application to preview and watch the generated video, and can also share the video within the application or on platforms such as social media.

[1102] Input: Destination identifier

[1103] Output: Preview and share your footage

[1104] This will enable high-quality video to be quickly generated, viewed, and shared based on user-input ideas and preferences.

[1105] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1106] An embodiment of the present invention is described below: The system uses a generative AI model and an emotion engine to generate a video script based on a user's linguistic instructions and recognized emotions, and creates and saves a video file based on the video script.

[1107] 1. Creating and Sending a Request

[1108] The user inputs their ideas and preferences into the interface. For example, they specify the idea of ​​"an adventure with a space travel image" and their preference of "upbeat, upbeat music." At this point, the emotion engine is activated to recognize the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to obtain emotional data. This information is converted into a request object by the terminal and sent to the server.

[1109] 2. Receiving and parsing the request

[1110] The server receives user requests and emotional data, extracts information on ideas, preferences, and recognized emotions, and prepares this information as input to a generative AI model.

[1111] 3. Generate the script

[1112] The server invokes a generative AI model to generate a video script based on the user's ideas, preferences, and emotions. The generative AI model analyzes these inputs and outputs a script that reflects the video's scenes, narration, musical details, and emotional tone.

[1113] 4. Image Generation

[1114] The server's video generation engine creates the video based on the generated script. This process includes rendering the scene and integrating music. The video generation engine creates high-quality video files according to the instructions in the script.

[1115] 5. Storage of footage and issuance of identifiers

[1116] The server stores the completed video file in a database. During this storage process, an identifier is issued to ensure the video file is properly managed.

[1117] Specific examples

[1118] For example, consider a user who wants to create a video with the theme of "an adventure with the image of space travel." The user inputs the theme and preferences into the interface, while the emotion engine recognizes the user's excited facial expression and tone of voice. The server receives this request and emotion data and generates a script using a generative AI model. The script reflects the theme of "an adventure with the image of space travel" and the excited tone. The video generation engine then generates a video based on the script and stores it in a database. Finally, the user can view their video using an identifier returned by the server. This identifier allows the user to access the generated video through a web browser or application.

[1119] The processing flow will be explained below.

[1120] Step 1:

[1121] The user inputs their ideas and preferences into the user interface. For example, they might specify a theme such as "an adventure with a space travel image" and preferences such as "a cheerful atmosphere and up-tempo music." The emotion engine also analyzes the user's facial expressions and voice via a camera and microphone built into the user interface.

[1122] Step 2:

[1123] The device receives the user's input data and the emotion data from the emotion engine, and combines them to generate a request object (user_request), which contains all the themes, preferences, and emotion data.

[1124] Step 3:

[1125] The terminal sends the generated request object to the server, which includes the user's ideas, preferences, and emotion data.

[1126] Step 4:

[1127] The server receives the request object from the user and analyzes its contents, which includes extracting the user's themes, preferences, and sentiment data.

[1128] Step 5:

[1129] Based on the extracted information, the server prepares input data (ai_input) for the generative AI model, which includes the user's themes, preferences, and emotional data.

[1130] Step 6:

[1131] The server invokes a generative AI model to generate a video script based on the user's themes, preferences, and emotions. The generative AI model analyzes these input data and outputs a video script, which includes video scenes, narration, music, and a tone that reflects the user's emotions.

[1132] Step 7:

[1133] The server's video generation engine creates the video based on the generated script. This process includes rendering the scenes described in the script and integrating music. The video generation engine creates high-quality video files according to the instructions in the script.

[1134] Step 8:

[1135] The server stores the completed video file in a database. During this storage process, an identifier is issued to ensure the video file is properly managed.

[1136] Step 9:

[1137] The server returns to the user an identifier for the video stored in the database, allowing the user to access the generated video.

[1138] Step 10:

[1139] The user uses the returned identifier to watch the video on their device, which allows them to stream or download their video content through a web browser or dedicated app.

[1140] Example 2

[1141] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1142] Conventional video generation systems generate scripts based solely on the user's verbal instructions, making it difficult to reflect the user's emotions and detailed preferences. Furthermore, when the management of generated video files and the issuance of identifiers are done manually, this increases the administrative workload and increases the likelihood of errors.

[1143] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for generating a video script based on a user's linguistic instructions and recognized emotions using a generative AI model, a video generation engine means for creating a video file based on the generated video script, a database means for saving the generated video file and issuing an identifier for the saving destination, and means for analyzing the user's input data and emotion data, generating it as a request object, and transmitting it to the server. This enables video generation that reflects the user's emotions and detailed preferences, and further enables efficient automatic management of the generated video files and the issuance of identifiers.

[1144] A "generative artificial intelligence model" is a model that uses artificial intelligence technology to generate video scripts and other content based on user input data and emotional data.

[1145] A "video script" is a set of instructions that describes the scene structure, narration, musical details, and emotional tone for creating a video.

[1146] A "video generation engine" is a software or hardware system for creating the actual video based on the generated video script.

[1147] "Database" refers to a storage device and associated management system for storing and managing generated video files.

[1148] A "request object" is a data structure that includes user input data and emotion data and is generated to be sent to a server.

[1149] The "emotion engine" is a system that analyzes the user's facial expressions and tone of voice, recognizes emotions, and generates and provides that data.

[1150] An "identifier" is a unique code or number that uniquely identifies the generated video file.

[1151] "Input data" is data, including ideas and preferences, that a user provides to the system.

[1152] An embodiment of the present invention is a system for generating a video script based on a user's linguistic instructions and recognized emotions using a generative AI model and an emotion engine, and creating and saving a high-quality video file based on the script. The system includes a server, a terminal, and various user interfaces.

[1153] The user interface (e.g., a PC or smartphone application) provides a means for the user to input their ideas and preferences. The emotion engine is used to analyze the user's facial expressions and tone of voice to obtain emotion data. This emotion data, along with the user's input data, is packaged into a request object and sent from the device to the server.

[1154] The server receives the user's request object and analyzes the user's ideas, preferences, and emotional data using a video script generation method based on a generative AI model. The generative AI model is often operated using cloud computing resources. Based on the input data, the model generates a specific video script including scene composition, narration, musical details, and emotional tone.

[1155] The generated video script is then converted into a video file by a server-based video generation engine, which consists of advanced software and hardware for scene rendering and music integration. The video generation engine performs the necessary data calculations and processing according to the instructions to create a high-quality video file.

[1156] The generated video file is stored in a database on the server. During the storage process, an identifier is generated to uniquely identify the video file. This identifier allows users to access the completed video file. The identifier is used in web browsers and applications, providing users with an easy way to view the video.

[1157] Specific examples

[1158] For example, consider a case where a user wants to generate a video with the theme of "adventure with the image of space travel." The user inputs the theme and preferences (e.g., "upbeat atmosphere, up-tempo music") into the interface, and the emotion engine recognizes the user's excited facial expression and tone of voice. Next, the user's input data and emotion data are packaged into a request object and sent to the server.

[1159] The server receives this request and generates a script using a generative AI model. The generative AI model outputs a script that reflects an exciting tone in line with the theme of "an adventure with the image of space travel." Based on this script, a video generation engine renders the scene and integrates music and narration to create a video file. Finally, the server stores the video file in a database and returns an identifier to the user. The user can use this identifier to watch the video they generated.

[1160] Prompt Sentence Examples

[1161] "The theme specified by the user is 'an adventure with the image of space travel,' and their preference is 'a cheerful atmosphere, up-tempo music.' We also recognized an excited tone as the user's emotion. Please generate a video script based on this."

[1162] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1163] Step 1: User Input and Emotion Recognition

[1164] The user uses the interface to input their ideas and preferences. For example, the user may specify a theme of "adventure with space travel imagery" and a preference for "upbeat, upbeat music."

[1165] The emotion engine analyzes the user's facial expressions and tone of voice in real time to obtain emotion data. For example, if the user has an excited expression, that data is obtained by the emotion engine.

[1166] Input: Your ideas, preferences, facial expressions and tone of voice.

[1167] Output: User input data and emotion data.

[1168] Step 2: Generate and send a request

[1169] The terminal collects the acquired user input data and emotion data and organizes them into a request object, which includes the user's ideas, preferences, and emotion data.

[1170] The terminal sends a request object to the server.

[1171] Input: User input data and emotion data.

[1172] Output: The request object that is sent to the server.

[1173] Step 3: Receiving and parsing the request

[1174] The server receives the request object.

[1175] The server extracts the user's ideas, preferences, and sentiment data from the request object, checking the data for consistency and preprocessing the data if necessary.

[1176] Input: A request object.

[1177] Output: Parsed ideas, preferences, and sentiment data.

[1178] Step 4: Generate the script

[1179] The server invokes the generative AI model and generates a script using the parsed data as input.

[1180] The generative AI model outputs a specific video script containing detailed scenes, narration, musical details, and emotional tone based on user input and emotional data.

[1181] Input: Parsed idea, preference, and sentiment data.

[1182] Output: The generated video script.

[1183] Step 5: Generate the footage

[1184] The server's video generation engine creates video files based on the generated script, and performs processes such as rendering scenes and integrating music and narration.

[1185] The image generation engine follows the instructions of the script and performs the necessary data calculations and processing to output high-quality images.

[1186] Input: The generated video script.

[1187] Output: The generated video file.

[1188] Step 6: Storing the footage and issuing an identifier

[1189] The server stores the completed video files in a database, a process that may also involve compression and format conversion of the video files.

[1190] The server issues a unique identifier for the stored video file and returns this identifier to the user, allowing them to access the video at a later time.

[1191] Input: The generated video file.

[1192] Output: The file stored in the database and the issued identifier.

[1193] (Application example 2)

[1194] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1195] Modern virtual stores face the challenge of generating and displaying promotional videos in real time that reflect users' interests and emotions. Furthermore, there is no established method for reflecting user emotions in the generation of such videos, making it difficult to customize the user experience with conventional content generation systems. Furthermore, while video file management and high-speed display are required, achieving these goals requires advanced technology.

[1196] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1197] In this invention, the server includes means for generating a video script based on the user's linguistic instructions and recognized emotions, a video generation engine for creating a video file based on the generated video script, and a database for saving the generated video file and issuing an identifier for the saving destination, thereby enabling the real-time generation and display of promotional videos according to the user's interests and emotions.

[1198] A "generative artificial intelligence model" is an artificial intelligence model for generating a video script based on a user's instructions and emotions.

[1199] The "video generation engine" is an engine that creates a video file based on the generated video script.

[1200] The "database" is a system that stores the generated video files and issues identifiers for the storage locations.

[1201] The "emotion engine" is a system that analyzes the user's facial expressions and tone of voice to obtain emotional data.

[1202] The "video display means" is a means for generating promotional video relating to a designated product in real time and showing it on the spot.

[1203] A "user interface" is an interface through which a user inputs instructions and preferences into a system.

[1204] "Scene rendering" is the computational process for generating multiple video scenes.

[1205] "Integrated music" means selecting music that is appropriate for the video scene and playing it in an integrated manner as a whole.

[1206] This invention is a system that combines a generative AI model and an emotion engine to generate and display videos in real time based on user instructions and emotions. This system enables users to automatically generate promotional videos and display them on the spot while browsing products in a virtual store.

[1207] Hardware and Software Configuration

[1208] The system includes the following main hardware and software:

[1209] Hardware:

[1210] 1. Camera: Used to capture the user's facial expressions.

[1211] 2. Microphone: Used to capture the tone of the user's voice.

[1212] 3. Smartphone or head-mounted display (HMD): Used to run applications and play videos.

[1213] software:

[1214] 1. OpenCV: Analyzes camera footage and extracts facial expression data.

[1215] 2. TensorFlow: Used to build and run emotion engines and generative AI models.

[1216] 3. JSON: Used for data exchange and storage.

[1217] 4. FFmpeg: A library for creating footage based on the generated script.

[1218] Overall system processing flow

[1219] 1. User Interface: The user selects products in the virtual store and inputs the promotion theme and preferences. For example, assume that the user prefers "up-tempo music" with a theme of "adventure with space travel images."

[1220] 2. Emotion engine: The camera and microphone are used to capture the user's facial expressions and tone of voice, which are then analyzed to obtain emotional data. For example, if the user is excited, this data is input into the system.

[1221] 3. Generative AI model: Generate a video script based on the user's input of themes, preferences, and emotional data. Generate the script using a generative AI model (e.g., GPT-3).

[1222] 4. Video Generation Engine: Generates high-quality video based on the generated script. This process includes scene rendering and music integration.

[1223] 5. Database: Stores the generated video files and issues identifiers that users can use to access the videos.

[1224] 6. Video display means: The generated video is played in real time on a smartphone or HMD. For example, a promotional video of an "adventure inspired by space travel" is played in a virtual store, which is expected to arouse users' interest in the product.

[1225] Specific examples

[1226] For example, suppose a user requests an "adventure"-themed video with an excited expression while browsing products in a virtual store. In this case, the emotion engine analyzes the user's excited facial expression and tone of voice, and provides the prompt "an adventure with the image of space travel" as input data to the generative AI model (GPT-3). The video generation engine generates a video based on the script, saves it in a database, and issues an identifier. Finally, the video is played in real time on a smartphone or HMD.

[1227] Prompt Sentence Examples

[1228] "I'd like some exciting, up-tempo music that evokes the image of space travel and adventure."

[1229] This system makes it possible to provide promotional videos in real time based on the user's emotions and individual preferences.

[1230] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1231] Step 1:

[1232] Users access the virtual store via their smartphone or HMD and select products. At this time, the user inputs the theme and preferences of the promotional video. For example, if a user inputs the instructions "adventure with a space travel image" and "up-tempo music," this data is sent to the system.

[1233] Step 2:

[1234] The device uses a camera and microphone to capture the user's facial expressions and tone of voice in real time, which then generates emotion data. For example, if the user is excited, their facial expressions and tone of voice are analyzed and the emotion data for "excitement" is generated.

[1235] Step 3:

[1236] The server receives user-entered themes, preferences, and emotional data. This data is organized in JSON format and prepared as input for the generative AI model. Examples of input data include "adventure with space travel images," "up-tempo music," and "excitement."

[1237] Step 4:

[1238] The server generates a video script using a generative AI model. Specifically, a generative AI model using TensorFlow (e.g., GPT-3) analyzes the input data and outputs a video script based on the theme of "an adventure with the image of space travel" and the emotions of "up-tempo music" and "excitement."

[1239] Step 5:

[1240] The server's video generation engine creates video files based on the generated script, using software such as FFmpeg to render the scene and integrate music. A high-quality video file is generated as the output.

[1241] Step 6:

[1242] The server stores the generated video file in a database and issues a file identifier, which is required for users to access the video later and ensures that the video file is properly managed in the database.

[1243] Step 7:

[1244] The device receives the identifier issued by the server and plays the video generated in real time. While the user browses products in the virtual store, a promotional video for an "adventure inspired by space travel" is played, providing an experience based on the user's interests and emotions.

[1245] The above steps will realize a system that generates and displays promotional videos in real time that reflect the user's preferences and emotions.

[1246] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1247] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1248] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1249] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1250] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1251] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1252] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1253] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1254] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1255] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1256] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1257] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1258] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1259] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1260] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1261] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1262] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1263] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1264] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1265] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1266] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1267] The following is further disclosed regarding the above embodiment.

[1268] (Claim 1)

[1269] means for generating a video script based on a user's linguistic instructions using a generative artificial intelligence model;

[1270] a video generation engine means for creating a video file based on the generated video script;

[1271] a database means for storing the generated video file and issuing an identifier for the storage destination;

[1272] A system including:

[1273] (Claim 2)

[1274] 10. The system of claim 1, wherein the system renders multiple scenes and integrates music in generating a video file.

[1275] (Claim 3)

[1276] 10. The system of claim 1, further comprising means for generating input data to the generative artificial intelligence model based on user-provided ideas and preferences.

[1277] "Example 1"

[1278] (Claim 1)

[1279] a means for a user to input ideas and preferences into the interface as verbal instructions;

[1280] a means for converting user input into a request object and sending the request to the server;

[1281] a means for the server to receive the request and extract the idea and preference information;

[1282] means for generating a video script based on a user's ideas and preferences using a generative AI model;

[1283] means including a video generation engine for creating a video file based on the generated video script;

[1284] A means for storing the generated video file in a database and issuing an identifier for the storage location;

[1285] A system including:

[1286] (Claim 2)

[1287] 10. The system of claim 1, wherein the video generation engine renders scenes based on a script and integrates music and narration.

[1288] (Claim 3)

[1289] 10. The system of claim 1, wherein the server further comprises means for preparing input data to the generative AI model.

[1290] "Application Example 1"

[1291] (Claim 1)

[1292] means for generating a video script based on a user's linguistic instructions using a generative artificial intelligence model;

[1293] a video generation engine means for creating a video file based on the generated video script;

[1294] a database means for storing the generated video file and issuing an identifier for the storage destination;

[1295] a means for a user to input ideas and preferences using a smartphone;

[1296] A means for formatting the input data as input data to a generative AI model;

[1297] a means for providing a user with a preview and sharing of the generated video file;

[1298] A system including:

[1299] (Claim 2)

[1300] 10. The system of claim 1, wherein the system renders multiple scenes and integrates music in generating a video file.

[1301] (Claim 3)

[1302] 10. The system of claim 1, further comprising means for generating input data to the generative artificial intelligence model based on user-provided ideas and preferences.

[1303] "Example 2: Combining Emotion Engines"

[1304] (Claim 1)

[1305] means for generating a video script based on the user's linguistic instructions and the recognized emotions using a generative artificial intelligence model;

[1306] a video generation engine means for creating a video file based on the generated video script;

[1307] a database means for storing the generated video file and issuing an identifier for the storage destination;

[1308] means for analyzing the user's input data and emotion data, generating a request object, and transmitting the request object to a server;

[1309] A system including:

[1310] (Claim 2)

[1311] 10. The system of claim 1, wherein the system renders multiple scenes and integrates music in generating a video file.

[1312] (Claim 3)

[1313] 10. The system of claim 1, further comprising means for generating input data to the generative artificial intelligence model based on user-provided ideas and preferences as well as perceived affective data of the user.

[1314] "Application example 2 when combining emotion engines"

[1315] (Claim 1)

[1316] means for generating a video script based on the user's linguistic instructions and the recognized emotions using a generative artificial intelligence model;

[1317] a video generation engine means for creating a video file based on the generated video script;

[1318] a database means for storing the generated video file and issuing an identifier for the storage destination;

[1319] emotion engine means for acquiring emotion data by analyzing a user's facial expression and tone of voice;

[1320] a video display means for generating a promotional video related to a designated product in real time and showing it on the spot;

[1321] A system including:

[1322] (Claim 2)

[1323] 10. The system of claim 1, wherein the system renders multiple scenes and integrates music in generating a video file.

[1324] (Claim 3)

[1325] 10. The system of claim 1, further comprising means for generating input data to the generative artificial intelligence model based on user-provided ideas and preferences. [Explanation of symbols]

[1326] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for generating a video script based on a user's linguistic instructions using a generative artificial intelligence model; a video generation engine means for creating a video file based on the generated video script; a database means for storing the generated video file and issuing an identifier for the storage destination; A system including:

2. The system of claim 1 , which renders multiple scenes and integrates music in generating a video file.

3. 10. The system of claim 1, further comprising means for generating input data to the generative artificial intelligence model based on user-provided ideas and preferences.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A