System

The system automates video production using generative AI, enabling non-experts to efficiently create high-quality videos by submitting requests, providing feedback, and finalizing videos with user input.

JP2026021187APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122869
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional video production requires advanced technology and specialized knowledge, making it difficult for ordinary users to produce professional-quality videos efficiently.

Method used

A system that automates video production using generative artificial intelligence, allowing users to submit requests, receive data, generate scripts and structures, provide feedback on preview videos, and finalize high-quality videos with user input.

Benefits of technology

Enables non-experts to quickly create high-quality videos by automating the process and incorporating real-time feedback, improving work efficiency and video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021187000001_ABST
    Figure 2026021187000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for a user to send a request for video production; means for a server to receive the user's request and provide related data to generative artificial intelligence; means for the generative artificial intelligence to generate a script and a proposed configuration of a video based on the provided data; means for the server to generate a simplified preview video based on the generated proposed configuration and provide the preview video to the user; means for the user to provide feedback on the preview video; means for the generative artificial intelligence to generate a final version of the video based on the feedback; and means for the server to provide the final version of the video to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Traditional video production requires advanced technology and a long time, making it difficult for ordinary users without specialized knowledge to produce professional-quality videos. Furthermore, the manual editing and material selection process requires a lot of time and effort, reducing work efficiency. There is a need for a method that solves these problems and allows anyone to quickly and easily produce high-quality videos. [Means for solving the problem]

[0005] The present invention provides a system including: a means for a user to send a video production request; a means for a server to receive the user's request and provide relevant data to a generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a proposed structure based on the provided data; a means for the server to generate a simple preview video based on the generated structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for the generative artificial intelligence to generate a final version of the video based on the feedback; and a means for the server to provide the final version of the video to the user. This system automates the video production process, allowing even non-experts to quickly produce high-quality videos. Furthermore, the generative artificial intelligence includes means for acquiring and incorporating additional information from external data sources and means for automatically integrating text, images, video clips, and audio narration, thereby enabling even higher-quality and more efficient video production.

[0006] "User" refers to the user who submits a video production request and receives the generated preview video and final video.

[0007] "Server" refers to the computer system that receives requests from users, provides data to the generative artificial intelligence, and provides previews and final versions of the generated video to users.

[0008] "Generative AI" refers to artificial intelligence technology that generates video scripts, plot plans, and final videos based on provided data.

[0009] A "request" refers to the act of a user entering basic information about the video they wish to create (title, purpose, summary of content, etc.) and sending it to the server.

[0010] "Related data" refers to the information and materials (images, video clips, text, etc.) that generative AI needs to create a video.

[0011] A "script" refers to text that explains the content of a video and a specific structural document that directs the flow of scenes.

[0012] "Structure plan" refers to the actual video production plan, including the overall scene arrangement, visual effects, text placement, etc. of the video based on the script.

[0013] A "preview video" is an early stage video generated by generative artificial intelligence, a simplified version of the video that users can review and provide feedback on.

[0014] "Feedback" refers to the act of a user checking a preview video and providing correction requests or improvement suggestions to the server.

[0015] "Final video" refers to the complete video that is finally created by the generative artificial intelligence, incorporating user feedback. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention relates to a system that enables anyone to easily and quickly create high-quality videos. This system involves a process in which a user submits a video production request, a server receives the request and provides data to a generative artificial intelligence (AI), the AI ​​generates a script and a draft structure for the video, the server provides a preview version to the user, and finally generates and provides the final version of the video that incorporates user feedback.

[0038] Specifically, this includes the following processes:

[0039] Users submit requests for video production using a dedicated application or web portal. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[0040] The server analyzes the requests received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[0041] Generative AI generates a video script and structure proposal based on the information provided. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0042] The server generates a simple preview video based on the proposed composition by the generative AI and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[0043] Users can view the preview video and provide feedback and corrections as needed. This feedback is entered via the user's dashboard. Users can provide correction instructions for specific scenes or text.

[0044] The server receives user feedback and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user feedback.

[0045] Finally, the server provides the final video to the user, which is stored on the server and available for download or streaming from the user's dashboard.

[0046] Specific examples

[0047] Example 1: Corporate PR video production

[0048] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[0049] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0050] Based on this information, generative AI creates a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending."

[0051] The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. The user reviews the preview video and provides feedback such as "Make the introduction a little shorter" or "Explain the benefits in more detail."

[0052] The server receives user feedback, and the generative AI generates the final video that reflects the feedback.

[0053] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[0054] In this way, this invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing generative AI, the video production process can be automated, significantly improving work efficiency.

[0055] The processing flow will be explained below.

[0056] Step 1:

[0057] Users access a dedicated application or web portal to submit a video production request. Specifically, users enter the required information, such as the title, purpose, summary of the content, and the materials they want to use, and then click the "Submit Request" button.

[0058] Step 2:

[0059] The server analyzes the request received from the user and collects relevant data based on the request information (e.g., materials from internal databases, additional information from external data sources, etc.).

[0060] Step 3:

[0061] The server provides the collected relevant data to the generative artificial intelligence (AI), including images, text, and existing video clips.

[0062] Step 4:

[0063] Generative AI generates a video script and structure proposal based on the provided data. Specifically, it uses natural language processing technology to generate text descriptions for each scene, image processing technology to determine the placement and effects of images, and obtains additional information from external data sources as needed.

[0064] Step 5:

[0065] The server generates a simple preview video based on the generated script and plot plan, which includes scene stitching, text overlays, and some visual effects.

[0066] Step 6:

[0067] The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready, after which the user can access the dashboard to view the preview video.

[0068] Step 7:

[0069] Users can check the preview video and provide feedback and correction requests through the dashboard. Specifically, users can enter comments for each scene and text, and specify additional corrections and improvements.

[0070] Step 8:

[0071] The server receives feedback from users and reflects it in the generative AI. The server analyzes the feedback and notifies the generative AI of any necessary changes.

[0072] Step 9:

[0073] The generative AI incorporates the feedback it receives and generates the final video. The AI ​​then learns from the data again and makes adjustments based on the user's requests.

[0074] Step 10:

[0075] The server uploads the final video to the user's dashboard and notifies the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[0076] In this way, collaboration between users, servers, and generative AI streamlines the video production process, enabling anyone to quickly and easily create high-quality videos.

[0077] Example 1

[0078] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0079] Traditionally, video production required specialized knowledge and skills, making it difficult for non-expert users to create high-quality videos in a short amount of time. Furthermore, the feedback process was inefficient, making it difficult to incorporate user feedback in real time. This resulted in increased costs and time for video production, and reduced operational efficiency.

[0080] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0081] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a draft structure based on the provided data; a means for the server to generate a simple preview video based on the generated draft structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for the generative artificial intelligence to generate a final version of the video based on the feedback; a means for the server to provide the final version of the video to the user; a means for the user to input and send feedback in real time; and a means for the server to receive the user's real-time feedback and reflect it in the generative artificial intelligence. This enables even non-expert users to produce high-quality videos in a short amount of time, and makes the video production process more efficient by reflecting feedback in real time.

[0082] "User" refers to the person who submits a video production request and provides feedback.

[0083] "Request" refers to information regarding requests and specifications submitted by a User for video production.

[0084] "Server" refers to a device or system that receives user requests, provides relevant data to the generative artificial intelligence, and manages the generated videos and feedback.

[0085] "Generative AI" refers to an AI technology that generates video scripts and plot plans based on provided data.

[0086] "Associated data" refers to information necessary for video production, such as text, images, video clips, audio narration, etc.

[0087] A "script" refers to a document or text that specifically outlines the scene structure and content of a video.

[0088] A "synopsis plan" refers to a plan that shows the overall structure of the video and details of each scene.

[0089] A "preview video" refers to a simple video created based on the generated composition plan.

[0090] "Feedback" refers to requests for corrections or opinions provided by users regarding preview videos.

[0091] "Final video" refers to a high-quality video that has been completed with user feedback reflected.

[0092] "Real-time feedback" refers to feedback sent by a user immediately on the spot.

[0093] The present invention relates to a system that allows users to easily and quickly create high-quality videos. This system involves a process in which a user submits a video production request, a server receives the request and provides data to a generative artificial intelligence (AI), the AI ​​generates a script and a draft structure for the video, the server provides a preview version to the user, and finally generates and provides the final version of the video that incorporates user feedback.

[0094] Specifically, it is configured as follows:

[0095] First, a user submits a request for video production using a dedicated application or web portal. In this request, the user enters basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[0096] The server then analyzes the request received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles, social media posts, etc.). The generative AI used by the server includes, for example, OpenAI's GPT-4 and AI with image processing technology.

[0097] Generative AI generates a video script and structure proposal based on the information provided. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0098] The server then generates a simple preview video based on the proposed composition by the generative AI and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server then uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[0099] Users can preview the video and provide feedback or suggestions for corrections as needed. This feedback is entered via the user's dashboard. Users can also provide correction suggestions for specific scenes or text and provide feedback in real time.

[0100] The server receives user feedback and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user feedback.

[0101] Finally, the server provides the final video to the user, which is stored on the server and available for download or streaming from the user's dashboard.

[0102] Specific examples

[0103] Example 1: Corporate PR video production

[0104] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[0105] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0106] Based on this information, generative AI creates a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending."

[0107] The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. The user reviews the preview video and provides feedback such as "Make the introduction a little shorter" or "Explain the benefits in more detail."

[0108] The server receives user feedback, and the generative AI generates the final video that reflects the feedback.

[0109] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[0110] Prompt Sentence Examples

[0111] "We would like to create a video to introduce a new product. The target users are businessmen in their 30s living in urban areas. The product's features are durability and beautiful design. Could you create a promotional video that emphasizes these points?"

[0112] In this way, this invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing generative AI, the video production process can be automated, significantly improving work efficiency.

[0113] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0114] Step 1:

[0115] A user submits a request for video production.

[0116] Specific operation: The user accesses a dedicated application or web portal and enters basic information such as the title, purpose, content summary, and materials to be used in the video production request form. When the user clicks the send button, the request is sent to the server.

[0117] Input: Information such as title "New product introduction", purpose "Customer promotion", and content summary "Highlight the features and benefits of the new product".

[0118] Output: The request data sent to the server.

[0119] Step 2:

[0120] The server receives and analyzes the user's request.

[0121] Specific operation: The server receives the request information from the user, analyzes its content, extracts the necessary relevant data based on the request content, and collects the necessary information from internal databases and external data sources.

[0122] Input: User request data.

[0123] Output: Relevant data to feed into the generative AI (e.g., product images, existing promotional video clips, text information).

[0124] Step 3:

[0125] Generative AI generates a video script and structure based on relevant data.

[0126] Specific operation: Based on the analysis results, the server provides prompts and related data to the generative AI. The generative AI then uses the provided information to create a video script and structure. The AI ​​uses natural language processing and image processing techniques to determine scene composition, text placement, and visual effects, and also generates text for the voice narration.

[0127] Input: relevant data, prompt statement.

[0128] Output: A video script and outline.

[0129] Step 4:

[0130] The server generates a preview video based on the generated configuration plan and provides it to the user.

[0131] Specific operation: The server generates a simple preview video based on the script and composition plan provided by the generative AI. The preview video includes scene stitching, text overlays, etc. The server uploads the generated preview video to the user's dashboard and notifies the user that it is ready.

[0132] Input: Video script and outline.

[0133] Output: The preview video that is provided to the user.

[0134] Step 5:

[0135] The user provides feedback on the preview video.

[0136] Specific operations: Users can check the preview video on the dashboard, input correction requests and feedback, provide correction instructions for specific scenes and text, and submit. Users can also submit feedback in real time.

[0137] Input: Preview video, user feedback.

[0138] Output: Feedback data sent to the server.

[0139] Step 6:

[0140] The server receives the feedback and reflects it in the generative AI.

[0141] Specific operation: The server receives feedback from the user and sends it to the generative AI. The generative AI reflects the feedback and revises the video script and composition plan. The server generates a new preview video based on the revised composition plan and provides it to the user again.

[0142] Input: User feedback, revised script and structure of the video.

[0143] Output: Revised preview video.

[0144] Step 7:

[0145] The server provides the final video to the user.

[0146] How it works: The server generates a high-quality final video based on the final script and plot plan. The generated video is stored on the server and uploaded to the user's dashboard. The user can then watch the final video by downloading or streaming.

[0147] Input: Final video script and outline.

[0148] Output: The final video that is delivered to the user.

[0149] (Application example 1)

[0150] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0151] Conventional video production systems have made it difficult for users without specialized knowledge and advanced skills to quickly produce high-quality videos. Furthermore, they lacked an interface for individual users to easily create and distribute personalized video content. Furthermore, they lacked a prompt input mechanism for generating video composition plans customized to the user's needs.

[0152] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0153] In this invention, the server includes: means for a user to send a video production request; means for providing related data to a generative artificial intelligence; means for generating a video script and a proposed composition based on the provided data; means for generating a simple preview video based on the proposed composition and providing it to the user; means for the user to provide feedback on the preview video; means for generating a final version of the video based on the feedback; means for providing the final version of the video to the user; means for the user to upload the generated video to a distribution service; and means for the user to customize the proposed composition by inputting a prompt. This makes it possible to quickly and easily create high-quality videos and easily distribute personalized video content without specialized knowledge or advanced technology.

[0154] "Means for users to submit video production requests" means the ability for users to use their own devices to input video production requests and details through a specific format or interface and send them to the server.

[0155] "Generative AI" refers to AI technology that automatically generates video scripts and plots based on provided data and information. It utilizes natural language processing and image processing technologies.

[0156] A "simple preview video" is a prototype video generated based on an initial video script and structure created by generative artificial intelligence, for users to check and provide feedback.

[0157] The "means for users to provide feedback on the preview video" refers to an interface or function that allows users to check the preview video, input requests for improvement or opinions about the content, and send the input to the server.

[0158] The "means of generating the final version of the video" is a function in which the generative artificial intelligence re-edits and re-structures the video based on feedback provided by the user, generating the final, completed version of the video.

[0159] The "means by which the server provides the final video to the user" refers to the function by which the server stores the generated final video and provides it for download or streaming through an interface accessible to the user.

[0160] "Means for uploading user-generated videos to a distribution service" refers to a function that allows users to easily upload and publish completed videos to a content distribution platform.

[0161] A "prompt sentence" is a text-based sentence that a user uses to input specific instructions or requests to a generative artificial intelligence when generating a script or plot plan for a video.

[0162] The present invention relates to a system that allows anyone to easily create high-quality videos and upload them to a content distribution service. Hereinafter, an embodiment of the invention will be described in detail.

[0163] The system mainly consists of a terminal where users send requests, a server that receives and processes video production requests, generative artificial intelligence (AI), and related databases and external data sources. The following describes a specific implementation of the process in which a user makes a video production request, AI generates a video script and composition plan based on that request, and ultimately provides a high-quality video.

[0164] 1. User submits request:

[0165] Users submit video production requests using a dedicated smartphone application or web portal, where they enter details such as the title of the video they want to create, its purpose, a summary of the content, and the materials they will use.

[0166] 2. Server reception and AI provision:

[0167] The request sent by the user is received by the server, which analyzes the information and provides the generative AI with relevant data, which is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[0168] 3. Video script and structure generation:

[0169] Generative AI generates a video script and structure proposal based on the provided data. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0170] 4. Generate preview video:

[0171] The server creates a simple preview video based on the generated composition plan and provides it to the user. This preview video includes scene stitching, text overlays, etc., so that the user can check it.

[0172] 5. User feedback processing:

[0173] Users can view the preview video and provide feedback and suggestions for revisions, which are entered via the user's dashboard. This feedback can include specific instructions such as "make the introduction a little shorter" or "explain the benefits in more detail."

[0174] 6. Generate the final video:

[0175] The server receives user feedback and applies it to the generative AI. The AI ​​then retrains and generates a revised version of the video that reflects the user's feedback. Finally, the server provides the final version of the video to the user.

[0176] 7. Uploading and Publishing Videos:

[0177] Users can upload the generated video to a content distribution service, and can customize the video's structure by entering specific prompts.

[0178] Hardware and software used:

[0179] Hardware: Smartphones, server computers

[0180] Software: Python, Flask (for server processing), generative AI module, video editing module

[0181] Examples:

[0182] An example prompt for a user to create a promotional video for a new product:

[0183] Title: New Product Review

[0184] Purpose: To inform the audience

[0185] Summary: Highlight the features and benefits of your new product and compare it with competing products. Include specific usage scenarios if possible.

[0186] Using this prompt, the generative AI generates specific and effective video content.

[0187] As described above, this system makes it easy for anyone to create and distribute high-quality videos, automating the video production process and significantly improving work efficiency.

[0188] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0189] Step 1:

[0190] Users submit video production requests via a dedicated smartphone application or web portal.

[0191] Input: Information about the video title, purpose, summary, and materials used that the user enters into the application.

[0192] Specific operation: When the user enters the required information into the form and clicks the submit button, this data is sent from the application to the server.

[0193] Output: The user's request information arrives at the server.

[0194] Step 2:

[0195] The server receives the user's request, analyzes it, and provides the relevant data to the generative artificial intelligence.

[0196] Input: Request information sent by the user.

[0197] What it does: The server analyzes the request and gathers the necessary data from its internal database. If necessary, it retrieves additional data from external sources (e.g., news articles or social media posts) and provides it to the generative AI.

[0198] Output: Relevant data is provided to the generative artificial intelligence.

[0199] Step 3:

[0200] Generative AI generates a video script and structure based on the data provided.

[0201] Input: Relevant data provided to the generative artificial intelligence.

[0202] Specific operation: AI uses natural language processing and image processing technology to determine the video's scene composition, text placement, visual effects, etc., and also generates text for the voice narration.

[0203] Output: A script and outline of the generated video.

[0204] Step 4:

[0205] The server generates a simple preview video based on the generated configuration plan and provides it to the user.

[0206] Input: A video script and plot plan generated by generative artificial intelligence.

[0207] What happens: The server uses the video editing module to create a preview video, stitching scenes together and adding text overlays, uploading the generated preview video to the user's dashboard, and sending a notification that it's ready.

[0208] Output: A simple preview video provided to the user.

[0209] Step 5:

[0210] The user provides feedback on the preview video.

[0211] Input: The preview video the user watched and their feedback (e.g., a shorter introduction, a more detailed explanation of the benefits, etc.).

[0212] Specific action: A user enters comments or correction requests into the feedback form on the dashboard and clicks the submit button.

[0213] Output: User feedback reaches the server.

[0214] Step 6:

[0215] Based on the feedback, generative artificial intelligence generates the final version of the video.

[0216] Input: Feedback information from the user.

[0217] How it works: The server provides feedback to the generative AI, which then re-learns the content and generates a revised script and structure based on the feedback. The server then uses the video editing module to generate the final video.

[0218] Output: The final video generated.

[0219] Step 7:

[0220] The server provides the final video to the user.

[0221] Input: The final video.

[0222] What it does: Saves the final video to our servers and provides a link for users to download or stream it from their dashboard. Notifies users that their video is complete.

[0223] Output: The final video that is delivered to the user.

[0224] Step 8:

[0225] User-generated videos can be uploaded to content distribution services.

[0226] Input: Final video and streaming service account information.

[0227] What it does: Users click the "Upload" button in the application, follow a few simple steps, and the final video is uploaded to the distribution platform of their choice.

[0228] Output: The video published to a distribution platform.

[0229] Step 9:

[0230] Users can customize the video composition by entering a prompt.

[0231] Input: The specific prompt text that the user enters.

[0232] How it works: A user enters a prompt into a dashboard or application, which is then provided to the generative AI, which then generates content based on the prompt and customizes it to fit the user's needs.

[0233] Output: A video based on your customized composition.

[0234] By following the above steps, users can quickly create and distribute high-quality videos without having specialized knowledge.

[0235] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0236] The present invention relates to a system that enables users to easily and quickly create high-quality videos. This system involves a process in which the user submits a video production request, a server receives the request, and provides the data to a generative artificial intelligence (AI). The AI ​​then generates a script and a draft structure for the video, and the server provides a preview version to the user. Finally, the system generates and provides a final version of the video that incorporates user feedback. Furthermore, the present invention aims to achieve high satisfaction by incorporating an emotion engine that recognizes user emotions, thereby reflecting the user's emotions in the video production process.

[0237] Specifically, this includes the following processes:

[0238] Users submit requests for video production using a dedicated application or web portal. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[0239] The server analyzes the requests received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[0240] The generative AI generates a video script and structure proposal based on the information provided. The AI ​​uses natural language processing and image processing techniques to determine the video's scene composition, text placement, visual effects, etc. It also obtains additional information from external data sources as needed. Here, the emotion engine analyzes the user's reactions and emotions and provides this information to the generative AI, which then generates content that is more tailored to the user.

[0241] The server generates a simple preview video based on the generated script and plot plan and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[0242] When users watch the preview video, the emotion engine measures their reactions in real time. The emotions (e.g., satisfaction, dissatisfaction, excitement, etc.) felt by the user while watching the preview video are recorded. Users can provide feedback via the dashboard. Specifically, users can enter comments for each scene or text and specify additional corrections or improvements. This feedback also includes the emotional data measured by the emotion engine.

[0243] The server receives user feedback and emotional data and reflects it in the generative AI. The generative AI then retrains and generates a revised version of the video that reflects the user's feedback and emotions. The AI ​​then edits the video to emphasize the positive emotions felt by the user and reduce negative emotions.

[0244] Finally, the server uploads the final video to the user's dashboard and sends a notification to the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[0245] Specific examples

[0246] Example 1: Corporate PR video production

[0247] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[0248] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0249] The generative AI uses this information to create a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending." The emotion engine analyzes the user's emotions and provides them to the generative AI, which then reflects the content according to the user's preferences.

[0250] The server creates a preview video based on the generated configuration and provides it to the user via a dashboard. The emotions (e.g., excitement, anticipation, doubt) that arise when the user watches the preview video are recorded.

[0251] Users provide feedback such as "Make the introduction shorter" or "Explain the benefits in more detail," and data from the sentiment engine is also sent.

[0252] The server receives user feedback, and the generative AI reflects the feedback and emotional data to generate the final video.

[0253] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[0254] In this way, the present invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing the emotion engine, more personalized content that reflects the user's emotions is generated, thereby improving user satisfaction.

[0255] The processing flow will be explained below.

[0256] Step 1:

[0257] Users access a dedicated application or web portal to submit a video production request. Specifically, users enter the required information, such as the title, purpose, summary of the content, and the materials they want to use, and then click the "Submit Request" button.

[0258] Step 2:

[0259] The server analyzes the request received from the user and collects relevant data based on the request information (e.g., materials from internal databases, additional information from external data sources, etc.).

[0260] Step 3:

[0261] The server provides the collected relevant data to the generative artificial intelligence (AI), including images, text, and existing video clips.

[0262] Step 4:

[0263] Generative AI generates a video script and structure proposal based on the provided data. Specifically, it uses natural language processing technology to generate text descriptions for each scene, image processing technology to determine the placement and effects of images, and obtains additional information from external data sources as needed.

[0264] Step 5:

[0265] The emotion engine measures users' reactions in real time. While users are submitting video creation requests, the emotion engine analyzes their facial expressions, voice, and other biometric information to generate emotion data.

[0266] Step 6:

[0267] The server generates a simple preview video based on the generated script and plot plan, which includes scene stitching, text overlays, and some visual effects.

[0268] Step 7:

[0269] The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready, after which the user can access the dashboard to view the preview video.

[0270] Step 8:

[0271] Users can view the preview video and provide feedback and correction requests via the dashboard. Specifically, users can enter comments for each scene and text, specifying additional corrections and improvements. This feedback also includes emotional data measured by the emotion engine.

[0272] Step 9:

[0273] The server receives feedback and emotion data from users and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user's feedback and emotion.

[0274] Step 10:

[0275] The server uploads the final video to the user's dashboard and notifies the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[0276] In this way, collaboration between users, servers, generative AI, and the emotion engine allows even non-experts to quickly and easily create high-quality videos. Furthermore, the inclusion of the emotion engine allows users' emotions to be reflected in the video production process, enabling the provision of more personalized content with a high level of satisfaction.

[0277] Example 2

[0278] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0279] In the past, users needed specialized knowledge and skills to quickly create high-quality videos. This meant that video production required a lot of time and money, making it difficult for average users to achieve this. It was also difficult to personalize the video content to match the user's emotions, making it difficult to increase user satisfaction.

[0280] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0281] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a draft structure based on the provided data; a means for the server to generate a simple preview video based on the generated draft structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for an emotion engine to analyze and acquire user emotion data; a means for the generative artificial intelligence to generate a revised version of the video based on the feedback and emotion data; a means for the server to provide the revised version of the video to the user; and a means for the server to generate a final version of the video and provide it to the user. This enables users to quickly create high-quality videos without specialized knowledge or skills, and further enables personalization of video content that reflects the user's emotions, thereby improving user satisfaction.

[0282] "User" refers to an individual or organization that uses the System to make a video production request.

[0283] "Server" refers to the computer system that receives and processes user requests, provides relevant data to the generative artificial intelligence, and generates videos and manages feedback.

[0284] "Generative AI" refers to AI technology that generates a video script and structure based on provided data, and then creates revised and final versions of the video.

[0285] An "emotion engine" refers to a technology that analyzes user reactions and emotions, acquires them as data, and provides this to generative artificial intelligence.

[0286] A "video production request" refers to the action of sending information to the server, including basic information about the video the user wants to produce (e.g., title, purpose, summary of content, materials to be used, etc.).

[0287] "Related Data" means data necessary to generate a video script and story (e.g., user-provided information, data from internal databases, data from external data sources).

[0288] A "script" refers to the story or script of a video that generative artificial intelligence creates based on the data provided.

[0289] "Composition proposal" refers to the scene composition and layout proposal for a video that is determined by generative artificial intelligence based on the data provided.

[0290] "Preview video" refers to a simple video generated by the server based on the script and plot plan created by generative AI, which is used by users to confirm and provide feedback.

[0291] "Feedback" refers to opinions and correction requests provided by users regarding preview videos.

[0292] "Revised video" refers to a video edited by generative artificial intelligence based on user feedback and emotional data.

[0293] "Final video" refers to the final completed video created by the generative artificial intelligence that the server provides to the user.

[0294] "Database" refers to a system that resides within a server and stores and manages related data.

[0295] "External data sources" refers to information sources that can be obtained from outside, such as news articles and social media posts.

[0296] This invention relates to a system that enables users to quickly create high-quality videos without specialized knowledge or skills. This system is composed of a server, generative artificial intelligence, an emotion engine, a dedicated application, and a web portal.

[0297] Hardware and software used

[0298] Hardware: Servers (computer systems with high-performance computing resources), terminals (user computers, smartphones, etc.)

[0299] Software: Dedicated applications, web portals, generative AI (e.g., OpenAI's GPT model, image generation model, etc.), emotion recognition engines (e.g., Affectiva, Microsoft Azure Emotion API, etc.)

[0300] server

[0301] The server receives video production requests submitted by users using a dedicated application or web portal. The request includes basic information such as the video title, purpose, summary, and materials used. The server analyzes this information and prepares relevant data to provide to the generative artificial intelligence. This relevant data includes information from the server's internal database and information from external data sources (such as news articles and social media posts).

[0302] Generative Artificial Intelligence

[0303] The generative AI generates a video script and structure proposal based on data provided by the server. It uses natural language processing and image processing techniques to determine the video's scene composition, text placement, visual effects, etc. It also uses data from the emotion engine to generate personalized content that reflects the user's emotions.

[0304] Emotion Engine

[0305] The emotion engine measures and analyzes the user's emotions in real time while they are watching the preview video. Emotional data such as satisfaction, dissatisfaction, and excitement felt by the user is recorded and provided to the generative AI. This allows content that reflects the user's emotions to be generated.

[0306] Dedicated application / web portal

[0307] Users can use a dedicated application or web portal to submit video production requests, view the generated preview videos and final videos, and provide specific feedback on the preview videos to the server through the application or web portal.

[0308] Example: Production of corporate PR videos

[0309] 1. A user requests the creation of a promotional video for a new product. At that time, the user enters information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and benefits of the new product."

[0310] 2. The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0311] 3. The generative AI uses this information to create a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending." The emotion engine analyzes the user's emotional data and provides it to the generative AI, which then reflects the content according to the user's preferences.

[0312] 4. The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. As the user watches the video, the emotion engine records emotions such as excitement, anticipation, and doubt.

[0313] 5. The user provides feedback such as a short introduction and detailed benefit explanation, along with sentiment engine data.

[0314] 6. The server receives the user feedback, and the generative AI generates a revised version of the video based on the feedback and emotional data.

[0315] 7. Once the final video is completed, the server will provide it to the user, who can then download the high-quality final PR video from their dashboard and use it for promotional purposes.

[0316] Example prompts for generative AI models

[0317] "Write a script for a promotional video highlighting the features of a new product. The title should be 'New Product Introduction' and the purpose should be to promote it to customers. Use the following information: product images, text information, and existing promotional video clips."

[0318] "Edit your videos to reflect your users' emotional data and build excitement and anticipation."

[0319] In this way, the present invention allows users to quickly and easily create high-quality videos. In addition, by using an emotion engine, personalized video content that reflects the user's emotions is generated, improving user satisfaction.

[0320] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0321] Step 1:

[0322] The user submits a request for video production using a dedicated application or web portal. The user enters basic information such as the video title, purpose, and content summary, and clicks the submit button. The input data includes the video title "New product introduction," purpose "Customer promotion," and content summary "Highlighting the features and benefits of the new product." This is then sent to the server.

[0323] Step 2:

[0324] The server receives the request sent by the user and analyzes the content. The analyzed data includes the title, purpose, and content summary. Based on this data, the server prepares related data to provide to the generative artificial intelligence (AI). Specifically, it collects product images, existing promotional video clips, and text information from the server's internal database. This becomes the input data provided to the generative AI.

[0325] Step 3:

[0326] The server provides the prepared relevant data to the generative AI. The provided data includes product images, existing video clips, and text information. Based on this input data, the generative AI generates a video script and a proposed structure. Specific data processing involves using natural language processing technology to create the script, and image processing technology to determine scene composition, text placement, and visual effects. The output is a video script and a proposed structure.

[0327] Step 4:

[0328] The server generates a simple preview video based on the generated script and proposed structure. Generating the preview video involves stitching together scenes, overlaying text, and adding simple visual effects. The script and proposed structure output by the generative AI are used as input data. This results in a preview video that users can watch and confirm.

[0329] Step 5:

[0330] The server uploads the preview video to the user's dashboard, makes it available for viewing, and notifies the user when the preview video is ready. The input data includes the preview video, and this is the output data provided to the user via the dashboard.

[0331] Step 6:

[0332] Users watch preview videos on the dashboard. While watching, the emotion engine measures the user's emotional data in real time. Specifically, the emotion engine analyzes and records the user's satisfaction, dissatisfaction, excitement, etc. from facial expressions and voice. This becomes the input data sent to the server as emotional data.

[0333] Step 7:

[0334] Users provide specific feedback on the preview video to the server via the dashboard. This feedback includes specific instructions such as shortening the introduction or explaining the benefits in more detail. In addition, emotional data measured by the emotion engine is also sent to the server. This becomes the input data received by the server.

[0335] Step 8:

[0336] The server receives user feedback and emotional data and reflects it in the generative AI. The generative AI generates a revised version of the video based on this input data. Specific data calculations include changes to the script and composition plan based on the feedback, and scene editing using the emotional data. This results in the output of a revised version of the video.

[0337] Step 9:

[0338] The server uploads the generated modified video to the user's dashboard, where it is available for viewing and download, and notifies the user that the modified video is ready. The input data includes the modified video, which is the output data provided to the user.

[0339] Step 10:

[0340] The user can then re-watch the revised video from the dashboard for a final check. If necessary, they can provide further feedback. If there is further feedback, the server receives it and reflects it back into the generative AI. By repeating this process, the final video is completed.

[0341] Step 11:

[0342] The server generates the final video and uploads it to the user's dashboard. The server notifies the user when the final video is ready. The input data includes the final script, plot plan, and edited footage, and the output data is the final video.

[0343] Step 12:

[0344] Users will receive a notification and can download or stream the final video from their dashboard, allowing them to use the final, high-quality video for promotional or other uses.

[0345] Through the above processing steps, the system enables users to quickly create high-quality, personalized videos without having specialized knowledge or skills.

[0346] (Application example 2)

[0347] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0348] Conventional video production systems require non-expert users to have high technical knowledge and skills to quickly create high-quality videos, which requires a lot of time and effort. It is also difficult to create videos that reflect the user's emotions, making it difficult to generate content that satisfies the user. Furthermore, they are unable to quickly generate product introduction videos, especially in virtual stores, making it difficult to effectively promote products.

[0349] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0350] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a proposed composition based on the provided data; a means including an emotion engine that recognizes the user's emotions and reflects them in the video production process; and a means for generating and providing product introduction videos within a virtual store. This allows users to quickly create high-quality videos without non-specialized knowledge or skills, and the generation of personalized content that reflects emotions improves user satisfaction. Furthermore, the rapid generation of product introduction videos within a virtual store enables effective product advertising.

[0351] "Means for users to submit video production requests" refers to an interface through which users can input the information required to request video production and send it to the server. This may include a web portal, a smartphone app, or dedicated software.

[0352] "Means by which the server receives a user's request and provides relevant data to the generative AI" refers to the process by which the server receives a request sent by a user and appropriately conveys it to the generative AI, including querying a database or retrieving data from an external data source.

[0353] "Means for generative AI to generate a video script and composition plan based on provided data" refers to the process in which generative AI analyzes user request data provided by a server and, based on that, determines the video's scene composition, text placement, visual effects, etc.

[0354] "Means for the server to generate a simple preview video based on the generated composition plan and provide it to the user" refers to the process by which the server creates a visualized preview video based on the composition plan of the video created by the generative artificial intelligence and presents it to the user.

[0355] The "means for users to provide feedback on the preview video" refers to an interface that allows users to view the preview video and send their opinions or requests for corrections to the server. This includes a comment function and a rating system.

[0356] "Means for the generative AI to generate the final version of the video based on the feedback" refers to the process by which the generative AI receives feedback from users and creates the final version of the video that reflects that feedback. This process includes analyzing the user's emotional data using an emotion engine.

[0357] The "means by which the server provides the final video to the user" refers to the process by which the server generates and provides the final video file to the user through an interface such as a dashboard, which the user can then view by downloading or streaming.

[0358] The "Emotion Engine" is a system that measures the emotions of users when they watch videos in real time and collects the data. This makes it possible to create content that reflects the user's emotions during the video generation process.

[0359] The "means for generating and providing a product introduction video within a virtual store" refers to a means for generating an introduction video for a product selected by a user within a virtual store environment, visualizing it on the spot, and providing it to the user, thereby enabling the user to instantly understand the content of the product within the virtual store.

[0360] This invention relates to a system that allows users to quickly and easily create high-quality videos, and in particular, aims to facilitate the creation and provision of product introduction videos in a virtual store. This system has the following configuration and functions.

[0361] Users submit requests for video production using a dedicated application that runs on a smartphone or head-mounted display. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials used, etc.). This communicates the specifications of the video the user requires to the system.

[0362] The server receives requests sent by users, analyzes them, and provides relevant data to a generative artificial intelligence (AI). Based on the information provided by the server, the generative AI generates a video script and a proposed structure. The generative AI uses natural language processing and image processing technologies to determine scene composition, text placement, visual effects, and more. It also obtains additional information from external data sources (e.g., news articles and social media posts) as needed and incorporates it into the script and proposed structure.

[0363] As part of the generation process, an emotion engine analyzes the user's reactions and emotions and provides them to the generative AI, which then personalizes the video content based on the user's preferences. The emotion engine analyzes facial and voice data in real time as the user types their request.

[0364] Next, the server generates a simple preview video based on the generated script and plot plan and provides it to the user. The user can watch the preview video in the application, and the emotions they felt while watching it (e.g., excitement, dissatisfaction, anticipation, etc.) are recorded. The user provides feedback on the preview video and sends it to the server via the dashboard. This feedback also includes emotional data measured by the emotion engine.

[0365] The server receives user feedback and emotion data, and the generative AI retrains based on that data to generate the final video. The generative AI edits the video to emphasize the positive emotions felt by the user and reduce negative emotions. Finally, the server uploads the final video to the user's dashboard and sends a notification. After receiving the notification, the user can download or stream the final video from their dashboard.

[0366] A specific use case is creating a video to introduce a new drone product in a virtual store. Users can submit a request such as, "Please create a promotional video for our new drone product, highlighting its features and benefits. Please make it appealing and reflect the user's emotions." Based on this request, the generative AI creates the video, and the emotion engine is used to provide a video that reflects the user's feedback.

[0367] This system allows users to quickly create high-quality videos without specialized knowledge or skills, and generates personalized content that reflects emotions, improving user satisfaction.It also enables the rapid generation of product introduction videos in virtual stores, enabling effective product advertising.

[0368] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0369] Step 1:

[0370] The user submits a request for video production using a dedicated application.

[0371] Input: The user enters basic information into the application, such as the title, purpose, summary, and materials used.

[0372] Specific operation: The information entered by the user is sent to the server by clicking the "Submit" button.

[0373] Output: The input request information is transmitted to the server.

[0374] Step 2:

[0375] The server receives the user's request and provides the relevant data to the generative artificial intelligence.

[0376] Input: The request data sent by the user.

[0377] What happens: The server analyzes the request data and collects relevant data (e.g., images and text from a database, or additional information from external data sources).

[0378] Output: A relevant dataset that is fed into the generative artificial intelligence.

[0379] Step 3:

[0380] Generative AI generates a video script and structure based on the data provided.

[0381] Input: Relevant data provided by the server.

[0382] How it works: Generative AI uses natural language processing and image processing technologies to determine scene composition, text placement, and visual effects.

[0383] Output: Video script and outline.

[0384] Step 4:

[0385] The server generates a simple preview video based on the generated configuration plan and provides it to the user.

[0386] Input: A script and plot draft created by a generative artificial intelligence.

[0387] What it does: The server generates a preview video by stitching scenes, overlaying text, and combining basic visual effects, then uploads the preview video to the user's dashboard and sends a notification.

[0388] Output: The preview video created.

[0389] Step 5:

[0390] The user provides feedback on the preview video.

[0391] Input: User preview video views and reactions.

[0392] How it works: Users use the dashboard to enter feedback, submit comments and correction requests, and the emotion engine measures and captures their reactions (e.g., facial expressions and voice) as data.

[0393] Output: Feedback and emotion data.

[0394] Step 6:

[0395] Based on the feedback, the server uses generative artificial intelligence to generate the final version of the video.

[0396] Input: User-provided feedback and sentiment data.

[0397] How it works: The server passes this data to a generative AI, which then re-edits and re-generates the video, emphasizing positive emotions and reducing negative ones.

[0398] Output: The final video.

[0399] Step 7:

[0400] The server provides the final video to the user.

[0401] Input: The final video created by the generative artificial intelligence.

[0402] What happens: The server uploads the final video to the user's dashboard and sends a completion notification, which the user can then download or stream.

[0403] Output: Final video and completion notification.

[0404] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0405] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0406] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0407] [Second embodiment]

[0408] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0409] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0410] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0411] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0412] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0413] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0414] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0415] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0416] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0417] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0418] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0419] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0420] This invention relates to a system that enables anyone to easily and quickly create high-quality videos. This system involves a process in which a user submits a video production request, a server receives the request and provides data to a generative artificial intelligence (AI), the AI ​​generates a script and a draft structure for the video, the server provides a preview version to the user, and finally generates and provides the final version of the video that incorporates user feedback.

[0421] Specifically, this includes the following processes:

[0422] Users submit requests for video production using a dedicated application or web portal. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[0423] The server analyzes the requests received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[0424] Generative AI generates a video script and structure proposal based on the information provided. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0425] The server generates a simple preview video based on the proposed composition by the generative AI and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[0426] Users can view the preview video and provide feedback and corrections as needed. This feedback is entered via the user's dashboard. Users can provide correction instructions for specific scenes or text.

[0427] The server receives user feedback and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user feedback.

[0428] Finally, the server provides the final video to the user, which is stored on the server and available for download or streaming from the user's dashboard.

[0429] Specific examples

[0430] Example 1: Corporate PR video production

[0431] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[0432] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0433] Based on this information, generative AI creates a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending."

[0434] The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. The user reviews the preview video and provides feedback such as "Make the introduction a little shorter" or "Explain the benefits in more detail."

[0435] The server receives user feedback, and the generative AI generates the final video that reflects the feedback.

[0436] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[0437] In this way, this invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing generative AI, the video production process can be automated, significantly improving work efficiency.

[0438] The processing flow will be explained below.

[0439] Step 1:

[0440] Users access a dedicated application or web portal to submit a video production request. Specifically, users enter the required information, such as the title, purpose, summary of the content, and the materials they want to use, and then click the "Submit Request" button.

[0441] Step 2:

[0442] The server analyzes the request received from the user and collects relevant data based on the request information (e.g., materials from internal databases, additional information from external data sources, etc.).

[0443] Step 3:

[0444] The server provides the collected relevant data to the generative artificial intelligence (AI), including images, text, and existing video clips.

[0445] Step 4:

[0446] Generative AI generates a video script and structure proposal based on the provided data. Specifically, it uses natural language processing technology to generate text descriptions for each scene, image processing technology to determine the placement and effects of images, and obtains additional information from external data sources as needed.

[0447] Step 5:

[0448] The server generates a simple preview video based on the generated script and plot plan, which includes scene stitching, text overlays, and some visual effects.

[0449] Step 6:

[0450] The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready, after which the user can access the dashboard to view the preview video.

[0451] Step 7:

[0452] Users can check the preview video and provide feedback and correction requests through the dashboard. Specifically, users can enter comments for each scene and text, and specify additional corrections and improvements.

[0453] Step 8:

[0454] The server receives feedback from users and reflects it in the generative AI. The server analyzes the feedback and notifies the generative AI of any necessary changes.

[0455] Step 9:

[0456] The generative AI incorporates the feedback it receives and generates the final video. The AI ​​then learns from the data again and makes adjustments based on the user's requests.

[0457] Step 10:

[0458] The server uploads the final video to the user's dashboard and notifies the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[0459] In this way, collaboration between users, servers, and generative AI streamlines the video production process, enabling anyone to quickly and easily create high-quality videos.

[0460] Example 1

[0461] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0462] Traditionally, video production required specialized knowledge and skills, making it difficult for non-expert users to create high-quality videos in a short amount of time. Furthermore, the feedback process was inefficient, making it difficult to incorporate user feedback in real time. This resulted in increased costs and time for video production, and reduced operational efficiency.

[0463] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0464] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a draft structure based on the provided data; a means for the server to generate a simple preview video based on the generated draft structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for the generative artificial intelligence to generate a final version of the video based on the feedback; a means for the server to provide the final version of the video to the user; a means for the user to input and send feedback in real time; and a means for the server to receive the user's real-time feedback and reflect it in the generative artificial intelligence. This enables even non-expert users to produce high-quality videos in a short amount of time, and makes the video production process more efficient by reflecting feedback in real time.

[0465] "User" refers to the person who submits a video production request and provides feedback.

[0466] "Request" refers to information regarding requests and specifications submitted by a User for video production.

[0467] "Server" refers to a device or system that receives user requests, provides relevant data to the generative artificial intelligence, and manages the generated videos and feedback.

[0468] "Generative AI" refers to an AI technology that generates video scripts and plot plans based on provided data.

[0469] "Associated data" refers to information necessary for video production, such as text, images, video clips, audio narration, etc.

[0470] A "script" refers to a document or text that specifically outlines the scene structure and content of a video.

[0471] A "synopsis plan" refers to a plan that shows the overall structure of the video and details of each scene.

[0472] A "preview video" refers to a simple video created based on the generated composition plan.

[0473] "Feedback" refers to requests for corrections or opinions provided by users regarding preview videos.

[0474] "Final video" refers to a high-quality video that has been completed with user feedback reflected.

[0475] "Real-time feedback" refers to feedback sent by a user immediately on the spot.

[0476] The present invention relates to a system that allows users to easily and quickly create high-quality videos. This system involves a process in which a user submits a video production request, a server receives the request and provides data to a generative artificial intelligence (AI), the AI ​​generates a script and a draft structure for the video, the server provides a preview version to the user, and finally generates and provides the final version of the video that incorporates user feedback.

[0477] Specifically, it is configured as follows:

[0478] First, a user submits a request for video production using a dedicated application or web portal. In this request, the user enters basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[0479] The server then analyzes the request received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles, social media posts, etc.). The generative AI used by the server includes, for example, OpenAI's GPT-4 and AI with image processing technology.

[0480] Generative AI generates a video script and structure proposal based on the information provided. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0481] The server then generates a simple preview video based on the proposed composition by the generative AI and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server then uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[0482] Users can preview the video and provide feedback or suggestions for corrections as needed. This feedback is entered via the user's dashboard. Users can also provide correction suggestions for specific scenes or text and provide feedback in real time.

[0483] The server receives user feedback and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user feedback.

[0484] Finally, the server provides the final video to the user, which is stored on the server and available for download or streaming from the user's dashboard.

[0485] Specific examples

[0486] Example 1: Corporate PR video production

[0487] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[0488] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0489] Based on this information, generative AI creates a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending."

[0490] The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. The user reviews the preview video and provides feedback such as "Make the introduction a little shorter" or "Explain the benefits in more detail."

[0491] The server receives user feedback, and the generative AI generates the final video that reflects the feedback.

[0492] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[0493] Prompt Sentence Examples

[0494] "We would like to create a video to introduce a new product. The target users are businessmen in their 30s living in urban areas. The product's features are durability and beautiful design. Could you create a promotional video that emphasizes these points?"

[0495] In this way, this invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing generative AI, the video production process can be automated, significantly improving work efficiency.

[0496] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0497] Step 1:

[0498] A user submits a request for video production.

[0499] Specific operation: The user accesses a dedicated application or web portal and enters basic information such as the title, purpose, content summary, and materials to be used in the video production request form. When the user clicks the send button, the request is sent to the server.

[0500] Input: Information such as title "New product introduction", purpose "Customer promotion", and content summary "Highlight the features and benefits of the new product".

[0501] Output: The request data sent to the server.

[0502] Step 2:

[0503] The server receives and analyzes the user's request.

[0504] Specific operation: The server receives the request information from the user, analyzes its content, extracts the necessary relevant data based on the request content, and collects the necessary information from internal databases and external data sources.

[0505] Input: User request data.

[0506] Output: Relevant data to feed into the generative AI (e.g., product images, existing promotional video clips, text information).

[0507] Step 3:

[0508] Generative AI generates a video script and structure based on relevant data.

[0509] Specific operation: Based on the analysis results, the server provides prompts and related data to the generative AI. The generative AI then uses the provided information to create a video script and structure. The AI ​​uses natural language processing and image processing techniques to determine scene composition, text placement, and visual effects, and also generates text for the voice narration.

[0510] Input: relevant data, prompt statement.

[0511] Output: A video script and outline.

[0512] Step 4:

[0513] The server generates a preview video based on the generated configuration plan and provides it to the user.

[0514] Specific operation: The server generates a simple preview video based on the script and composition plan provided by the generative AI. The preview video includes scene stitching, text overlays, etc. The server uploads the generated preview video to the user's dashboard and notifies the user that it is ready.

[0515] Input: Video script and outline.

[0516] Output: The preview video that is provided to the user.

[0517] Step 5:

[0518] The user provides feedback on the preview video.

[0519] Specific operations: Users can check the preview video on the dashboard, input correction requests and feedback, provide correction instructions for specific scenes and text, and submit. Users can also submit feedback in real time.

[0520] Input: Preview video, user feedback.

[0521] Output: Feedback data sent to the server.

[0522] Step 6:

[0523] The server receives the feedback and reflects it in the generative AI.

[0524] Specific operation: The server receives feedback from the user and sends it to the generative AI. The generative AI reflects the feedback and revises the video script and composition plan. The server generates a new preview video based on the revised composition plan and provides it to the user again.

[0525] Input: User feedback, revised script and structure of the video.

[0526] Output: Revised preview video.

[0527] Step 7:

[0528] The server provides the final video to the user.

[0529] How it works: The server generates a high-quality final video based on the final script and plot plan. The generated video is stored on the server and uploaded to the user's dashboard. The user can then watch the final video by downloading or streaming.

[0530] Input: Final video script and outline.

[0531] Output: The final video that is delivered to the user.

[0532] (Application example 1)

[0533] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0534] Conventional video production systems have made it difficult for users without specialized knowledge and advanced skills to quickly produce high-quality videos. Furthermore, they lacked an interface for individual users to easily create and distribute personalized video content. Furthermore, they lacked a prompt input mechanism for generating video composition plans customized to the user's needs.

[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0536] In this invention, the server includes: means for a user to send a video production request; means for providing related data to a generative artificial intelligence; means for generating a video script and a proposed composition based on the provided data; means for generating a simple preview video based on the proposed composition and providing it to the user; means for the user to provide feedback on the preview video; means for generating a final version of the video based on the feedback; means for providing the final version of the video to the user; means for the user to upload the generated video to a distribution service; and means for the user to customize the proposed composition by inputting a prompt. This makes it possible to quickly and easily create high-quality videos and easily distribute personalized video content without specialized knowledge or advanced technology.

[0537] "Means for users to submit video production requests" means the ability for users to use their own devices to input video production requests and details through a specific format or interface and send them to the server.

[0538] "Generative AI" refers to AI technology that automatically generates video scripts and plots based on provided data and information. It utilizes natural language processing and image processing technologies.

[0539] A "simple preview video" is a prototype video generated based on an initial video script and structure created by generative artificial intelligence, for users to check and provide feedback.

[0540] The "means for users to provide feedback on the preview video" refers to an interface or function that allows users to check the preview video, input requests for improvement or opinions about the content, and send the input to the server.

[0541] The "means of generating the final version of the video" is a function in which the generative artificial intelligence re-edits and re-structures the video based on feedback provided by the user, generating the final, completed version of the video.

[0542] The "means by which the server provides the final video to the user" refers to the function by which the server stores the generated final video and provides it for download or streaming through an interface accessible to the user.

[0543] "Means for uploading user-generated videos to a distribution service" refers to a function that allows users to easily upload and publish completed videos to a content distribution platform.

[0544] A "prompt sentence" is a text-based sentence that a user uses to input specific instructions or requests to a generative artificial intelligence when generating a script or plot plan for a video.

[0545] The present invention relates to a system that allows anyone to easily create high-quality videos and upload them to a content distribution service. Hereinafter, an embodiment of the invention will be described in detail.

[0546] The system mainly consists of a terminal where users send requests, a server that receives and processes video production requests, generative artificial intelligence (AI), and related databases and external data sources. The following describes a specific implementation of the process in which a user makes a video production request, AI generates a video script and composition plan based on that request, and ultimately provides a high-quality video.

[0547] 1. User submits request:

[0548] Users submit video production requests using a dedicated smartphone application or web portal, where they enter details such as the title of the video they want to create, its purpose, a summary of the content, and the materials they will use.

[0549] 2. Server reception and AI provision:

[0550] The request sent by the user is received by the server, which analyzes the information and provides the generative AI with relevant data, which is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[0551] 3. Video script and structure generation:

[0552] Generative AI generates a video script and structure proposal based on the provided data. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0553] 4. Generate preview video:

[0554] The server creates a simple preview video based on the generated composition plan and provides it to the user. This preview video includes scene stitching, text overlays, etc., so that the user can check it.

[0555] 5. User feedback processing:

[0556] Users can view the preview video and provide feedback and suggestions for revisions, which are entered via the user's dashboard. This feedback can include specific instructions such as "make the introduction a little shorter" or "explain the benefits in more detail."

[0557] 6. Generate the final video:

[0558] The server receives user feedback and applies it to the generative AI. The AI ​​then retrains and generates a revised version of the video that reflects the user's feedback. Finally, the server provides the final version of the video to the user.

[0559] 7. Uploading and Publishing Videos:

[0560] Users can upload the generated video to a content distribution service, and can customize the video's structure by entering specific prompts.

[0561] Hardware and software used:

[0562] Hardware: Smartphones, server computers

[0563] Software: Python, Flask (for server processing), generative AI module, video editing module

[0564] Examples:

[0565] An example prompt for a user to create a promotional video for a new product:

[0566] Title: New Product Review

[0567] Purpose: To inform the audience

[0568] Summary: Highlight the features and benefits of your new product and compare it with competing products. Include specific usage scenarios if possible.

[0569] Using this prompt, the generative AI generates specific and effective video content.

[0570] As described above, this system makes it easy for anyone to create and distribute high-quality videos, automating the video production process and significantly improving work efficiency.

[0571] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0572] Step 1:

[0573] Users submit video production requests via a dedicated smartphone application or web portal.

[0574] Input: Information about the video title, purpose, summary, and materials used that the user enters into the application.

[0575] Specific operation: When the user enters the required information into the form and clicks the submit button, this data is sent from the application to the server.

[0576] Output: The user's request information arrives at the server.

[0577] Step 2:

[0578] The server receives the user's request, analyzes it, and provides the relevant data to the generative artificial intelligence.

[0579] Input: Request information sent by the user.

[0580] What it does: The server analyzes the request and gathers the necessary data from its internal database. If necessary, it retrieves additional data from external sources (e.g., news articles or social media posts) and provides it to the generative AI.

[0581] Output: Relevant data is provided to the generative artificial intelligence.

[0582] Step 3:

[0583] Generative AI generates a video script and structure based on the data provided.

[0584] Input: Relevant data provided to the generative artificial intelligence.

[0585] Specific operation: AI uses natural language processing and image processing technology to determine the video's scene composition, text placement, visual effects, etc., and also generates text for the voice narration.

[0586] Output: A script and outline of the generated video.

[0587] Step 4:

[0588] The server generates a simple preview video based on the generated configuration plan and provides it to the user.

[0589] Input: A video script and plot plan generated by generative artificial intelligence.

[0590] What happens: The server uses the video editing module to create a preview video, stitching scenes together and adding text overlays, uploading the generated preview video to the user's dashboard, and sending a notification that it's ready.

[0591] Output: A simple preview video provided to the user.

[0592] Step 5:

[0593] The user provides feedback on the preview video.

[0594] Input: The preview video the user watched and their feedback (e.g., a shorter introduction, a more detailed explanation of the benefits, etc.).

[0595] Specific action: A user enters comments or correction requests into the feedback form on the dashboard and clicks the submit button.

[0596] Output: User feedback reaches the server.

[0597] Step 6:

[0598] Based on the feedback, generative artificial intelligence generates the final version of the video.

[0599] Input: Feedback information from the user.

[0600] How it works: The server provides feedback to the generative AI, which then re-learns the content and generates a revised script and structure based on the feedback. The server then uses the video editing module to generate the final video.

[0601] Output: The final video generated.

[0602] Step 7:

[0603] The server provides the final video to the user.

[0604] Input: The final video.

[0605] What it does: Saves the final video to our servers and provides a link for users to download or stream it from their dashboard. Notifies users that their video is complete.

[0606] Output: The final video that is delivered to the user.

[0607] Step 8:

[0608] User-generated videos can be uploaded to content distribution services.

[0609] Input: Final video and streaming service account information.

[0610] What it does: Users click the "Upload" button in the application, follow a few simple steps, and the final video is uploaded to the distribution platform of their choice.

[0611] Output: The video published to a distribution platform.

[0612] Step 9:

[0613] Users can customize the video composition by entering a prompt.

[0614] Input: The specific prompt text that the user enters.

[0615] How it works: A user enters a prompt into a dashboard or application, which is then provided to the generative AI, which then generates content based on the prompt and customizes it to fit the user's needs.

[0616] Output: A video based on your customized composition.

[0617] By following the above steps, users can quickly create and distribute high-quality videos without having specialized knowledge.

[0618] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0619] The present invention relates to a system that enables users to easily and quickly create high-quality videos. This system involves a process in which the user submits a video production request, a server receives the request, and provides the data to a generative artificial intelligence (AI). The AI ​​then generates a script and a draft structure for the video, and the server provides a preview version to the user. Finally, the system generates and provides a final version of the video that incorporates user feedback. Furthermore, the present invention aims to achieve high satisfaction by incorporating an emotion engine that recognizes user emotions, thereby reflecting the user's emotions in the video production process.

[0620] Specifically, this includes the following processes:

[0621] Users submit requests for video production using a dedicated application or web portal. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[0622] The server analyzes the requests received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[0623] The generative AI generates a video script and structure proposal based on the information provided. The AI ​​uses natural language processing and image processing techniques to determine the video's scene composition, text placement, visual effects, etc. It also obtains additional information from external data sources as needed. Here, the emotion engine analyzes the user's reactions and emotions and provides this information to the generative AI, which then generates content that is more tailored to the user.

[0624] The server generates a simple preview video based on the generated script and plot plan and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[0625] When users watch the preview video, the emotion engine measures their reactions in real time. The emotions (e.g., satisfaction, dissatisfaction, excitement, etc.) felt by the user while watching the preview video are recorded. Users can provide feedback via the dashboard. Specifically, users can enter comments for each scene or text and specify additional corrections or improvements. This feedback also includes the emotional data measured by the emotion engine.

[0626] The server receives user feedback and emotional data and reflects it in the generative AI. The generative AI then retrains and generates a revised version of the video that reflects the user's feedback and emotions. The AI ​​then edits the video to emphasize the positive emotions felt by the user and reduce negative emotions.

[0627] Finally, the server uploads the final video to the user's dashboard and sends a notification to the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[0628] Specific examples

[0629] Example 1: Corporate PR video production

[0630] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[0631] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0632] The generative AI uses this information to create a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending." The emotion engine analyzes the user's emotions and provides them to the generative AI, which then reflects the content according to the user's preferences.

[0633] The server creates a preview video based on the generated configuration and provides it to the user via a dashboard. The emotions (e.g., excitement, anticipation, doubt) that arise when the user watches the preview video are recorded.

[0634] Users provide feedback such as "Make the introduction shorter" or "Explain the benefits in more detail," and data from the sentiment engine is also sent.

[0635] The server receives user feedback, and the generative AI reflects the feedback and emotional data to generate the final video.

[0636] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[0637] In this way, the present invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing the emotion engine, more personalized content that reflects the user's emotions is generated, thereby improving user satisfaction.

[0638] The processing flow will be explained below.

[0639] Step 1:

[0640] Users access a dedicated application or web portal to submit a video production request. Specifically, users enter the required information, such as the title, purpose, summary of the content, and the materials they want to use, and then click the "Submit Request" button.

[0641] Step 2:

[0642] The server analyzes the request received from the user and collects relevant data based on the request information (e.g., materials from internal databases, additional information from external data sources, etc.).

[0643] Step 3:

[0644] The server provides the collected relevant data to the generative artificial intelligence (AI), including images, text, and existing video clips.

[0645] Step 4:

[0646] Generative AI generates a video script and structure proposal based on the provided data. Specifically, it uses natural language processing technology to generate text descriptions for each scene, image processing technology to determine the placement and effects of images, and obtains additional information from external data sources as needed.

[0647] Step 5:

[0648] The emotion engine measures users' reactions in real time. While users are submitting video creation requests, the emotion engine analyzes their facial expressions, voice, and other biometric information to generate emotion data.

[0649] Step 6:

[0650] The server generates a simple preview video based on the generated script and plot plan, which includes scene stitching, text overlays, and some visual effects.

[0651] Step 7:

[0652] The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready, after which the user can access the dashboard to view the preview video.

[0653] Step 8:

[0654] Users can view the preview video and provide feedback and correction requests via the dashboard. Specifically, users can enter comments for each scene and text, specifying additional corrections and improvements. This feedback also includes emotional data measured by the emotion engine.

[0655] Step 9:

[0656] The server receives feedback and emotion data from users and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user's feedback and emotion.

[0657] Step 10:

[0658] The server uploads the final video to the user's dashboard and notifies the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[0659] In this way, collaboration between users, servers, generative AI, and the emotion engine allows even non-experts to quickly and easily create high-quality videos. Furthermore, the inclusion of the emotion engine allows users' emotions to be reflected in the video production process, enabling the provision of more personalized content with a high level of satisfaction.

[0660] Example 2

[0661] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0662] In the past, users needed specialized knowledge and skills to quickly create high-quality videos. This meant that video production required a lot of time and money, making it difficult for average users to achieve this. It was also difficult to personalize the video content to match the user's emotions, making it difficult to increase user satisfaction.

[0663] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0664] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a draft structure based on the provided data; a means for the server to generate a simple preview video based on the generated draft structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for an emotion engine to analyze and acquire user emotion data; a means for the generative artificial intelligence to generate a revised version of the video based on the feedback and emotion data; a means for the server to provide the revised version of the video to the user; and a means for the server to generate a final version of the video and provide it to the user. This enables users to quickly create high-quality videos without specialized knowledge or skills, and further enables personalization of video content that reflects the user's emotions, thereby improving user satisfaction.

[0665] "User" refers to an individual or organization that uses the System to make a video production request.

[0666] "Server" refers to the computer system that receives and processes user requests, provides relevant data to the generative artificial intelligence, and generates videos and manages feedback.

[0667] "Generative AI" refers to AI technology that generates a video script and structure based on provided data, and then creates revised and final versions of the video.

[0668] An "emotion engine" refers to a technology that analyzes user reactions and emotions, acquires them as data, and provides this to generative artificial intelligence.

[0669] A "video production request" refers to the action of sending information to the server, including basic information about the video the user wants to produce (e.g., title, purpose, summary of content, materials to be used, etc.).

[0670] "Related Data" means data necessary to generate a video script and story (e.g., user-provided information, data from internal databases, data from external data sources).

[0671] A "script" refers to the story or script of a video that generative artificial intelligence creates based on the data provided.

[0672] "Composition proposal" refers to the scene composition and layout proposal for a video that is determined by generative artificial intelligence based on the data provided.

[0673] "Preview video" refers to a simple video generated by the server based on the script and plot plan created by generative AI, which is used by users to confirm and provide feedback.

[0674] "Feedback" refers to opinions and correction requests provided by users regarding preview videos.

[0675] "Revised video" refers to a video edited by generative artificial intelligence based on user feedback and emotional data.

[0676] "Final video" refers to the final completed video created by the generative artificial intelligence that the server provides to the user.

[0677] "Database" refers to a system that resides within a server and stores and manages related data.

[0678] "External data sources" refers to information sources that can be obtained from outside, such as news articles and social media posts.

[0679] This invention relates to a system that enables users to quickly create high-quality videos without specialized knowledge or skills. This system is composed of a server, generative artificial intelligence, an emotion engine, a dedicated application, and a web portal.

[0680] Hardware and software used

[0681] Hardware: Servers (computer systems with high-performance computing resources), terminals (user computers, smartphones, etc.)

[0682] Software: Dedicated applications, web portals, generative AI (e.g., OpenAI's GPT model, image generation model, etc.), emotion recognition engines (e.g., Affectiva, Microsoft Azure Emotion API, etc.)

[0683] server

[0684] The server receives video production requests submitted by users using a dedicated application or web portal. The request includes basic information such as the video title, purpose, summary, and materials used. The server analyzes this information and prepares relevant data to provide to the generative artificial intelligence. This relevant data includes information from the server's internal database and information from external data sources (such as news articles and social media posts).

[0685] Generative Artificial Intelligence

[0686] The generative AI generates a video script and structure proposal based on data provided by the server. It uses natural language processing and image processing techniques to determine the video's scene composition, text placement, visual effects, etc. It also uses data from the emotion engine to generate personalized content that reflects the user's emotions.

[0687] Emotion Engine

[0688] The emotion engine measures and analyzes the user's emotions in real time while they are watching the preview video. Emotional data such as satisfaction, dissatisfaction, and excitement felt by the user is recorded and provided to the generative AI. This allows content that reflects the user's emotions to be generated.

[0689] Dedicated application / web portal

[0690] Users can use a dedicated application or web portal to submit video production requests, view the generated preview videos and final videos, and provide specific feedback on the preview videos to the server through the application or web portal.

[0691] Example: Production of corporate PR videos

[0692] 1. A user requests the creation of a promotional video for a new product. At that time, the user enters information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and benefits of the new product."

[0693] 2. The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0694] 3. The generative AI uses this information to create a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending." The emotion engine analyzes the user's emotional data and provides it to the generative AI, which then reflects the content according to the user's preferences.

[0695] 4. The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. As the user watches the video, the emotion engine records emotions such as excitement, anticipation, and doubt.

[0696] 5. The user provides feedback such as a short introduction and detailed benefit explanation, along with sentiment engine data.

[0697] 6. The server receives the user feedback, and the generative AI generates a revised version of the video based on the feedback and emotional data.

[0698] 7. Once the final video is completed, the server will provide it to the user, who can then download the high-quality final PR video from their dashboard and use it for promotional purposes.

[0699] Example prompts for generative AI models

[0700] "Write a script for a promotional video highlighting the features of a new product. The title should be 'New Product Introduction' and the purpose should be to promote it to customers. Use the following information: product images, text information, and existing promotional video clips."

[0701] "Edit your videos to reflect your users' emotional data and build excitement and anticipation."

[0702] In this way, the present invention allows users to quickly and easily create high-quality videos. In addition, by using an emotion engine, personalized video content that reflects the user's emotions is generated, improving user satisfaction.

[0703] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0704] Step 1:

[0705] The user submits a request for video production using a dedicated application or web portal. The user enters basic information such as the video title, purpose, and content summary, and clicks the submit button. The input data includes the video title "New product introduction," purpose "Customer promotion," and content summary "Highlighting the features and benefits of the new product." This is then sent to the server.

[0706] Step 2:

[0707] The server receives the request sent by the user and analyzes the content. The analyzed data includes the title, purpose, and content summary. Based on this data, the server prepares related data to provide to the generative artificial intelligence (AI). Specifically, it collects product images, existing promotional video clips, and text information from the server's internal database. This becomes the input data provided to the generative AI.

[0708] Step 3:

[0709] The server provides the prepared relevant data to the generative AI. The provided data includes product images, existing video clips, and text information. Based on this input data, the generative AI generates a video script and a proposed structure. Specific data processing involves using natural language processing technology to create the script, and image processing technology to determine scene composition, text placement, and visual effects. The output is a video script and a proposed structure.

[0710] Step 4:

[0711] The server generates a simple preview video based on the generated script and proposed structure. Generating the preview video involves stitching together scenes, overlaying text, and adding simple visual effects. The script and proposed structure output by the generative AI are used as input data. This results in a preview video that users can watch and confirm.

[0712] Step 5:

[0713] The server uploads the preview video to the user's dashboard, makes it available for viewing, and notifies the user when the preview video is ready. The input data includes the preview video, and this is the output data provided to the user via the dashboard.

[0714] Step 6:

[0715] Users watch preview videos on the dashboard. While watching, the emotion engine measures the user's emotional data in real time. Specifically, the emotion engine analyzes and records the user's satisfaction, dissatisfaction, excitement, etc. from facial expressions and voice. This becomes the input data sent to the server as emotional data.

[0716] Step 7:

[0717] Users provide specific feedback on the preview video to the server via the dashboard. This feedback includes specific instructions such as shortening the introduction or explaining the benefits in more detail. In addition, emotional data measured by the emotion engine is also sent to the server. This becomes the input data received by the server.

[0718] Step 8:

[0719] The server receives user feedback and emotional data and reflects it in the generative AI. The generative AI generates a revised version of the video based on this input data. Specific data calculations include changes to the script and composition plan based on the feedback, and scene editing using the emotional data. This results in the output of a revised version of the video.

[0720] Step 9:

[0721] The server uploads the generated modified video to the user's dashboard, where it is available for viewing and download, and notifies the user that the modified video is ready. The input data includes the modified video, which is the output data provided to the user.

[0722] Step 10:

[0723] The user can then re-watch the revised video from the dashboard for a final check. If necessary, they can provide further feedback. If there is further feedback, the server receives it and reflects it back into the generative AI. By repeating this process, the final video is completed.

[0724] Step 11:

[0725] The server generates the final video and uploads it to the user's dashboard. The server notifies the user when the final video is ready. The input data includes the final script, plot plan, and edited footage, and the output data is the final video.

[0726] Step 12:

[0727] Users will receive a notification and can download or stream the final video from their dashboard, allowing them to use the final, high-quality video for promotional or other uses.

[0728] Through the above processing steps, the system enables users to quickly create high-quality, personalized videos without having specialized knowledge or skills.

[0729] (Application example 2)

[0730] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0731] Conventional video production systems require non-expert users to have high technical knowledge and skills to quickly create high-quality videos, which requires a lot of time and effort. It is also difficult to create videos that reflect the user's emotions, making it difficult to generate content that satisfies the user. Furthermore, they are unable to quickly generate product introduction videos, especially in virtual stores, making it difficult to effectively promote products.

[0732] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0733] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a proposed composition based on the provided data; a means including an emotion engine that recognizes the user's emotions and reflects them in the video production process; and a means for generating and providing product introduction videos within a virtual store. This allows users to quickly create high-quality videos without non-specialized knowledge or skills, and the generation of personalized content that reflects emotions improves user satisfaction. Furthermore, the rapid generation of product introduction videos within a virtual store enables effective product advertising.

[0734] "Means for users to submit video production requests" refers to an interface through which users can input the information required to request video production and send it to the server. This may include a web portal, a smartphone app, or dedicated software.

[0735] "Means by which the server receives a user's request and provides relevant data to the generative AI" refers to the process by which the server receives a request sent by a user and appropriately conveys it to the generative AI, including querying a database or retrieving data from an external data source.

[0736] "Means for generative AI to generate a video script and composition plan based on provided data" refers to the process in which generative AI analyzes user request data provided by a server and, based on that, determines the video's scene composition, text placement, visual effects, etc.

[0737] "Means for the server to generate a simple preview video based on the generated composition plan and provide it to the user" refers to the process by which the server creates a visualized preview video based on the composition plan of the video created by the generative artificial intelligence and presents it to the user.

[0738] The "means for users to provide feedback on the preview video" refers to an interface that allows users to view the preview video and send their opinions or requests for corrections to the server. This includes a comment function and a rating system.

[0739] "Means for the generative AI to generate the final version of the video based on the feedback" refers to the process by which the generative AI receives feedback from users and creates the final version of the video that reflects that feedback. This process includes analyzing the user's emotional data using an emotion engine.

[0740] The "means by which the server provides the final video to the user" refers to the process by which the server generates and provides the final video file to the user through an interface such as a dashboard, which the user can then view by downloading or streaming.

[0741] The "Emotion Engine" is a system that measures the emotions of users when they watch videos in real time and collects the data. This makes it possible to create content that reflects the user's emotions during the video generation process.

[0742] The "means for generating and providing a product introduction video within a virtual store" refers to a means for generating an introduction video for a product selected by a user within a virtual store environment, visualizing it on the spot, and providing it to the user, thereby enabling the user to instantly understand the content of the product within the virtual store.

[0743] This invention relates to a system that allows users to quickly and easily create high-quality videos, and in particular, aims to facilitate the creation and provision of product introduction videos in a virtual store. This system has the following configuration and functions.

[0744] Users submit requests for video production using a dedicated application that runs on a smartphone or head-mounted display. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials used, etc.). This communicates the specifications of the video the user requires to the system.

[0745] The server receives requests sent by users, analyzes them, and provides relevant data to a generative artificial intelligence (AI). Based on the information provided by the server, the generative AI generates a video script and a proposed structure. The generative AI uses natural language processing and image processing technologies to determine scene composition, text placement, visual effects, and more. It also obtains additional information from external data sources (e.g., news articles and social media posts) as needed and incorporates it into the script and proposed structure.

[0746] As part of the generation process, an emotion engine analyzes the user's reactions and emotions and provides them to the generative AI, which then personalizes the video content based on the user's preferences. The emotion engine analyzes facial and voice data in real time as the user types their request.

[0747] Next, the server generates a simple preview video based on the generated script and plot plan and provides it to the user. The user can watch the preview video in the application, and the emotions they felt while watching it (e.g., excitement, dissatisfaction, anticipation, etc.) are recorded. The user provides feedback on the preview video and sends it to the server via the dashboard. This feedback also includes emotional data measured by the emotion engine.

[0748] The server receives user feedback and emotion data, and the generative AI retrains based on that data to generate the final video. The generative AI edits the video to emphasize the positive emotions felt by the user and reduce negative emotions. Finally, the server uploads the final video to the user's dashboard and sends a notification. After receiving the notification, the user can download or stream the final video from their dashboard.

[0749] A specific use case is creating a video to introduce a new drone product in a virtual store. Users can submit a request such as, "Please create a promotional video for our new drone product, highlighting its features and benefits. Please make it appealing and reflect the user's emotions." Based on this request, the generative AI creates the video, and the emotion engine is used to provide a video that reflects the user's feedback.

[0750] This system allows users to quickly create high-quality videos without specialized knowledge or skills, and generates personalized content that reflects emotions, improving user satisfaction.It also enables the rapid generation of product introduction videos in virtual stores, enabling effective product advertising.

[0751] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0752] Step 1:

[0753] The user submits a request for video production using a dedicated application.

[0754] Input: The user enters basic information into the application, such as the title, purpose, summary, and materials used.

[0755] Specific operation: The information entered by the user is sent to the server by clicking the "Submit" button.

[0756] Output: The input request information is transmitted to the server.

[0757] Step 2:

[0758] The server receives the user's request and provides the relevant data to the generative artificial intelligence.

[0759] Input: The request data sent by the user.

[0760] What happens: The server analyzes the request data and collects relevant data (e.g., images and text from a database, or additional information from external data sources).

[0761] Output: A relevant dataset that is fed into the generative artificial intelligence.

[0762] Step 3:

[0763] Generative AI generates a video script and structure based on the data provided.

[0764] Input: Relevant data provided by the server.

[0765] How it works: Generative AI uses natural language processing and image processing technologies to determine scene composition, text placement, and visual effects.

[0766] Output: Video script and outline.

[0767] Step 4:

[0768] The server generates a simple preview video based on the generated configuration plan and provides it to the user.

[0769] Input: A script and plot draft created by a generative artificial intelligence.

[0770] What it does: The server generates a preview video by stitching scenes, overlaying text, and combining basic visual effects, then uploads the preview video to the user's dashboard and sends a notification.

[0771] Output: The preview video created.

[0772] Step 5:

[0773] The user provides feedback on the preview video.

[0774] Input: User preview video views and reactions.

[0775] How it works: Users use the dashboard to enter feedback, submit comments and correction requests, and the emotion engine measures and captures their reactions (e.g., facial expressions and voice) as data.

[0776] Output: Feedback and emotion data.

[0777] Step 6:

[0778] Based on the feedback, the server uses generative artificial intelligence to generate the final version of the video.

[0779] Input: User-provided feedback and sentiment data.

[0780] How it works: The server passes this data to a generative AI, which then re-edits and re-generates the video, emphasizing positive emotions and reducing negative ones.

[0781] Output: The final video.

[0782] Step 7:

[0783] The server provides the final video to the user.

[0784] Input: The final video created by the generative artificial intelligence.

[0785] What happens: The server uploads the final video to the user's dashboard and sends a completion notification, which the user can then download or stream.

[0786] Output: Final video and completion notification.

[0787] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0788] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0789] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0790] [Third embodiment]

[0791] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0792] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0793] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0794] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0795] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0796] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0797] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0798] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0799] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0800] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0801] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0802] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0803] This invention relates to a system that enables anyone to easily and quickly create high-quality videos. This system involves a process in which a user submits a video production request, a server receives the request and provides data to a generative artificial intelligence (AI), the AI ​​generates a script and a draft structure for the video, the server provides a preview version to the user, and finally generates and provides the final version of the video that incorporates user feedback.

[0804] Specifically, this includes the following processes:

[0805] Users submit requests for video production using a dedicated application or web portal. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[0806] The server analyzes the requests received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[0807] Generative AI generates a video script and structure proposal based on the information provided. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0808] The server generates a simple preview video based on the proposed composition by the generative AI and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[0809] Users can view the preview video and provide feedback and corrections as needed. This feedback is entered via the user's dashboard. Users can provide correction instructions for specific scenes or text.

[0810] The server receives user feedback and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user feedback.

[0811] Finally, the server provides the final video to the user, which is stored on the server and available for download or streaming from the user's dashboard.

[0812] Specific examples

[0813] Example 1: Corporate PR video production

[0814] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[0815] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0816] Based on this information, generative AI creates a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending."

[0817] The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. The user reviews the preview video and provides feedback such as "Make the introduction a little shorter" or "Explain the benefits in more detail."

[0818] The server receives user feedback, and the generative AI generates the final video that reflects the feedback.

[0819] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[0820] In this way, this invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing generative AI, the video production process can be automated, significantly improving work efficiency.

[0821] The processing flow will be explained below.

[0822] Step 1:

[0823] Users access a dedicated application or web portal to submit a video production request. Specifically, users enter the required information, such as the title, purpose, summary of the content, and the materials they want to use, and then click the "Submit Request" button.

[0824] Step 2:

[0825] The server analyzes the request received from the user and collects relevant data based on the request information (e.g., materials from internal databases, additional information from external data sources, etc.).

[0826] Step 3:

[0827] The server provides the collected relevant data to the generative artificial intelligence (AI), including images, text, and existing video clips.

[0828] Step 4:

[0829] Generative AI generates a video script and structure proposal based on the provided data. Specifically, it uses natural language processing technology to generate text descriptions for each scene, image processing technology to determine the placement and effects of images, and obtains additional information from external data sources as needed.

[0830] Step 5:

[0831] The server generates a simple preview video based on the generated script and plot plan, which includes scene stitching, text overlays, and some visual effects.

[0832] Step 6:

[0833] The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready, after which the user can access the dashboard to view the preview video.

[0834] Step 7:

[0835] Users can check the preview video and provide feedback and correction requests through the dashboard. Specifically, users can enter comments for each scene and text, and specify additional corrections and improvements.

[0836] Step 8:

[0837] The server receives feedback from users and reflects it in the generative AI. The server analyzes the feedback and notifies the generative AI of any necessary changes.

[0838] Step 9:

[0839] The generative AI incorporates the feedback it receives and generates the final video. The AI ​​then learns from the data again and makes adjustments based on the user's requests.

[0840] Step 10:

[0841] The server uploads the final video to the user's dashboard and notifies the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[0842] In this way, collaboration between users, servers, and generative AI streamlines the video production process, enabling anyone to quickly and easily create high-quality videos.

[0843] Example 1

[0844] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0845] Traditionally, video production required specialized knowledge and skills, making it difficult for non-expert users to create high-quality videos in a short amount of time. Furthermore, the feedback process was inefficient, making it difficult to incorporate user feedback in real time. This resulted in increased costs and time for video production, and reduced operational efficiency.

[0846] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0847] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a draft structure based on the provided data; a means for the server to generate a simple preview video based on the generated draft structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for the generative artificial intelligence to generate a final version of the video based on the feedback; a means for the server to provide the final version of the video to the user; a means for the user to input and send feedback in real time; and a means for the server to receive the user's real-time feedback and reflect it in the generative artificial intelligence. This enables even non-expert users to produce high-quality videos in a short amount of time, and makes the video production process more efficient by reflecting feedback in real time.

[0848] "User" refers to the person who submits a video production request and provides feedback.

[0849] "Request" refers to information regarding requests and specifications submitted by a User for video production.

[0850] "Server" refers to a device or system that receives user requests, provides relevant data to the generative artificial intelligence, and manages the generated videos and feedback.

[0851] "Generative AI" refers to an AI technology that generates video scripts and plot plans based on provided data.

[0852] "Associated data" refers to information necessary for video production, such as text, images, video clips, audio narration, etc.

[0853] A "script" refers to a document or text that specifically outlines the scene structure and content of a video.

[0854] A "synopsis plan" refers to a plan that shows the overall structure of the video and details of each scene.

[0855] A "preview video" refers to a simple video created based on the generated composition plan.

[0856] "Feedback" refers to requests for corrections or opinions provided by users regarding preview videos.

[0857] "Final video" refers to a high-quality video that has been completed with user feedback reflected.

[0858] "Real-time feedback" refers to feedback sent by a user immediately on the spot.

[0859] The present invention relates to a system that allows users to easily and quickly create high-quality videos. This system involves a process in which a user submits a video production request, a server receives the request and provides data to a generative artificial intelligence (AI), the AI ​​generates a script and a draft structure for the video, the server provides a preview version to the user, and finally generates and provides the final version of the video that incorporates user feedback.

[0860] Specifically, it is configured as follows:

[0861] First, a user submits a request for video production using a dedicated application or web portal. In this request, the user enters basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[0862] The server then analyzes the request received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles, social media posts, etc.). The generative AI used by the server includes, for example, OpenAI's GPT-4 and AI with image processing technology.

[0863] Generative AI generates a video script and structure proposal based on the information provided. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0864] The server then generates a simple preview video based on the proposed composition by the generative AI and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server then uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[0865] Users can preview the video and provide feedback or suggestions for corrections as needed. This feedback is entered via the user's dashboard. Users can also provide correction suggestions for specific scenes or text and provide feedback in real time.

[0866] The server receives user feedback and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user feedback.

[0867] Finally, the server provides the final video to the user, which is stored on the server and available for download or streaming from the user's dashboard.

[0868] Specific examples

[0869] Example 1: Corporate PR video production

[0870] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[0871] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[0872] Based on this information, generative AI creates a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending."

[0873] The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. The user reviews the preview video and provides feedback such as "Make the introduction a little shorter" or "Explain the benefits in more detail."

[0874] The server receives user feedback, and the generative AI generates the final video that reflects the feedback.

[0875] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[0876] Prompt Sentence Examples

[0877] "We would like to create a video to introduce a new product. The target users are businessmen in their 30s living in urban areas. The product's features are durability and beautiful design. Could you create a promotional video that emphasizes these points?"

[0878] In this way, this invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing generative AI, the video production process can be automated, significantly improving work efficiency.

[0879] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0880] Step 1:

[0881] A user submits a request for video production.

[0882] Specific operation: The user accesses a dedicated application or web portal and enters basic information such as the title, purpose, content summary, and materials to be used in the video production request form. When the user clicks the send button, the request is sent to the server.

[0883] Input: Information such as title "New product introduction", purpose "Customer promotion", and content summary "Highlight the features and benefits of the new product".

[0884] Output: The request data sent to the server.

[0885] Step 2:

[0886] The server receives and analyzes the user's request.

[0887] Specific operation: The server receives the request information from the user, analyzes its content, extracts the necessary relevant data based on the request content, and collects the necessary information from internal databases and external data sources.

[0888] Input: User request data.

[0889] Output: Relevant data to feed into the generative AI (e.g., product images, existing promotional video clips, text information).

[0890] Step 3:

[0891] Generative AI generates a video script and structure based on relevant data.

[0892] Specific operation: Based on the analysis results, the server provides prompts and related data to the generative AI. The generative AI then uses the provided information to create a video script and structure. The AI ​​uses natural language processing and image processing techniques to determine scene composition, text placement, and visual effects, and also generates text for the voice narration.

[0893] Input: relevant data, prompt statement.

[0894] Output: A video script and outline.

[0895] Step 4:

[0896] The server generates a preview video based on the generated configuration plan and provides it to the user.

[0897] Specific operation: The server generates a simple preview video based on the script and composition plan provided by the generative AI. The preview video includes scene stitching, text overlays, etc. The server uploads the generated preview video to the user's dashboard and notifies the user that it is ready.

[0898] Input: Video script and outline.

[0899] Output: The preview video that is provided to the user.

[0900] Step 5:

[0901] The user provides feedback on the preview video.

[0902] Specific operations: Users can check the preview video on the dashboard, input correction requests and feedback, provide correction instructions for specific scenes and text, and submit. Users can also submit feedback in real time.

[0903] Input: Preview video, user feedback.

[0904] Output: Feedback data sent to the server.

[0905] Step 6:

[0906] The server receives the feedback and reflects it in the generative AI.

[0907] Specific operation: The server receives feedback from the user and sends it to the generative AI. The generative AI reflects the feedback and revises the video script and composition plan. The server generates a new preview video based on the revised composition plan and provides it to the user again.

[0908] Input: User feedback, revised script and structure of the video.

[0909] Output: Revised preview video.

[0910] Step 7:

[0911] The server provides the final video to the user.

[0912] How it works: The server generates a high-quality final video based on the final script and plot plan. The generated video is stored on the server and uploaded to the user's dashboard. The user can then watch the final video by downloading or streaming.

[0913] Input: Final video script and outline.

[0914] Output: The final video that is delivered to the user.

[0915] (Application example 1)

[0916] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0917] Conventional video production systems have made it difficult for users without specialized knowledge and advanced skills to quickly produce high-quality videos. Furthermore, they lacked an interface for individual users to easily create and distribute personalized video content. Furthermore, they lacked a prompt input mechanism for generating video composition plans customized to the user's needs.

[0918] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0919] In this invention, the server includes: means for a user to send a video production request; means for providing related data to a generative artificial intelligence; means for generating a video script and a proposed composition based on the provided data; means for generating a simple preview video based on the proposed composition and providing it to the user; means for the user to provide feedback on the preview video; means for generating a final version of the video based on the feedback; means for providing the final version of the video to the user; means for the user to upload the generated video to a distribution service; and means for the user to customize the proposed composition by inputting a prompt. This makes it possible to quickly and easily create high-quality videos and easily distribute personalized video content without specialized knowledge or advanced technology.

[0920] "Means for users to submit video production requests" means the ability for users to use their own devices to input video production requests and details through a specific format or interface and send them to the server.

[0921] "Generative AI" refers to AI technology that automatically generates video scripts and plots based on provided data and information. It utilizes natural language processing and image processing technologies.

[0922] A "simple preview video" is a prototype video generated based on an initial video script and structure created by generative artificial intelligence, for users to check and provide feedback.

[0923] The "means for users to provide feedback on the preview video" refers to an interface or function that allows users to check the preview video, input requests for improvement or opinions about the content, and send the input to the server.

[0924] The "means of generating the final version of the video" is a function in which the generative artificial intelligence re-edits and re-structures the video based on feedback provided by the user, generating the final, completed version of the video.

[0925] The "means by which the server provides the final video to the user" refers to the function by which the server stores the generated final video and provides it for download or streaming through an interface accessible to the user.

[0926] "Means for uploading user-generated videos to a distribution service" refers to a function that allows users to easily upload and publish completed videos to a content distribution platform.

[0927] A "prompt sentence" is a text-based sentence that a user uses to input specific instructions or requests to a generative artificial intelligence when generating a script or plot plan for a video.

[0928] The present invention relates to a system that allows anyone to easily create high-quality videos and upload them to a content distribution service. Hereinafter, an embodiment of the invention will be described in detail.

[0929] The system mainly consists of a terminal where users send requests, a server that receives and processes video production requests, generative artificial intelligence (AI), and related databases and external data sources. The following describes a specific implementation of the process in which a user makes a video production request, AI generates a video script and composition plan based on that request, and ultimately provides a high-quality video.

[0930] 1. User submits request:

[0931] Users submit video production requests using a dedicated smartphone application or web portal, where they enter details such as the title of the video they want to create, its purpose, a summary of the content, and the materials they will use.

[0932] 2. Server reception and AI provision:

[0933] The request sent by the user is received by the server, which analyzes the information and provides the generative AI with relevant data, which is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[0934] 3. Video script and structure generation:

[0935] Generative AI generates a video script and structure proposal based on the provided data. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[0936] 4. Generate preview video:

[0937] The server creates a simple preview video based on the generated composition plan and provides it to the user. This preview video includes scene stitching, text overlays, etc., so that the user can check it.

[0938] 5. User feedback processing:

[0939] Users can view the preview video and provide feedback and suggestions for revisions, which are entered via the user's dashboard. This feedback can include specific instructions such as "make the introduction a little shorter" or "explain the benefits in more detail."

[0940] 6. Generate the final video:

[0941] The server receives user feedback and applies it to the generative AI. The AI ​​then retrains and generates a revised version of the video that reflects the user's feedback. Finally, the server provides the final version of the video to the user.

[0942] 7. Uploading and Publishing Videos:

[0943] Users can upload the generated video to a content distribution service, and can customize the video's structure by entering specific prompts.

[0944] Hardware and software used:

[0945] Hardware: Smartphones, server computers

[0946] Software: Python, Flask (for server processing), generative AI module, video editing module

[0947] Examples:

[0948] An example prompt for a user to create a promotional video for a new product:

[0949] Title: New Product Review

[0950] Purpose: To inform the audience

[0951] Summary: Highlight the features and benefits of your new product and compare it with competing products. Include specific usage scenarios if possible.

[0952] Using this prompt, the generative AI generates specific and effective video content.

[0953] As described above, this system makes it easy for anyone to create and distribute high-quality videos, automating the video production process and significantly improving work efficiency.

[0954] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0955] Step 1:

[0956] Users submit video production requests via a dedicated smartphone application or web portal.

[0957] Input: Information about the video title, purpose, summary, and materials used that the user enters into the application.

[0958] Specific operation: When the user enters the required information into the form and clicks the submit button, this data is sent from the application to the server.

[0959] Output: The user's request information arrives at the server.

[0960] Step 2:

[0961] The server receives the user's request, analyzes it, and provides the relevant data to the generative artificial intelligence.

[0962] Input: Request information sent by the user.

[0963] What it does: The server analyzes the request and gathers the necessary data from its internal database. If necessary, it retrieves additional data from external sources (e.g., news articles or social media posts) and provides it to the generative AI.

[0964] Output: Relevant data is provided to the generative artificial intelligence.

[0965] Step 3:

[0966] Generative AI generates a video script and structure based on the data provided.

[0967] Input: Relevant data provided to the generative artificial intelligence.

[0968] Specific operation: AI uses natural language processing and image processing technology to determine the video's scene composition, text placement, visual effects, etc., and also generates text for the voice narration.

[0969] Output: A script and outline of the generated video.

[0970] Step 4:

[0971] The server generates a simple preview video based on the generated configuration plan and provides it to the user.

[0972] Input: A video script and plot plan generated by generative artificial intelligence.

[0973] What happens: The server uses the video editing module to create a preview video, stitching scenes together and adding text overlays, uploading the generated preview video to the user's dashboard, and sending a notification that it's ready.

[0974] Output: A simple preview video provided to the user.

[0975] Step 5:

[0976] The user provides feedback on the preview video.

[0977] Input: The preview video the user watched and their feedback (e.g., a shorter introduction, a more detailed explanation of the benefits, etc.).

[0978] Specific action: A user enters comments or correction requests into the feedback form on the dashboard and clicks the submit button.

[0979] Output: User feedback reaches the server.

[0980] Step 6:

[0981] Based on the feedback, generative artificial intelligence generates the final version of the video.

[0982] Input: Feedback information from the user.

[0983] How it works: The server provides feedback to the generative AI, which then re-learns the content and generates a revised script and structure based on the feedback. The server then uses the video editing module to generate the final video.

[0984] Output: The final video generated.

[0985] Step 7:

[0986] The server provides the final video to the user.

[0987] Input: The final video.

[0988] What it does: Saves the final video to our servers and provides a link for users to download or stream it from their dashboard. Notifies users that their video is complete.

[0989] Output: The final video that is delivered to the user.

[0990] Step 8:

[0991] User-generated videos can be uploaded to content distribution services.

[0992] Input: Final video and streaming service account information.

[0993] What it does: Users click the "Upload" button in the application, follow a few simple steps, and the final video is uploaded to the distribution platform of their choice.

[0994] Output: The video published to a distribution platform.

[0995] Step 9:

[0996] Users can customize the video composition by entering a prompt.

[0997] Input: The specific prompt text that the user enters.

[0998] How it works: A user enters a prompt into a dashboard or application, which is then provided to the generative AI, which then generates content based on the prompt and customizes it to fit the user's needs.

[0999] Output: A video based on your customized composition.

[1000] By following the above steps, users can quickly create and distribute high-quality videos without having specialized knowledge.

[1001] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1002] The present invention relates to a system that enables users to easily and quickly create high-quality videos. This system involves a process in which the user submits a video production request, a server receives the request, and provides the data to a generative artificial intelligence (AI). The AI ​​then generates a script and a draft structure for the video, and the server provides a preview version to the user. Finally, the system generates and provides a final version of the video that incorporates user feedback. Furthermore, the present invention aims to achieve high satisfaction by incorporating an emotion engine that recognizes user emotions, thereby reflecting the user's emotions in the video production process.

[1003] Specifically, this includes the following processes:

[1004] Users submit requests for video production using a dedicated application or web portal. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[1005] The server analyzes the requests received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[1006] The generative AI generates a video script and structure proposal based on the information provided. The AI ​​uses natural language processing and image processing techniques to determine the video's scene composition, text placement, visual effects, etc. It also obtains additional information from external data sources as needed. Here, the emotion engine analyzes the user's reactions and emotions and provides this information to the generative AI, which then generates content that is more tailored to the user.

[1007] The server generates a simple preview video based on the generated script and plot plan and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[1008] When users watch the preview video, the emotion engine measures their reactions in real time. The emotions (e.g., satisfaction, dissatisfaction, excitement, etc.) felt by the user while watching the preview video are recorded. Users can provide feedback via the dashboard. Specifically, users can enter comments for each scene or text and specify additional corrections or improvements. This feedback also includes the emotional data measured by the emotion engine.

[1009] The server receives user feedback and emotional data and reflects it in the generative AI. The generative AI then retrains and generates a revised version of the video that reflects the user's feedback and emotions. The AI ​​then edits the video to emphasize the positive emotions felt by the user and reduce negative emotions.

[1010] Finally, the server uploads the final video to the user's dashboard and sends a notification to the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[1011] Specific examples

[1012] Example 1: Corporate PR video production

[1013] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[1014] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[1015] The generative AI uses this information to create a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending." The emotion engine analyzes the user's emotions and provides them to the generative AI, which then reflects the content according to the user's preferences.

[1016] The server creates a preview video based on the generated configuration and provides it to the user via a dashboard. The emotions (e.g., excitement, anticipation, doubt) that arise when the user watches the preview video are recorded.

[1017] Users provide feedback such as "Make the introduction shorter" or "Explain the benefits in more detail," and data from the sentiment engine is also sent.

[1018] The server receives user feedback, and the generative AI reflects the feedback and emotional data to generate the final video.

[1019] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[1020] In this way, the present invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing the emotion engine, more personalized content that reflects the user's emotions is generated, thereby improving user satisfaction.

[1021] The processing flow will be explained below.

[1022] Step 1:

[1023] Users access a dedicated application or web portal to submit a video production request. Specifically, users enter the required information, such as the title, purpose, summary of the content, and the materials they want to use, and then click the "Submit Request" button.

[1024] Step 2:

[1025] The server analyzes the request received from the user and collects relevant data based on the request information (e.g., materials from internal databases, additional information from external data sources, etc.).

[1026] Step 3:

[1027] The server provides the collected relevant data to the generative artificial intelligence (AI), including images, text, and existing video clips.

[1028] Step 4:

[1029] Generative AI generates a video script and structure proposal based on the provided data. Specifically, it uses natural language processing technology to generate text descriptions for each scene, image processing technology to determine the placement and effects of images, and obtains additional information from external data sources as needed.

[1030] Step 5:

[1031] The emotion engine measures users' reactions in real time. While users are submitting video creation requests, the emotion engine analyzes their facial expressions, voice, and other biometric information to generate emotion data.

[1032] Step 6:

[1033] The server generates a simple preview video based on the generated script and plot plan, which includes scene stitching, text overlays, and some visual effects.

[1034] Step 7:

[1035] The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready, after which the user can access the dashboard to view the preview video.

[1036] Step 8:

[1037] Users can view the preview video and provide feedback and correction requests via the dashboard. Specifically, users can enter comments for each scene and text, specifying additional corrections and improvements. This feedback also includes emotional data measured by the emotion engine.

[1038] Step 9:

[1039] The server receives feedback and emotion data from users and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user's feedback and emotion.

[1040] Step 10:

[1041] The server uploads the final video to the user's dashboard and notifies the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[1042] In this way, collaboration between users, servers, generative AI, and the emotion engine allows even non-experts to quickly and easily create high-quality videos. Furthermore, the inclusion of the emotion engine allows users' emotions to be reflected in the video production process, enabling the provision of more personalized content with a high level of satisfaction.

[1043] Example 2

[1044] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1045] In the past, users needed specialized knowledge and skills to quickly create high-quality videos. This meant that video production required a lot of time and money, making it difficult for average users to achieve this. It was also difficult to personalize the video content to match the user's emotions, making it difficult to increase user satisfaction.

[1046] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1047] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a draft structure based on the provided data; a means for the server to generate a simple preview video based on the generated draft structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for an emotion engine to analyze and acquire user emotion data; a means for the generative artificial intelligence to generate a revised version of the video based on the feedback and emotion data; a means for the server to provide the revised version of the video to the user; and a means for the server to generate a final version of the video and provide it to the user. This enables users to quickly create high-quality videos without specialized knowledge or skills, and further enables personalization of video content that reflects the user's emotions, thereby improving user satisfaction.

[1048] "User" refers to an individual or organization that uses the System to make a video production request.

[1049] "Server" refers to the computer system that receives and processes user requests, provides relevant data to the generative artificial intelligence, and generates videos and manages feedback.

[1050] "Generative AI" refers to AI technology that generates a video script and structure based on provided data, and then creates revised and final versions of the video.

[1051] An "emotion engine" refers to a technology that analyzes user reactions and emotions, acquires them as data, and provides this to generative artificial intelligence.

[1052] A "video production request" refers to the action of sending information to the server, including basic information about the video the user wants to produce (e.g., title, purpose, summary of content, materials to be used, etc.).

[1053] "Related Data" means data necessary to generate a video script and story (e.g., user-provided information, data from internal databases, data from external data sources).

[1054] A "script" refers to the story or script of a video that generative artificial intelligence creates based on the data provided.

[1055] "Composition proposal" refers to the scene composition and layout proposal for a video that is determined by generative artificial intelligence based on the data provided.

[1056] "Preview video" refers to a simple video generated by the server based on the script and plot plan created by generative AI, which is used by users to confirm and provide feedback.

[1057] "Feedback" refers to opinions and correction requests provided by users regarding preview videos.

[1058] "Revised video" refers to a video edited by generative artificial intelligence based on user feedback and emotional data.

[1059] "Final video" refers to the final completed video created by the generative artificial intelligence that the server provides to the user.

[1060] "Database" refers to a system that resides within a server and stores and manages related data.

[1061] "External data sources" refers to information sources that can be obtained from outside, such as news articles and social media posts.

[1062] This invention relates to a system that enables users to quickly create high-quality videos without specialized knowledge or skills. This system is composed of a server, generative artificial intelligence, an emotion engine, a dedicated application, and a web portal.

[1063] Hardware and software used

[1064] Hardware: Servers (computer systems with high-performance computing resources), terminals (user computers, smartphones, etc.)

[1065] Software: Dedicated applications, web portals, generative AI (e.g., OpenAI's GPT model, image generation model, etc.), emotion recognition engines (e.g., Affectiva, Microsoft Azure Emotion API, etc.)

[1066] server

[1067] The server receives video production requests submitted by users using a dedicated application or web portal. The request includes basic information such as the video title, purpose, summary, and materials used. The server analyzes this information and prepares relevant data to provide to the generative artificial intelligence. This relevant data includes information from the server's internal database and information from external data sources (such as news articles and social media posts).

[1068] Generative Artificial Intelligence

[1069] The generative AI generates a video script and structure proposal based on data provided by the server. It uses natural language processing and image processing techniques to determine the video's scene composition, text placement, visual effects, etc. It also uses data from the emotion engine to generate personalized content that reflects the user's emotions.

[1070] Emotion Engine

[1071] The emotion engine measures and analyzes the user's emotions in real time while they are watching the preview video. Emotional data such as satisfaction, dissatisfaction, and excitement felt by the user is recorded and provided to the generative AI. This allows content that reflects the user's emotions to be generated.

[1072] Dedicated application / web portal

[1073] Users can use a dedicated application or web portal to submit video production requests, view the generated preview videos and final videos, and provide specific feedback on the preview videos to the server through the application or web portal.

[1074] Example: Production of corporate PR videos

[1075] 1. A user requests the creation of a promotional video for a new product. At that time, the user enters information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and benefits of the new product."

[1076] 2. The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[1077] 3. The generative AI uses this information to create a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending." The emotion engine analyzes the user's emotional data and provides it to the generative AI, which then reflects the content according to the user's preferences.

[1078] 4. The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. As the user watches the video, the emotion engine records emotions such as excitement, anticipation, and doubt.

[1079] 5. The user provides feedback such as a short introduction and detailed benefit explanation, along with sentiment engine data.

[1080] 6. The server receives the user feedback, and the generative AI generates a revised version of the video based on the feedback and emotional data.

[1081] 7. Once the final video is completed, the server will provide it to the user, who can then download the high-quality final PR video from their dashboard and use it for promotional purposes.

[1082] Example prompts for generative AI models

[1083] "Write a script for a promotional video highlighting the features of a new product. The title should be 'New Product Introduction' and the purpose should be to promote it to customers. Use the following information: product images, text information, and existing promotional video clips."

[1084] "Edit your videos to reflect your users' emotional data and build excitement and anticipation."

[1085] In this way, the present invention allows users to quickly and easily create high-quality videos. In addition, by using an emotion engine, personalized video content that reflects the user's emotions is generated, improving user satisfaction.

[1086] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1087] Step 1:

[1088] The user submits a request for video production using a dedicated application or web portal. The user enters basic information such as the video title, purpose, and content summary, and clicks the submit button. The input data includes the video title "New product introduction," purpose "Customer promotion," and content summary "Highlighting the features and benefits of the new product." This is then sent to the server.

[1089] Step 2:

[1090] The server receives the request sent by the user and analyzes the content. The analyzed data includes the title, purpose, and content summary. Based on this data, the server prepares related data to provide to the generative artificial intelligence (AI). Specifically, it collects product images, existing promotional video clips, and text information from the server's internal database. This becomes the input data provided to the generative AI.

[1091] Step 3:

[1092] The server provides the prepared relevant data to the generative AI. The provided data includes product images, existing video clips, and text information. Based on this input data, the generative AI generates a video script and a proposed structure. Specific data processing involves using natural language processing technology to create the script, and image processing technology to determine scene composition, text placement, and visual effects. The output is a video script and a proposed structure.

[1093] Step 4:

[1094] The server generates a simple preview video based on the generated script and proposed structure. Generating the preview video involves stitching together scenes, overlaying text, and adding simple visual effects. The script and proposed structure output by the generative AI are used as input data. This results in a preview video that users can watch and confirm.

[1095] Step 5:

[1096] The server uploads the preview video to the user's dashboard, makes it available for viewing, and notifies the user when the preview video is ready. The input data includes the preview video, and this is the output data provided to the user via the dashboard.

[1097] Step 6:

[1098] Users watch preview videos on the dashboard. While watching, the emotion engine measures the user's emotional data in real time. Specifically, the emotion engine analyzes and records the user's satisfaction, dissatisfaction, excitement, etc. from facial expressions and voice. This becomes the input data sent to the server as emotional data.

[1099] Step 7:

[1100] Users provide specific feedback on the preview video to the server via the dashboard. This feedback includes specific instructions such as shortening the introduction or explaining the benefits in more detail. In addition, emotional data measured by the emotion engine is also sent to the server. This becomes the input data received by the server.

[1101] Step 8:

[1102] The server receives user feedback and emotional data and reflects it in the generative AI. The generative AI generates a revised version of the video based on this input data. Specific data calculations include changes to the script and composition plan based on the feedback, and scene editing using the emotional data. This results in the output of a revised version of the video.

[1103] Step 9:

[1104] The server uploads the generated modified video to the user's dashboard, where it is available for viewing and download, and notifies the user that the modified video is ready. The input data includes the modified video, which is the output data provided to the user.

[1105] Step 10:

[1106] The user can then re-watch the revised video from the dashboard for a final check. If necessary, they can provide further feedback. If there is further feedback, the server receives it and reflects it back into the generative AI. By repeating this process, the final video is completed.

[1107] Step 11:

[1108] The server generates the final video and uploads it to the user's dashboard. The server notifies the user when the final video is ready. The input data includes the final script, plot plan, and edited footage, and the output data is the final video.

[1109] Step 12:

[1110] Users will receive a notification and can download or stream the final video from their dashboard, allowing them to use the final, high-quality video for promotional or other uses.

[1111] Through the above processing steps, the system enables users to quickly create high-quality, personalized videos without having specialized knowledge or skills.

[1112] (Application example 2)

[1113] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1114] Conventional video production systems require non-expert users to have high technical knowledge and skills to quickly create high-quality videos, which requires a lot of time and effort. It is also difficult to create videos that reflect the user's emotions, making it difficult to generate content that satisfies the user. Furthermore, they are unable to quickly generate product introduction videos, especially in virtual stores, making it difficult to effectively promote products.

[1115] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1116] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a proposed composition based on the provided data; a means including an emotion engine that recognizes the user's emotions and reflects them in the video production process; and a means for generating and providing product introduction videos within a virtual store. This allows users to quickly create high-quality videos without non-specialized knowledge or skills, and the generation of personalized content that reflects emotions improves user satisfaction. Furthermore, the rapid generation of product introduction videos within a virtual store enables effective product advertising.

[1117] "Means for users to submit video production requests" refers to an interface through which users can input the information required to request video production and send it to the server. This may include a web portal, a smartphone app, or dedicated software.

[1118] "Means by which the server receives a user's request and provides relevant data to the generative AI" refers to the process by which the server receives a request sent by a user and appropriately conveys it to the generative AI, including querying a database or retrieving data from an external data source.

[1119] "Means for generative AI to generate a video script and composition plan based on provided data" refers to the process in which generative AI analyzes user request data provided by a server and, based on that, determines the video's scene composition, text placement, visual effects, etc.

[1120] "Means for the server to generate a simple preview video based on the generated composition plan and provide it to the user" refers to the process by which the server creates a visualized preview video based on the composition plan of the video created by the generative artificial intelligence and presents it to the user.

[1121] The "means for users to provide feedback on the preview video" refers to an interface that allows users to view the preview video and send their opinions or requests for corrections to the server. This includes a comment function and a rating system.

[1122] "Means for the generative AI to generate the final version of the video based on the feedback" refers to the process by which the generative AI receives feedback from users and creates the final version of the video that reflects that feedback. This process includes analyzing the user's emotional data using an emotion engine.

[1123] The "means by which the server provides the final video to the user" refers to the process by which the server generates and provides the final video file to the user through an interface such as a dashboard, which the user can then view by downloading or streaming.

[1124] The "Emotion Engine" is a system that measures the emotions of users when they watch videos in real time and collects the data. This makes it possible to create content that reflects the user's emotions during the video generation process.

[1125] The "means for generating and providing a product introduction video within a virtual store" refers to a means for generating an introduction video for a product selected by a user within a virtual store environment, visualizing it on the spot, and providing it to the user, thereby enabling the user to instantly understand the content of the product within the virtual store.

[1126] This invention relates to a system that allows users to quickly and easily create high-quality videos, and in particular, aims to facilitate the creation and provision of product introduction videos in a virtual store. This system has the following configuration and functions.

[1127] Users submit requests for video production using a dedicated application that runs on a smartphone or head-mounted display. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials used, etc.). This communicates the specifications of the video the user requires to the system.

[1128] The server receives requests sent by users, analyzes them, and provides relevant data to a generative artificial intelligence (AI). Based on the information provided by the server, the generative AI generates a video script and a proposed structure. The generative AI uses natural language processing and image processing technologies to determine scene composition, text placement, visual effects, and more. It also obtains additional information from external data sources (e.g., news articles and social media posts) as needed and incorporates it into the script and proposed structure.

[1129] As part of the generation process, an emotion engine analyzes the user's reactions and emotions and provides them to the generative AI, which then personalizes the video content based on the user's preferences. The emotion engine analyzes facial and voice data in real time as the user types their request.

[1130] Next, the server generates a simple preview video based on the generated script and plot plan and provides it to the user. The user can watch the preview video in the application, and the emotions they felt while watching it (e.g., excitement, dissatisfaction, anticipation, etc.) are recorded. The user provides feedback on the preview video and sends it to the server via the dashboard. This feedback also includes emotional data measured by the emotion engine.

[1131] The server receives user feedback and emotion data, and the generative AI retrains based on that data to generate the final video. The generative AI edits the video to emphasize the positive emotions felt by the user and reduce negative emotions. Finally, the server uploads the final video to the user's dashboard and sends a notification. After receiving the notification, the user can download or stream the final video from their dashboard.

[1132] A specific use case is creating a video to introduce a new drone product in a virtual store. Users can submit a request such as, "Please create a promotional video for our new drone product, highlighting its features and benefits. Please make it appealing and reflect the user's emotions." Based on this request, the generative AI creates the video, and the emotion engine is used to provide a video that reflects the user's feedback.

[1133] This system allows users to quickly create high-quality videos without specialized knowledge or skills, and generates personalized content that reflects emotions, improving user satisfaction.It also enables the rapid generation of product introduction videos in virtual stores, enabling effective product advertising.

[1134] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1135] Step 1:

[1136] The user submits a request for video production using a dedicated application.

[1137] Input: The user enters basic information into the application, such as the title, purpose, summary, and materials used.

[1138] Specific operation: The information entered by the user is sent to the server by clicking the "Submit" button.

[1139] Output: The input request information is transmitted to the server.

[1140] Step 2:

[1141] The server receives the user's request and provides the relevant data to the generative artificial intelligence.

[1142] Input: The request data sent by the user.

[1143] What happens: The server analyzes the request data and collects relevant data (e.g., images and text from a database, or additional information from external data sources).

[1144] Output: A relevant dataset that is fed into the generative artificial intelligence.

[1145] Step 3:

[1146] Generative AI generates a video script and structure based on the data provided.

[1147] Input: Relevant data provided by the server.

[1148] How it works: Generative AI uses natural language processing and image processing technologies to determine scene composition, text placement, and visual effects.

[1149] Output: Video script and outline.

[1150] Step 4:

[1151] The server generates a simple preview video based on the generated configuration plan and provides it to the user.

[1152] Input: A script and plot draft created by a generative artificial intelligence.

[1153] What it does: The server generates a preview video by stitching scenes, overlaying text, and combining basic visual effects, then uploads the preview video to the user's dashboard and sends a notification.

[1154] Output: The preview video created.

[1155] Step 5:

[1156] The user provides feedback on the preview video.

[1157] Input: User preview video views and reactions.

[1158] How it works: Users use the dashboard to enter feedback, submit comments and correction requests, and the emotion engine measures and captures their reactions (e.g., facial expressions and voice) as data.

[1159] Output: Feedback and emotion data.

[1160] Step 6:

[1161] Based on the feedback, the server uses generative artificial intelligence to generate the final version of the video.

[1162] Input: User-provided feedback and sentiment data.

[1163] How it works: The server passes this data to a generative AI, which then re-edits and re-generates the video, emphasizing positive emotions and reducing negative ones.

[1164] Output: The final video.

[1165] Step 7:

[1166] The server provides the final video to the user.

[1167] Input: The final video created by the generative artificial intelligence.

[1168] What happens: The server uploads the final video to the user's dashboard and sends a completion notification, which the user can then download or stream.

[1169] Output: Final video and completion notification.

[1170] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1171] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1172] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1173] [Fourth embodiment]

[1174] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1175] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1176] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1177] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1178] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1179] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1180] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1181] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1182] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1183] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1184] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1185] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1186] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1187] This invention relates to a system that enables anyone to easily and quickly create high-quality videos. This system involves a process in which a user submits a video production request, a server receives the request and provides data to a generative artificial intelligence (AI), the AI ​​generates a script and a draft structure for the video, the server provides a preview version to the user, and finally generates and provides the final version of the video that incorporates user feedback.

[1188] Specifically, this includes the following processes:

[1189] Users submit requests for video production using a dedicated application or web portal. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[1190] The server analyzes the requests received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[1191] Generative AI generates a video script and structure proposal based on the information provided. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[1192] The server generates a simple preview video based on the proposed composition by the generative AI and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[1193] Users can view the preview video and provide feedback and corrections as needed. This feedback is entered via the user's dashboard. Users can provide correction instructions for specific scenes or text.

[1194] The server receives user feedback and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user feedback.

[1195] Finally, the server provides the final video to the user, which is stored on the server and available for download or streaming from the user's dashboard.

[1196] Specific examples

[1197] Example 1: Corporate PR video production

[1198] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[1199] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[1200] Based on this information, generative AI creates a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending."

[1201] The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. The user reviews the preview video and provides feedback such as "Make the introduction a little shorter" or "Explain the benefits in more detail."

[1202] The server receives user feedback, and the generative AI generates the final video that reflects the feedback.

[1203] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[1204] In this way, this invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing generative AI, the video production process can be automated, significantly improving work efficiency.

[1205] The processing flow will be explained below.

[1206] Step 1:

[1207] Users access a dedicated application or web portal to submit a video production request. Specifically, users enter the required information, such as the title, purpose, summary of the content, and the materials they want to use, and then click the "Submit Request" button.

[1208] Step 2:

[1209] The server analyzes the request received from the user and collects relevant data based on the request information (e.g., materials from internal databases, additional information from external data sources, etc.).

[1210] Step 3:

[1211] The server provides the collected relevant data to the generative artificial intelligence (AI), including images, text, and existing video clips.

[1212] Step 4:

[1213] Generative AI generates a video script and structure proposal based on the provided data. Specifically, it uses natural language processing technology to generate text descriptions for each scene, image processing technology to determine the placement and effects of images, and obtains additional information from external data sources as needed.

[1214] Step 5:

[1215] The server generates a simple preview video based on the generated script and plot plan, which includes scene stitching, text overlays, and some visual effects.

[1216] Step 6:

[1217] The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready, after which the user can access the dashboard to view the preview video.

[1218] Step 7:

[1219] Users can check the preview video and provide feedback and correction requests through the dashboard. Specifically, users can enter comments for each scene and text, and specify additional corrections and improvements.

[1220] Step 8:

[1221] The server receives feedback from users and reflects it in the generative AI. The server analyzes the feedback and notifies the generative AI of any necessary changes.

[1222] Step 9:

[1223] The generative AI incorporates the feedback it receives and generates the final video. The AI ​​then learns from the data again and makes adjustments based on the user's requests.

[1224] Step 10:

[1225] The server uploads the final video to the user's dashboard and notifies the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[1226] In this way, collaboration between users, servers, and generative AI streamlines the video production process, enabling anyone to quickly and easily create high-quality videos.

[1227] Example 1

[1228] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1229] Traditionally, video production required specialized knowledge and skills, making it difficult for non-expert users to create high-quality videos in a short amount of time. Furthermore, the feedback process was inefficient, making it difficult to incorporate user feedback in real time. This resulted in increased costs and time for video production, and reduced operational efficiency.

[1230] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1231] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a draft structure based on the provided data; a means for the server to generate a simple preview video based on the generated draft structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for the generative artificial intelligence to generate a final version of the video based on the feedback; a means for the server to provide the final version of the video to the user; a means for the user to input and send feedback in real time; and a means for the server to receive the user's real-time feedback and reflect it in the generative artificial intelligence. This enables even non-expert users to produce high-quality videos in a short amount of time, and makes the video production process more efficient by reflecting feedback in real time.

[1232] "User" refers to the person who submits a video production request and provides feedback.

[1233] "Request" refers to information regarding requests and specifications submitted by a User for video production.

[1234] "Server" refers to a device or system that receives user requests, provides relevant data to the generative artificial intelligence, and manages the generated videos and feedback.

[1235] "Generative AI" refers to an AI technology that generates video scripts and plot plans based on provided data.

[1236] "Associated data" refers to information necessary for video production, such as text, images, video clips, audio narration, etc.

[1237] A "script" refers to a document or text that specifically outlines the scene structure and content of a video.

[1238] A "synopsis plan" refers to a plan that shows the overall structure of the video and details of each scene.

[1239] A "preview video" refers to a simple video created based on the generated composition plan.

[1240] "Feedback" refers to requests for corrections or opinions provided by users regarding preview videos.

[1241] "Final video" refers to a high-quality video that has been completed with user feedback reflected.

[1242] "Real-time feedback" refers to feedback sent by a user immediately on the spot.

[1243] The present invention relates to a system that allows users to easily and quickly create high-quality videos. This system involves a process in which a user submits a video production request, a server receives the request and provides data to a generative artificial intelligence (AI), the AI ​​generates a script and a draft structure for the video, the server provides a preview version to the user, and finally generates and provides the final version of the video that incorporates user feedback.

[1244] Specifically, it is configured as follows:

[1245] First, a user submits a request for video production using a dedicated application or web portal. In this request, the user enters basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[1246] The server then analyzes the request received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles, social media posts, etc.). The generative AI used by the server includes, for example, OpenAI's GPT-4 and AI with image processing technology.

[1247] Generative AI generates a video script and structure proposal based on the information provided. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[1248] The server then generates a simple preview video based on the proposed composition by the generative AI and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server then uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[1249] Users can preview the video and provide feedback or suggestions for corrections as needed. This feedback is entered via the user's dashboard. Users can also provide correction suggestions for specific scenes or text and provide feedback in real time.

[1250] The server receives user feedback and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user feedback.

[1251] Finally, the server provides the final video to the user, which is stored on the server and available for download or streaming from the user's dashboard.

[1252] Specific examples

[1253] Example 1: Corporate PR video production

[1254] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[1255] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[1256] Based on this information, generative AI creates a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending."

[1257] The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. The user reviews the preview video and provides feedback such as "Make the introduction a little shorter" or "Explain the benefits in more detail."

[1258] The server receives user feedback, and the generative AI generates the final video that reflects the feedback.

[1259] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[1260] Prompt Sentence Examples

[1261] "We would like to create a video to introduce a new product. The target users are businessmen in their 30s living in urban areas. The product's features are durability and beautiful design. Could you create a promotional video that emphasizes these points?"

[1262] In this way, this invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing generative AI, the video production process can be automated, significantly improving work efficiency.

[1263] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1264] Step 1:

[1265] A user submits a request for video production.

[1266] Specific operation: The user accesses a dedicated application or web portal and enters basic information such as the title, purpose, content summary, and materials to be used in the video production request form. When the user clicks the send button, the request is sent to the server.

[1267] Input: Information such as title "New product introduction", purpose "Customer promotion", and content summary "Highlight the features and benefits of the new product".

[1268] Output: The request data sent to the server.

[1269] Step 2:

[1270] The server receives and analyzes the user's request.

[1271] Specific operation: The server receives the request information from the user, analyzes its content, extracts the necessary relevant data based on the request content, and collects the necessary information from internal databases and external data sources.

[1272] Input: User request data.

[1273] Output: Relevant data to feed into the generative AI (e.g., product images, existing promotional video clips, text information).

[1274] Step 3:

[1275] Generative AI generates a video script and structure based on relevant data.

[1276] Specific operation: Based on the analysis results, the server provides prompts and related data to the generative AI. The generative AI then uses the provided information to create a video script and structure. The AI ​​uses natural language processing and image processing techniques to determine scene composition, text placement, and visual effects, and also generates text for the voice narration.

[1277] Input: relevant data, prompt statement.

[1278] Output: A video script and outline.

[1279] Step 4:

[1280] The server generates a preview video based on the generated configuration plan and provides it to the user.

[1281] Specific operation: The server generates a simple preview video based on the script and composition plan provided by the generative AI. The preview video includes scene stitching, text overlays, etc. The server uploads the generated preview video to the user's dashboard and notifies the user that it is ready.

[1282] Input: Video script and outline.

[1283] Output: The preview video that is provided to the user.

[1284] Step 5:

[1285] The user provides feedback on the preview video.

[1286] Specific operations: Users can check the preview video on the dashboard, input correction requests and feedback, provide correction instructions for specific scenes and text, and submit. Users can also submit feedback in real time.

[1287] Input: Preview video, user feedback.

[1288] Output: Feedback data sent to the server.

[1289] Step 6:

[1290] The server receives the feedback and reflects it in the generative AI.

[1291] Specific operation: The server receives feedback from the user and sends it to the generative AI. The generative AI reflects the feedback and revises the video script and composition plan. The server generates a new preview video based on the revised composition plan and provides it to the user again.

[1292] Input: User feedback, revised script and structure of the video.

[1293] Output: Revised preview video.

[1294] Step 7:

[1295] The server provides the final video to the user.

[1296] How it works: The server generates a high-quality final video based on the final script and plot plan. The generated video is stored on the server and uploaded to the user's dashboard. The user can then watch the final video by downloading or streaming.

[1297] Input: Final video script and outline.

[1298] Output: The final video that is delivered to the user.

[1299] (Application example 1)

[1300] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1301] Conventional video production systems have made it difficult for users without specialized knowledge and advanced skills to quickly produce high-quality videos. Furthermore, they lacked an interface for individual users to easily create and distribute personalized video content. Furthermore, they lacked a prompt input mechanism for generating video composition plans customized to the user's needs.

[1302] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1303] In this invention, the server includes: means for a user to send a video production request; means for providing related data to a generative artificial intelligence; means for generating a video script and a proposed composition based on the provided data; means for generating a simple preview video based on the proposed composition and providing it to the user; means for the user to provide feedback on the preview video; means for generating a final version of the video based on the feedback; means for providing the final version of the video to the user; means for the user to upload the generated video to a distribution service; and means for the user to customize the proposed composition by inputting a prompt. This makes it possible to quickly and easily create high-quality videos and easily distribute personalized video content without specialized knowledge or advanced technology.

[1304] "Means for users to submit video production requests" means the ability for users to use their own devices to input video production requests and details through a specific format or interface and send them to the server.

[1305] "Generative AI" refers to AI technology that automatically generates video scripts and plots based on provided data and information. It utilizes natural language processing and image processing technologies.

[1306] A "simple preview video" is a prototype video generated based on an initial video script and structure created by generative artificial intelligence, for users to check and provide feedback.

[1307] The "means for users to provide feedback on the preview video" refers to an interface or function that allows users to check the preview video, input requests for improvement or opinions about the content, and send the input to the server.

[1308] The "means of generating the final version of the video" is a function in which the generative artificial intelligence re-edits and re-structures the video based on feedback provided by the user, generating the final, completed version of the video.

[1309] The "means by which the server provides the final video to the user" refers to the function by which the server stores the generated final video and provides it for download or streaming through an interface accessible to the user.

[1310] "Means for uploading user-generated videos to a distribution service" refers to a function that allows users to easily upload and publish completed videos to a content distribution platform.

[1311] A "prompt sentence" is a text-based sentence that a user uses to input specific instructions or requests to a generative artificial intelligence when generating a script or plot plan for a video.

[1312] The present invention relates to a system that allows anyone to easily create high-quality videos and upload them to a content distribution service. Hereinafter, an embodiment of the invention will be described in detail.

[1313] The system mainly consists of a terminal where users send requests, a server that receives and processes video production requests, generative artificial intelligence (AI), and related databases and external data sources. The following describes a specific implementation of the process in which a user makes a video production request, AI generates a video script and composition plan based on that request, and ultimately provides a high-quality video.

[1314] 1. User submits request:

[1315] Users submit video production requests using a dedicated smartphone application or web portal, where they enter details such as the title of the video they want to create, its purpose, a summary of the content, and the materials they will use.

[1316] 2. Server reception and AI provision:

[1317] The request sent by the user is received by the server, which analyzes the information and provides the generative AI with relevant data, which is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[1318] 3. Video script and structure generation:

[1319] Generative AI generates a video script and structure proposal based on the provided data. Using natural language processing and image processing techniques, the AI ​​determines the video's scene structure, text placement, visual effects, etc. It also generates text for the voice narration.

[1320] 4. Generate preview video:

[1321] The server creates a simple preview video based on the generated composition plan and provides it to the user. This preview video includes scene stitching, text overlays, etc., so that the user can check it.

[1322] 5. User feedback processing:

[1323] Users can view the preview video and provide feedback and suggestions for revisions, which are entered via the user's dashboard. This feedback can include specific instructions such as "make the introduction a little shorter" or "explain the benefits in more detail."

[1324] 6. Generate the final video:

[1325] The server receives user feedback and applies it to the generative AI. The AI ​​then retrains and generates a revised version of the video that reflects the user's feedback. Finally, the server provides the final version of the video to the user.

[1326] 7. Uploading and Publishing Videos:

[1327] Users can upload the generated video to a content distribution service, and can customize the video's structure by entering specific prompts.

[1328] Hardware and software used:

[1329] Hardware: Smartphones, server computers

[1330] Software: Python, Flask (for server processing), generative AI module, video editing module

[1331] Examples:

[1332] An example prompt for a user to create a promotional video for a new product:

[1333] Title: New Product Review

[1334] Purpose: To inform the audience

[1335] Summary: Highlight the features and benefits of your new product and compare it with competing products. Include specific usage scenarios if possible.

[1336] Using this prompt, the generative AI generates specific and effective video content.

[1337] As described above, this system makes it easy for anyone to create and distribute high-quality videos, automating the video production process and significantly improving work efficiency.

[1338] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1339] Step 1:

[1340] Users submit video production requests via a dedicated smartphone application or web portal.

[1341] Input: Information about the video title, purpose, summary, and materials used that the user enters into the application.

[1342] Specific operation: When the user enters the required information into the form and clicks the submit button, this data is sent from the application to the server.

[1343] Output: The user's request information arrives at the server.

[1344] Step 2:

[1345] The server receives the user's request, analyzes it, and provides the relevant data to the generative artificial intelligence.

[1346] Input: Request information sent by the user.

[1347] What it does: The server analyzes the request and gathers the necessary data from its internal database. If necessary, it retrieves additional data from external sources (e.g., news articles or social media posts) and provides it to the generative AI.

[1348] Output: Relevant data is provided to the generative artificial intelligence.

[1349] Step 3:

[1350] Generative AI generates a video script and structure based on the data provided.

[1351] Input: Relevant data provided to the generative artificial intelligence.

[1352] Specific operation: AI uses natural language processing and image processing technology to determine the video's scene composition, text placement, visual effects, etc., and also generates text for the voice narration.

[1353] Output: A script and outline of the generated video.

[1354] Step 4:

[1355] The server generates a simple preview video based on the generated configuration plan and provides it to the user.

[1356] Input: A video script and plot plan generated by generative artificial intelligence.

[1357] What happens: The server uses the video editing module to create a preview video, stitching scenes together and adding text overlays, uploading the generated preview video to the user's dashboard, and sending a notification that it's ready.

[1358] Output: A simple preview video provided to the user.

[1359] Step 5:

[1360] The user provides feedback on the preview video.

[1361] Input: The preview video the user watched and their feedback (e.g., a shorter introduction, a more detailed explanation of the benefits, etc.).

[1362] Specific action: A user enters comments or correction requests into the feedback form on the dashboard and clicks the submit button.

[1363] Output: User feedback reaches the server.

[1364] Step 6:

[1365] Based on the feedback, generative artificial intelligence generates the final version of the video.

[1366] Input: Feedback information from the user.

[1367] How it works: The server provides feedback to the generative AI, which then re-learns the content and generates a revised script and structure based on the feedback. The server then uses the video editing module to generate the final video.

[1368] Output: The final video generated.

[1369] Step 7:

[1370] The server provides the final video to the user.

[1371] Input: The final video.

[1372] What it does: Saves the final video to our servers and provides a link for users to download or stream it from their dashboard. Notifies users that their video is complete.

[1373] Output: The final video that is delivered to the user.

[1374] Step 8:

[1375] User-generated videos can be uploaded to content distribution services.

[1376] Input: Final video and streaming service account information.

[1377] What it does: Users click the "Upload" button in the application, follow a few simple steps, and the final video is uploaded to the distribution platform of their choice.

[1378] Output: The video published to a distribution platform.

[1379] Step 9:

[1380] Users can customize the video composition by entering a prompt.

[1381] Input: The specific prompt text that the user enters.

[1382] How it works: A user enters a prompt into a dashboard or application, which is then provided to the generative AI, which then generates content based on the prompt and customizes it to fit the user's needs.

[1383] Output: A video based on your customized composition.

[1384] By following the above steps, users can quickly create and distribute high-quality videos without having specialized knowledge.

[1385] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1386] The present invention relates to a system that enables users to easily and quickly create high-quality videos. This system involves a process in which the user submits a video production request, a server receives the request, and provides the data to a generative artificial intelligence (AI). The AI ​​then generates a script and a draft structure for the video, and the server provides a preview version to the user. Finally, the system generates and provides a final version of the video that incorporates user feedback. Furthermore, the present invention aims to achieve high satisfaction by incorporating an emotion engine that recognizes user emotions, thereby reflecting the user's emotions in the video production process.

[1387] Specifically, this includes the following processes:

[1388] Users submit requests for video production using a dedicated application or web portal. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials to be used, etc.). When the user enters the information and clicks the submit button, the information is sent to the server.

[1389] The server analyzes the requests received from the user and provides the relevant data required for the generative AI. This relevant data is collected from the server's internal database and, if necessary, from external data sources (e.g., news articles or social media posts).

[1390] The generative AI generates a video script and structure proposal based on the information provided. The AI ​​uses natural language processing and image processing techniques to determine the video's scene composition, text placement, visual effects, etc. It also obtains additional information from external data sources as needed. Here, the emotion engine analyzes the user's reactions and emotions and provides this information to the generative AI, which then generates content that is more tailored to the user.

[1391] The server generates a simple preview video based on the generated script and plot plan and provides it to the user. This preview video includes scene stitching, text overlays, etc. The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready.

[1392] When users watch the preview video, the emotion engine measures their reactions in real time. The emotions (e.g., satisfaction, dissatisfaction, excitement, etc.) felt by the user while watching the preview video are recorded. Users can provide feedback via the dashboard. Specifically, users can enter comments for each scene or text and specify additional corrections or improvements. This feedback also includes the emotional data measured by the emotion engine.

[1393] The server receives user feedback and emotional data and reflects it in the generative AI. The generative AI then retrains and generates a revised version of the video that reflects the user's feedback and emotions. The AI ​​then edits the video to emphasize the positive emotions felt by the user and reduce negative emotions.

[1394] Finally, the server uploads the final video to the user's dashboard and sends a notification to the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[1395] Specific examples

[1396] Example 1: Corporate PR video production

[1397] A user requests the server to create a promotional video for a new product, inputting information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and advantages of the new product."

[1398] The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[1399] The generative AI uses this information to create a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending." The emotion engine analyzes the user's emotions and provides them to the generative AI, which then reflects the content according to the user's preferences.

[1400] The server creates a preview video based on the generated configuration and provides it to the user via a dashboard. The emotions (e.g., excitement, anticipation, doubt) that arise when the user watches the preview video are recorded.

[1401] Users provide feedback such as "Make the introduction shorter" or "Explain the benefits in more detail," and data from the sentiment engine is also sent.

[1402] The server receives user feedback, and the generative AI reflects the feedback and emotional data to generate the final video.

[1403] After the final version of the video is completed, the server will provide it to the user, who can then download the final high-quality PR video from their dashboard and use it for promotional purposes.

[1404] In this way, the present invention enables users to quickly and easily create high-quality videos without specialized knowledge or skills. Furthermore, by utilizing the emotion engine, more personalized content that reflects the user's emotions is generated, thereby improving user satisfaction.

[1405] The processing flow will be explained below.

[1406] Step 1:

[1407] Users access a dedicated application or web portal to submit a video production request. Specifically, users enter the required information, such as the title, purpose, summary of the content, and the materials they want to use, and then click the "Submit Request" button.

[1408] Step 2:

[1409] The server analyzes the request received from the user and collects relevant data based on the request information (e.g., materials from internal databases, additional information from external data sources, etc.).

[1410] Step 3:

[1411] The server provides the collected relevant data to the generative artificial intelligence (AI), including images, text, and existing video clips.

[1412] Step 4:

[1413] Generative AI generates a video script and structure proposal based on the provided data. Specifically, it uses natural language processing technology to generate text descriptions for each scene, image processing technology to determine the placement and effects of images, and obtains additional information from external data sources as needed.

[1414] Step 5:

[1415] The emotion engine measures users' reactions in real time. While users are submitting video creation requests, the emotion engine analyzes their facial expressions, voice, and other biometric information to generate emotion data.

[1416] Step 6:

[1417] The server generates a simple preview video based on the generated script and plot plan, which includes scene stitching, text overlays, and some visual effects.

[1418] Step 7:

[1419] The server uploads the preview video to the user's dashboard and notifies the user that the preview is ready, after which the user can access the dashboard to view the preview video.

[1420] Step 8:

[1421] Users can view the preview video and provide feedback and correction requests via the dashboard. Specifically, users can enter comments for each scene and text, specifying additional corrections and improvements. This feedback also includes emotional data measured by the emotion engine.

[1422] Step 9:

[1423] The server receives feedback and emotion data from users and reflects it in the generative AI, which then retrains and generates a revised version of the video that reflects the user's feedback and emotion.

[1424] Step 10:

[1425] The server uploads the final video to the user's dashboard and notifies the user that the video is complete, allowing the user to download or stream the final video from their dashboard.

[1426] In this way, collaboration between users, servers, generative AI, and the emotion engine allows even non-experts to quickly and easily create high-quality videos. Furthermore, the inclusion of the emotion engine allows users' emotions to be reflected in the video production process, enabling the provision of more personalized content with a high level of satisfaction.

[1427] Example 2

[1428] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1429] In the past, users needed specialized knowledge and skills to quickly create high-quality videos. This meant that video production required a lot of time and money, making it difficult for average users to achieve this. It was also difficult to personalize the video content to match the user's emotions, making it difficult to increase user satisfaction.

[1430] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1431] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a draft structure based on the provided data; a means for the server to generate a simple preview video based on the generated draft structure and provide it to the user; a means for the user to provide feedback on the preview video; a means for an emotion engine to analyze and acquire user emotion data; a means for the generative artificial intelligence to generate a revised version of the video based on the feedback and emotion data; a means for the server to provide the revised version of the video to the user; and a means for the server to generate a final version of the video and provide it to the user. This enables users to quickly create high-quality videos without specialized knowledge or skills, and further enables personalization of video content that reflects the user's emotions, thereby improving user satisfaction.

[1432] "User" refers to an individual or organization that uses the System to make a video production request.

[1433] "Server" refers to the computer system that receives and processes user requests, provides relevant data to the generative artificial intelligence, and generates videos and manages feedback.

[1434] "Generative AI" refers to AI technology that generates a video script and structure based on provided data, and then creates revised and final versions of the video.

[1435] An "emotion engine" refers to a technology that analyzes user reactions and emotions, acquires them as data, and provides this to generative artificial intelligence.

[1436] A "video production request" refers to the action of sending information to the server, including basic information about the video the user wants to produce (e.g., title, purpose, summary of content, materials to be used, etc.).

[1437] "Related Data" means data necessary to generate a video script and story (e.g., user-provided information, data from internal databases, data from external data sources).

[1438] A "script" refers to the story or script of a video that generative artificial intelligence creates based on the data provided.

[1439] "Composition proposal" refers to the scene composition and layout proposal for a video that is determined by generative artificial intelligence based on the data provided.

[1440] "Preview video" refers to a simple video generated by the server based on the script and plot plan created by generative AI, which is used by users to confirm and provide feedback.

[1441] "Feedback" refers to opinions and correction requests provided by users regarding preview videos.

[1442] "Revised video" refers to a video edited by generative artificial intelligence based on user feedback and emotional data.

[1443] "Final video" refers to the final completed video created by the generative artificial intelligence that the server provides to the user.

[1444] "Database" refers to a system that resides within a server and stores and manages related data.

[1445] "External data sources" refers to information sources that can be obtained from outside, such as news articles and social media posts.

[1446] This invention relates to a system that enables users to quickly create high-quality videos without specialized knowledge or skills. This system is composed of a server, generative artificial intelligence, an emotion engine, a dedicated application, and a web portal.

[1447] Hardware and software used

[1448] Hardware: Servers (computer systems with high-performance computing resources), terminals (user computers, smartphones, etc.)

[1449] Software: Dedicated applications, web portals, generative AI (e.g., OpenAI's GPT model, image generation model, etc.), emotion recognition engines (e.g., Affectiva, Microsoft Azure Emotion API, etc.)

[1450] server

[1451] The server receives video production requests submitted by users using a dedicated application or web portal. The request includes basic information such as the video title, purpose, summary, and materials used. The server analyzes this information and prepares relevant data to provide to the generative artificial intelligence. This relevant data includes information from the server's internal database and information from external data sources (such as news articles and social media posts).

[1452] Generative Artificial Intelligence

[1453] The generative AI generates a video script and structure proposal based on data provided by the server. It uses natural language processing and image processing techniques to determine the video's scene composition, text placement, visual effects, etc. It also uses data from the emotion engine to generate personalized content that reflects the user's emotions.

[1454] Emotion Engine

[1455] The emotion engine measures and analyzes the user's emotions in real time while they are watching the preview video. Emotional data such as satisfaction, dissatisfaction, and excitement felt by the user is recorded and provided to the generative AI. This allows content that reflects the user's emotions to be generated.

[1456] Dedicated application / web portal

[1457] Users can use a dedicated application or web portal to submit video production requests, view the generated preview videos and final videos, and provide specific feedback on the preview videos to the server through the application or web portal.

[1458] Example: Production of corporate PR videos

[1459] 1. A user requests the creation of a promotional video for a new product. At that time, the user enters information such as the title "New product introduction," the purpose "Customer promotion," and the content summary "Highlighting the features and benefits of the new product."

[1460] 2. The server receives the request and provides the generative AI with relevant data (product images, existing promotional video clips, text information, etc.).

[1461] 3. The generative AI uses this information to create a video script and determines the composition of scenes such as "introduction," "product introduction," "emphasis on benefits," and "ending." The emotion engine analyzes the user's emotional data and provides it to the generative AI, which then reflects the content according to the user's preferences.

[1462] 4. The server creates a preview video based on the generated configuration and provides it to the user via the dashboard. As the user watches the video, the emotion engine records emotions such as excitement, anticipation, and doubt.

[1463] 5. The user provides feedback such as a short introduction and detailed benefit explanation, along with sentiment engine data.

[1464] 6. The server receives the user feedback, and the generative AI generates a revised version of the video based on the feedback and emotional data.

[1465] 7. Once the final video is completed, the server will provide it to the user, who can then download the high-quality final PR video from their dashboard and use it for promotional purposes.

[1466] Example prompts for generative AI models

[1467] "Write a script for a promotional video highlighting the features of a new product. The title should be 'New Product Introduction' and the purpose should be to promote it to customers. Use the following information: product images, text information, and existing promotional video clips."

[1468] "Edit your videos to reflect your users' emotional data and build excitement and anticipation."

[1469] In this way, the present invention allows users to quickly and easily create high-quality videos. In addition, by using an emotion engine, personalized video content that reflects the user's emotions is generated, improving user satisfaction.

[1470] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1471] Step 1:

[1472] The user submits a request for video production using a dedicated application or web portal. The user enters basic information such as the video title, purpose, and content summary, and clicks the submit button. The input data includes the video title "New product introduction," purpose "Customer promotion," and content summary "Highlighting the features and benefits of the new product." This is then sent to the server.

[1473] Step 2:

[1474] The server receives the request sent by the user and analyzes the content. The analyzed data includes the title, purpose, and content summary. Based on this data, the server prepares related data to provide to the generative artificial intelligence (AI). Specifically, it collects product images, existing promotional video clips, and text information from the server's internal database. This becomes the input data provided to the generative AI.

[1475] Step 3:

[1476] The server provides the prepared relevant data to the generative AI. The provided data includes product images, existing video clips, and text information. Based on this input data, the generative AI generates a video script and a proposed structure. Specific data processing involves using natural language processing technology to create the script, and image processing technology to determine scene composition, text placement, and visual effects. The output is a video script and a proposed structure.

[1477] Step 4:

[1478] The server generates a simple preview video based on the generated script and proposed structure. Generating the preview video involves stitching together scenes, overlaying text, and adding simple visual effects. The script and proposed structure output by the generative AI are used as input data. This results in a preview video that users can watch and confirm.

[1479] Step 5:

[1480] The server uploads the preview video to the user's dashboard, makes it available for viewing, and notifies the user when the preview video is ready. The input data includes the preview video, and this is the output data provided to the user via the dashboard.

[1481] Step 6:

[1482] Users watch preview videos on the dashboard. While watching, the emotion engine measures the user's emotional data in real time. Specifically, the emotion engine analyzes and records the user's satisfaction, dissatisfaction, excitement, etc. from facial expressions and voice. This becomes the input data sent to the server as emotional data.

[1483] Step 7:

[1484] Users provide specific feedback on the preview video to the server via the dashboard. This feedback includes specific instructions such as shortening the introduction or explaining the benefits in more detail. In addition, emotional data measured by the emotion engine is also sent to the server. This becomes the input data received by the server.

[1485] Step 8:

[1486] The server receives user feedback and emotional data and reflects it in the generative AI. The generative AI generates a revised version of the video based on this input data. Specific data calculations include changes to the script and composition plan based on the feedback, and scene editing using the emotional data. This results in the output of a revised version of the video.

[1487] Step 9:

[1488] The server uploads the generated modified video to the user's dashboard, where it is available for viewing and download, and notifies the user that the modified video is ready. The input data includes the modified video, which is the output data provided to the user.

[1489] Step 10:

[1490] The user can then re-watch the revised video from the dashboard for a final check. If necessary, they can provide further feedback. If there is further feedback, the server receives it and reflects it back into the generative AI. By repeating this process, the final video is completed.

[1491] Step 11:

[1492] The server generates the final video and uploads it to the user's dashboard. The server notifies the user when the final video is ready. The input data includes the final script, plot plan, and edited footage, and the output data is the final video.

[1493] Step 12:

[1494] Users will receive a notification and can download or stream the final video from their dashboard, allowing them to use the final, high-quality video for promotional or other uses.

[1495] Through the above processing steps, the system enables users to quickly create high-quality, personalized videos without having specialized knowledge or skills.

[1496] (Application example 2)

[1497] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1498] Conventional video production systems require non-expert users to have high technical knowledge and skills to quickly create high-quality videos, which requires a lot of time and effort. It is also difficult to create videos that reflect the user's emotions, making it difficult to generate content that satisfies the user. Furthermore, they are unable to quickly generate product introduction videos, especially in virtual stores, making it difficult to effectively promote products.

[1499] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1500] In this invention, the server includes: a means for a user to send a video production request; a means for the server to receive the user's request and provide related data to the generative artificial intelligence; a means for the generative artificial intelligence to generate a video script and a proposed composition based on the provided data; a means including an emotion engine that recognizes the user's emotions and reflects them in the video production process; and a means for generating and providing product introduction videos within a virtual store. This allows users to quickly create high-quality videos without non-specialized knowledge or skills, and the generation of personalized content that reflects emotions improves user satisfaction. Furthermore, the rapid generation of product introduction videos within a virtual store enables effective product advertising.

[1501] "Means for users to submit video production requests" refers to an interface through which users can input the information required to request video production and send it to the server. This may include a web portal, a smartphone app, or dedicated software.

[1502] "Means by which the server receives a user's request and provides relevant data to the generative AI" refers to the process by which the server receives a request sent by a user and appropriately conveys it to the generative AI, including querying a database or retrieving data from an external data source.

[1503] "Means for generative AI to generate a video script and composition plan based on provided data" refers to the process in which generative AI analyzes user request data provided by a server and, based on that, determines the video's scene composition, text placement, visual effects, etc.

[1504] "Means for the server to generate a simple preview video based on the generated composition plan and provide it to the user" refers to the process by which the server creates a visualized preview video based on the composition plan of the video created by the generative artificial intelligence and presents it to the user.

[1505] The "means for users to provide feedback on the preview video" refers to an interface that allows users to view the preview video and send their opinions or requests for corrections to the server. This includes a comment function and a rating system.

[1506] "Means for the generative AI to generate the final version of the video based on the feedback" refers to the process by which the generative AI receives feedback from users and creates the final version of the video that reflects that feedback. This process includes analyzing the user's emotional data using an emotion engine.

[1507] The "means by which the server provides the final video to the user" refers to the process by which the server generates and provides the final video file to the user through an interface such as a dashboard, which the user can then view by downloading or streaming.

[1508] The "Emotion Engine" is a system that measures the emotions of users when they watch videos in real time and collects the data. This makes it possible to create content that reflects the user's emotions during the video generation process.

[1509] The "means for generating and providing a product introduction video within a virtual store" refers to a means for generating an introduction video for a product selected by a user within a virtual store environment, visualizing it on the spot, and providing it to the user, thereby enabling the user to instantly understand the content of the product within the virtual store.

[1510] This invention relates to a system that allows users to quickly and easily create high-quality videos, and in particular, aims to facilitate the creation and provision of product introduction videos in a virtual store. This system has the following configuration and functions.

[1511] Users submit requests for video production using a dedicated application that runs on a smartphone or head-mounted display. In the request, they enter basic information about the video they want to create (e.g., title, purpose, summary of content, materials used, etc.). This communicates the specifications of the video the user requires to the system.

[1512] The server receives requests sent by users, analyzes them, and provides relevant data to a generative artificial intelligence (AI). Based on the information provided by the server, the generative AI generates a video script and a proposed structure. The generative AI uses natural language processing and image processing technologies to determine scene composition, text placement, visual effects, and more. It also obtains additional information from external data sources (e.g., news articles and social media posts) as needed and incorporates it into the script and proposed structure.

[1513] As part of the generation process, an emotion engine analyzes the user's reactions and emotions and provides them to the generative AI, which then personalizes the video content based on the user's preferences. The emotion engine analyzes facial and voice data in real time as the user types their request.

[1514] Next, the server generates a simple preview video based on the generated script and plot plan and provides it to the user. The user can watch the preview video in the application, and the emotions they felt while watching it (e.g., excitement, dissatisfaction, anticipation, etc.) are recorded. The user provides feedback on the preview video and sends it to the server via the dashboard. This feedback also includes emotional data measured by the emotion engine.

[1515] The server receives user feedback and emotion data, and the generative AI retrains based on that data to generate the final video. The generative AI edits the video to emphasize the positive emotions felt by the user and reduce negative emotions. Finally, the server uploads the final video to the user's dashboard and sends a notification. After receiving the notification, the user can download or stream the final video from their dashboard.

[1516] A specific use case is creating a video to introduce a new drone product in a virtual store. Users can submit a request such as, "Please create a promotional video for our new drone product, highlighting its features and benefits. Please make it appealing and reflect the user's emotions." Based on this request, the generative AI creates the video, and the emotion engine is used to provide a video that reflects the user's feedback.

[1517] This system allows users to quickly create high-quality videos without specialized knowledge or skills, and generates personalized content that reflects emotions, improving user satisfaction.It also enables the rapid generation of product introduction videos in virtual stores, enabling effective product advertising.

[1518] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1519] Step 1:

[1520] The user submits a request for video production using a dedicated application.

[1521] Input: The user enters basic information into the application, such as the title, purpose, summary, and materials used.

[1522] Specific operation: The information entered by the user is sent to the server by clicking the "Submit" button.

[1523] Output: The input request information is transmitted to the server.

[1524] Step 2:

[1525] The server receives the user's request and provides the relevant data to the generative artificial intelligence.

[1526] Input: The request data sent by the user.

[1527] What happens: The server analyzes the request data and collects relevant data (e.g., images and text from a database, or additional information from external data sources).

[1528] Output: A relevant dataset that is fed into the generative artificial intelligence.

[1529] Step 3:

[1530] Generative AI generates a video script and structure based on the data provided.

[1531] Input: Relevant data provided by the server.

[1532] How it works: Generative AI uses natural language processing and image processing technologies to determine scene composition, text placement, and visual effects.

[1533] Output: Video script and outline.

[1534] Step 4:

[1535] The server generates a simple preview video based on the generated configuration plan and provides it to the user.

[1536] Input: A script and plot draft created by a generative artificial intelligence.

[1537] What it does: The server generates a preview video by stitching scenes, overlaying text, and combining basic visual effects, then uploads the preview video to the user's dashboard and sends a notification.

[1538] Output: The preview video created.

[1539] Step 5:

[1540] The user provides feedback on the preview video.

[1541] Input: User preview video views and reactions.

[1542] How it works: Users use the dashboard to enter feedback, submit comments and correction requests, and the emotion engine measures and captures their reactions (e.g., facial expressions and voice) as data.

[1543] Output: Feedback and emotion data.

[1544] Step 6:

[1545] Based on the feedback, the server uses generative artificial intelligence to generate the final version of the video.

[1546] Input: User-provided feedback and sentiment data.

[1547] How it works: The server passes this data to a generative AI, which then re-edits and re-generates the video, emphasizing positive emotions and reducing negative ones.

[1548] Output: The final video.

[1549] Step 7:

[1550] The server provides the final video to the user.

[1551] Input: The final video created by the generative artificial intelligence.

[1552] What happens: The server uploads the final video to the user's dashboard and sends a completion notification, which the user can then download or stream.

[1553] Output: Final video and completion notification.

[1554] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1555] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1556] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1557] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1558] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1559] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1560] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1561] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1562] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1563] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1564] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1565] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1566] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1567] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1568] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1569] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1570] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1571] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1572] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1573] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1574] All publications, patent applications, and technical standards mentioned in...

Claims

1. a means for users to submit video production requests; A means for the server to receive a user request and provide relevant data to the generative artificial intelligence; A means for a generative artificial intelligence to generate a video script and a video composition plan based on the provided data; A means for the server to generate a simple preview video based on the generated configuration plan and provide it to the user; a means for users to provide feedback on the preview video; A means for generative artificial intelligence to generate the final version of the video based on the feedback; and a means by which the server provides the final video to the user; A system including:

2. The system of claim 1 , further comprising means for acquiring and incorporating additional information from external data sources when the generative artificial intelligence generates a script and a proposed structure for the video.

3. The system of claim 1 , further comprising means for automatically integrating text, images, video clips, and audio narration in the generation of the preview video and in the generation of the final version.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A