system
The system addresses limitations in conventional video generation by integrating user materials with generative AI, using computational vision and natural language processing to produce high-quality, shareable videos.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
Conventional video creation and posting systems face challenges in generating creative and high-quality content due to limited user materials and inadequate quality control, leading to difficulties in user engagement and platform usage.
A system that integrates user-selected media materials and selection information with generative artificial intelligence to create videos, incorporating computational vision and natural language processing for content generation, while detecting and filtering inappropriate content.
Enables the quick and intuitive production of high-quality, creative videos that users can easily share, ensuring appropriateness and enhancing user engagement on platforms.
Smart Images

Figure 2026071693000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional video creation and posting systems, it is difficult to generate creative and free content because the materials and technologies available to users are limited. In addition, the generated content may contain inappropriate elements, and there are quality control problems. As a result, there is a problem that users cannot easily produce satisfactory videos, and it is difficult to increase the number of users of video posting platforms.
Means for Solving the Problems
[0005] To solve the above problems, the present invention provides a means for transmitting images or videos acquired by a communication terminal to a server as media material, and for receiving and transmitting story and character selection information from the user. The server then integrates the received media material and selection information and sends instructions to a generating artificial intelligence. Subsequently, the generating artificial intelligence generates a video based on the received information and detects inappropriate content. Finally, the server reviews the generated video, and the communication terminal presents the generated video to the user, thereby realizing a system that can easily generate highly creative content and provide it to the user in an appropriate manner.
[0006] A "communication terminal" is an electronic device used by a user to acquire or transmit data, and includes devices such as smartphones and tablets.
[0007] "Image or video" refers to static or dynamic visual data, and is a media format acquired by a digital device.
[0008] "Media materials" refer to the source image or video data used to create a video.
[0009] "Generative artificial intelligence" refers to an algorithm or program that processes information like a human and generates new content based on user specifications.
[0010] "Story" refers to the narrative or series of events used to form the structure and plot of a video.
[0011] A "character" refers to a person, animal, or fictional being that acts according to the scenario within the video, and is an element selected by the user.
[0012] A "server" refers to a computer system that functions to store and process data and to provide information to other devices.
[0013] "Inappropriate content" refers to offensive or prohibited elements contained in a video, and any information deemed harmful to viewers.
[0014] "Review" refers to the process of evaluating or re-examining generated content and determining its appropriateness for publication.
[0015] "Presentation" refers to the act of displaying the generated video to the user and allowing them to visually confirm it. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0020] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention is a system centered around a user's communication terminal and server, utilizing generation AI to generate and manage creative and free-form videos. The functions and procedures of each component are described in detail below.
[0038] First, the user uses a communication device such as a smartphone to select images or videos they have taken or saved. This communication device has the function of sending the selected images or videos to a server. In terms of configuration, an application on the device provides the user interface, allowing the user to intuitively select the material they want to turn into a video.
[0039] Next, the user selects a story and characters to serve as a template for video generation via an application on their device. These choices are necessary to form the video's structure, visual elements, and scenario, and various templates are provided in advance within the application. The user can combine these options to create their own unique settings according to their preferences.
[0040] The selected materials and settings are sent from the terminal to the server. The server receives and integrates this input data and instructs the generative AI to generate a new video. The generative AI uses computer vision and natural language processing techniques to automatically create a video depicting the specified story and characters. The system also has a function to automatically detect inappropriate content, ensuring that no elements that may offend the user are generated.
[0041] Once the video generation is complete, it undergoes a verification process on the server and is then notified to the user's communication device. Upon receiving the notification, the user can preview the generated video within the application. This video is presented to the user as comprehensively edited content, including visual information and music or sound effects. The user can then choose to share the video on other platforms such as social media, or publish it within the platform itself.
[0042] (Specific example)
[0043] Suppose a user takes a photo of a beautiful landscape while traveling. This user launches a video generation app and selects the landscape photo. Next, the user chooses the theme "Fantasy Adventure" as the story theme and selects a dragon as the character. Based on this information, a generation AI summoned by the server creates an animation of a dragon flying through the sky and over mountains, accompanied by a grand background music. After checking the quality and appropriateness of the video, the user watches the finished video, is satisfied, and decides to share it with friends on social media. In this way, the present invention provides a form for users to intuitively and easily generate high-quality content.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The user launches an application on the communication terminal and selects images or videos from the terminal to be used as material for video generation. After the user's selection, the communication terminal packets the selected media files and prepares to send them to the server.
[0047] Step 2:
[0048] The device sends the image or video data selected by the user to the server. This transmission usually occurs in real time via an internet connection.
[0049] Step 3:
[0050] Users select story templates and character models presented through the application's UI, thereby concretizing their concept for the type of video they want to create.
[0051] Step 4:
[0052] The device sends the user's setting selections, along with story and character information, to the server in a defined data format.
[0053] Step 5:
[0054] The server integrates image or video data received from the terminal along with story and character selection information. This integrated data forms the basis for generating new videos.
[0055] Step 6:
[0056] The server inputs this integrated data into the generating AI and instructs it to begin the video generation process.
[0057] Step 7:
[0058] The generative AI utilizes computer vision and natural language processing technologies to generate videos based on integrated data. Simultaneously, the AI checks for inappropriate content during the generation process.
[0059] Step 8:
[0060] The server receives the video sent from the generating AI and reviews its content for final confirmation. If necessary, it can send correction instructions back to the generating AI.
[0061] Step 9:
[0062] The server converts the verified video data into the appropriate format and prepares it for storage in a form accessible to the user.
[0063] Step 10:
[0064] The server notifies the user that the video has been generated and is ready for publication. This notification is sent to the device and reflected in the application.
[0065] Step 11:
[0066] When the device receives a notification from the server, it provides the user with a preview of the newly generated video. The user can then review this preview and choose whether to publish or share it.
[0067] (Example 1)
[0068] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0069] Conventional video generation systems have the challenge of not being able to quickly generate videos that meet users' creative and individual requests. Furthermore, insufficient filtering of inappropriate content can potentially cause offense to users. Therefore, there is a need for efficient and high-quality video generation.
[0070] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0071] In this invention, the server includes means for integrating received media material and selection information and transmitting instructions to a generating artificial intelligence; means for the generating artificial intelligence to generate a video based on the received information and to detect and filter inappropriate content; and means for performing a process to examine the video generated by the generating artificial intelligence. This makes it possible to quickly generate safe and high-quality videos for users and improve the usability of the product.
[0072] A "communication terminal" is a device that has the function of allowing users to acquire images and videos and transmit that information to a computing device.
[0073] A "computational device" is a device that processes received media materials and selection information and sends instructions to the generating artificial intelligence.
[0074] "Generative artificial intelligence" is an artificial intelligence technology that generates videos based on received information and has the ability to detect and filter inappropriate content.
[0075] "Media material" refers to image and video data acquired by users using communication devices.
[0076] "Story and characters" refers to the visual and contextual elements necessary for video creation, as selected by the user.
[0077] A "prompt message" refers to text input used to provide specific instructions or requests to a generative artificial intelligence.
[0078] "Computational vision" refers to the ability of artificial intelligence to analyze image data and understand it as visual information.
[0079] "Natural language processing technology" refers to the ability of artificial intelligence to understand and process human language.
[0080] The "data storage unit" refers to a storage device that stores the generated video data and provides access to various information exchange platforms.
[0081] An "information exchange platform" refers to a platform that allows users to share videos they have created and communicate with other users.
[0082] This invention is a system for generating and managing creative videos, primarily using a communication terminal, a computing device (server), and generative artificial intelligence.
[0083] The user first selects images and videos they have taken using a communication device. This communication device includes mobile information devices such as smartphones and tablets. Accordingly, an application on the communication device provides a user interface, allowing the user to visually select the materials.
[0084] Next, the user selects the video's story and characters through the application. This selection is done through pre-prepared templates, allowing for a variety of video formats depending on the user's choices. The selected materials and input information are transmitted to a computing device via a communication protocol. The computing device processes the received information and sends instructions to the generating artificial intelligence.
[0085] Generative artificial intelligence utilizes computational vision and natural language processing technologies to generate videos based on a given prompt. The prompt might be provided in the form of, for example, "Generate a video based on an adventure story, combining a flying dragon with a mountain landscape." The generative AI analyzes this prompt and automatically creates the corresponding visual elements and animations.
[0086] Once video generation is complete, the computing device checks for any inappropriate content. The system performs this process using frame-by-frame image recognition technology. After confirming that a safe and high-quality video has been generated, a notification is sent to the user's communication terminal.
[0087] Ultimately, users can preview the generated video on their communication device and, if satisfied with the content, share it on information exchange platforms such as social media or other platforms. This makes it easy and intuitive for users to create high-quality content and share it with many people.
[0088] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0089] Step 1:
[0090] The user launches the application on the communication terminal and selects images or videos. In this step, the communication terminal provides a touch interface to acquire images, allowing the user to select materials from their own library. The input is the images or videos selected by the user, and the output is these data being selected and ready to be sent to the next step.
[0091] Step 2:
[0092] The user selects the story and characters within the application. The input data consists of selections made from the provided template options. The communication terminal is designed to allow visual selection through the user interface. The output is story and character data reflecting the user's selections.
[0093] Step 3:
[0094] The communication terminal transmits the selected materials and configuration information to the computing device. The input consists of the image selected in step 1 and the story and character information selected in step 2. The output is the transfer of this data to the computing device via a secure communication protocol.
[0095] Step 4:
[0096] The server integrates the received data and invokes the generative artificial intelligence to issue instructions for video generation. The input consists of media materials and selection information sent from the terminal. The computing device processes these integrally, using the generative AI model to create prompts for video generation. The output is the sending of these prompts to the generative artificial intelligence.
[0097] Step 5:
[0098] The generative artificial intelligence generates videos based on prompt text and filters out inappropriate content. The input is prompt text received from a server, and data processing utilizes computational vision and natural language processing techniques. The output is the generated video, which is the final video version after filtering.
[0099] Step 6:
[0100] The server analyzes the generated video to check for any inappropriate elements. The input is video data generated by the generation artificial intelligence. The output is a verified and safe video. This video is then ready to be sent to the user. The server uses image recognition technology to perform this analysis process.
[0101] Step 7:
[0102] A notification of the generated video is sent to the user's communication device. The input is notification data from the server, and the output is the notification to the user and a previewable video display. The user watches the video using the playback function within the application.
[0103] Step 8:
[0104] Users share videos they enjoy on social media and other information exchange platforms. Input is the user's operation of the share button, and output is the posting of the video to the selected platform. The communication terminal provides the application with sharing options to support this.
[0105] (Application Example 1)
[0106] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0107] In today's digital society, there is a need for individual users to easily generate high-quality visual artifacts and effectively share them on public platforms such as social networks. However, conventional visual artifact generation systems struggle to intuitively generate personalized content tailored to the user's needs and themes, and they lack sufficient processes to verify the inappropriateness of the generated content. Therefore, the challenge lies in providing a system that is easy for users to use and delivers personalized, high-quality content.
[0108] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0109] In this invention, the server includes means for integrating received information material and selected information and transmitting commands to a generative intelligence; means for the generative intelligence to generate a visual composite based on the received information and detect inappropriate content; and means for verifying the generated visual composite. This makes it possible to securely generate personalized, high-quality visual composites based on themes and visual experiences selected by individual users, and to easily share them on social media.
[0110] A "communication device" is a device used to acquire and transmit information, and is primarily a terminal operated by the user.
[0111] A "central device" is a centralized management device for receiving and processing information, such as a cloud server.
[0112] A "visual medium" is a collection of visually recognizable information, such as images and videos.
[0113] "Information material" refers to the data that is transmitted from a communication device to a central device, and includes visual media.
[0114] "Story" refers to a series of storylines or themes used in video creation.
[0115] "Featured elements" refer to elements such as characters and objects used within the video.
[0116] "Selection information" refers to information about the story and elements that the user has chosen.
[0117] "Generative intelligence" refers to artificial intelligence that has the ability to automatically generate visual constructs based on received information.
[0118] A "visual synthesizer" is a visual content product, such as images or videos, generated by generative intelligence.
[0119] "Inappropriate content" refers to content that violates public order and morals or that may cause offense to users.
[0120] A "social network" is an online platform for users to publish or share content they have created.
[0121] An "information storage medium" is a system resource or device for storing data such as generated visual composites.
[0122] The invention will now be described in terms of embodiments. This application provides a system that automatically generates highly personalized videos that can be shared on social networks. The entire system consists of a communication device, a central device, and generative intelligence.
[0123] The communication device, or user terminal, allows the user to acquire visual media and select narrative elements. Specifically, smartphones and tablets are envisioned, which provide the functionality to select visual media from camera rolls, etc., and transmit them to the central device. This creates an intuitive and easy-to-use interface for material selection for the user.
[0124] The central server receives information materials and selected information transmitted from communication devices and integrates them. The integrated information is sent to a generative intelligence, which uses artificial intelligence technology to analyze the content and generate new visual artifacts. Specifically, it uses a combination of image recognition technology such as Google® Cloud Vision API and natural language processing technology. Through this process, high-quality content is automatically generated that reflects themes and viewing experiences based on the materials.
[0125] Finally, the generated visual composite is reviewed on the server to ensure it does not contain any inappropriate content before being returned to the communication device. Users can then review and edit this visual composite through the application and share it as needed through social networks or other information sharing media.
[0126] For example, if the theme is landscape photos taken during a trip, and the user selects a story and fantasy characters related to "Summer Adventure," the generative AI will use this information to create a visually engaging adventure video. An example of a prompt input to the generative AI model would be, "Create a 3-minute animated video with a beach adventure theme."
[0127] The above describes the modes for carrying out the invention.
[0128] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0129] Step 1:
[0130] The device captures or selects visual media from the user. Images and videos are displayed on the screen through the device's interface, allowing the user to search and select materials by date of capture, tags, etc. The input is the visual media selected by the user, and the output is the data of the selection results.
[0131] Step 2:
[0132] The terminal transmits information about the visual medium selected by the user, along with the story and elements specified by the user, to the central device. Data communication is used for transmission. The input is the selection information of the story and elements set by the user, and the output is the integrated data transmitted to the central device.
[0133] Step 3:
[0134] The server receives visual media and configuration information transmitted from the terminal, integrates them, and prepares them for submission to the generative intelligence. The received data is formatted according to the specified format, and prompt statements are generated. The input is the visual media and selection information data from the terminal, and the output is the prompt statements and integrated data for the generative intelligence.
[0135] Step 4:
[0136] The generative intelligence generates visual artifacts based on prompt text and integrated data sent from the server. Here, natural language processing is used to analyze the narrative, and image recognition techniques are employed to analyze the visual elements. The input consists of prompt text and integrated data from the server, and the output is the generated visual artifact.
[0137] Step 5:
[0138] The server reviews the visual constructs generated by generative intelligence and checks for inappropriate content. An algorithm for detecting inappropriate elements is incorporated. The input is the generated visual construct, and the output is the reviewed visual construct. If there are no problems, it is approved and ready to be sent to the terminal.
[0139] Step 6:
[0140] The terminal receives the final visual composite sent from the server and provides the user with preview and editing functions. After reviewing the video, the user can add text and music as needed. The input is the visual composite from the server, and the output is the final content that the user views.
[0141] Step 7:
[0142] Users choose whether or not to share the generated content on social networks, and through this action, they share the visual composite with other users. Sharing is done through the application, and links or embed codes are generated. The input is the sharing instruction based on the user's action, and the output is the link information of the shared content.
[0143] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0144] This invention relates to a system that recognizes a user's emotions in real time and generates customized videos based on those emotions. Here, we describe a detailed embodiment of an innovative video generation system that combines an emotion engine.
[0145] First, the user operates a communication terminal and launches the video generation application. The user selects the image or video files to be used in the video they want to generate using the communication terminal and sends them to the server.
[0146] Next, the user selects from multiple templates via an interface on their communication terminal to set up the story and characters. The selected information, along with the results of the emotion engine's analysis, is sent to the server.
[0147] Here, the emotion engine comes into play, collecting facial expression data through the user's camera. This engine analyzes the user's facial feature points and can identify emotions in real time. The analyzed emotion data influences decisions about what kind of story development and character behavior to adopt in video production.
[0148] Next, the server integrates the received image or video data, story and character information, and real-time analyzed emotion data. Based on this integrated data, the server issues a command to the generating AI to create a video. The generating AI then generates a video optimized for the user's current emotional state based on the emotion analysis results.
[0149] The videos created by the generation AI are verified by the server and then sent to the communication terminal. Users can preview the generated videos on the communication terminal and adjust them as needed.
[0150] (Specific example)
[0151] For example, consider a scenario where a user wants to generate a video to relieve work stress. The user selects a story with a relaxation theme and gentle characters. The communication device inputs the user's smile and calm expressions into the emotion engine via the camera. After the emotion engine detects the user's relaxed state, the generation AI creates a video with calming music and a beautiful natural landscape as the background. Finally, this video is tailored specifically for the user and previewed on the communication device. The user can then watch this relaxing video and even share it with other users.
[0152] This invention enables users to generate and enjoy content adapted to their individual emotional states, thus allowing for a more personalized experience.
[0153] The following describes the processing flow.
[0154] Step 1:
[0155] The user launches a video generation app installed on their communication device. The app's UI is displayed, and the user selects the image or video saved for the video they want to generate.
[0156] Step 2:
[0157] The device sends the image or video data selected by the user to the server. At this point, the data is packetized and securely transferred to the server.
[0158] Step 3:
[0159] The user selects from story templates and character options provided using the communication terminal interface, thereby determining the concept of the video.
[0160] Step 4:
[0161] The device sends story and character selection information to the server. This data will be used for future video generation.
[0162] Step 5:
[0163] An emotion engine operates on the device and uses the user's camera to collect facial expression data. The emotion engine analyzes this data and recognizes emotions in real time.
[0164] Step 6:
[0165] The device sends the analyzed emotional data, along with story and character information, to the server. In this way, the user's emotional state is reflected in the video generation.
[0166] Step 7:
[0167] The server integrates all received data and sends instructions to the generating AI. Based on the integrated data, the generating AI creates a video adapted to the user's emotions.
[0168] Step 8:
[0169] The generation AI takes into account the results of emotion analysis and creates videos that match the specified story and characters. Music and special effects are added to the videos to enhance their quality.
[0170] Step 9:
[0171] The server receives the video sent from the generating AI and reviews its content. After reconfirming that there are no inappropriate elements, it retains the final version of the video.
[0172] Step 10:
[0173] The server sends the final version of the video to the user's device and notifies them that the generation is complete.
[0174] Step 11:
[0175] Users can preview the generated video on their device and review its content. They can make adjustments as needed, and if they are satisfied, they can save or share the video.
[0176] (Example 2)
[0177] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0178] Traditional video generation systems have a problem in automatically generating personalized content based on user emotions. In particular, the process of generating videos that reflect the user's real-time emotional state is complex and often requires manual adjustments. As a result, it has been difficult to quickly and efficiently obtain videos that satisfy users.
[0179] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0180] In this invention, the server includes means for integrating received media material, selection information, and emotion data; means for transmitting instructions to a generating artificial intelligence; and means for reviewing the generated video. This enables the automatic generation of videos optimized for the user's emotional state.
[0181] A "communication terminal" is a device used by a user to acquire images and videos, and to send and receive information with a server.
[0182] A "server" is a central processing unit that receives data from communication terminals, integrates emotional data and selection information, and manages instructions for the generating artificial intelligence.
[0183] "Generative artificial intelligence" is an algorithm that generates videos optimized for the user's emotional state based on information received from a server.
[0184] An "emotion analysis device" is a mechanism that analyzes the feature points of a user's face and identifies emotional data in real time.
[0185] A "prompt statement" is a command statement created by the server to instruct the AI model on the specifications for video generation.
[0186] A "video" is a dynamic visual medium generated based on user emotional data, and is content that appeals to both sight and hearing.
[0187] A "data storage device" is a storage device that stores generated video data and enables users to access and share it.
[0188] This invention is a system that recognizes a user's emotions in real time and generates a customized video based on those emotions. The following describes an embodiment of this system in detail.
[0189] The user launches a video generation application using a communication terminal. This terminal refers to a typical computer or mobile device equipped with a camera, which can acquire image and video files. Through the application, the user selects the media materials to be used for the video and sends them to the server.
[0190] This communication terminal is equipped with an emotion analysis device that collects user facial expression data. This device analyzes the user's facial feature points and has the function of identifying their emotional state in real time. The analysis results are transmitted from the communication terminal to a server.
[0191] The server integrates the received media materials, the story and character information selected by the user, and the sentiment analysis results. This integrated data is used as a prompt message to send instructions to the generation AI model. This prompt message is a set of commands that specify the video generation specifications and is written in a format that the generation AI model can understand.
[0192] The generative AI model generates videos appropriate to the user's emotional state based on the received prompt text. This AI model operates using machine learning algorithms combined with computer vision and natural language processing techniques. The generated videos are reviewed by a server and then sent to a communication terminal. On this terminal, the user can preview the video and adjust or modify it as needed.
[0193] As a concrete example, a user might want to generate a "relaxing video." The user selects a relaxation-themed template and uses the camera on their communication device to have their facial expressions detected by an emotion analysis device. The server sends a prompt message to the generation AI model: "Generate a video incorporating calming music and natural scenery." The generation AI model generates the video based on these instructions, and it is finally delivered to the communication device so that the user can view it.
[0194] In this way, this system can efficiently generate and deliver personalized video content that responds to the user's emotional state.
[0195] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0196] Step 1:
[0197] The user launches a video generation application using a communication terminal and selects image or video files within it. This serves as the input. The communication terminal acquires the selected media material as digital data and sends it to the server. Here, data format conversion and data transfer based on the communication protocol take place.
[0198] Step 2:
[0199] The user selects story and character templates using an interface on a communication terminal. The selected information becomes the input. The communication terminal sends this selection information to the server as digital data. In this step, the user's selected information is organized and packaged as data.
[0200] Step 3:
[0201] The emotion analysis device in the communication terminal captures the user's facial expressions with a camera and collects emotion data. The captured facial data becomes the input. This data is analyzed in real time, and the emotional state is identified based on the user's facial feature points. The analysis results are sent to the server as emotion data.
[0202] Step 4:
[0203] The server integrates the media material, selection information, and sentiment data received in steps 1 through 3. Based on this input data, it creates prompts for the generative AI model. The integrated data is then processed and formatted to create the video scenario.
[0204] Step 5:
[0205] The server sends a prompt to the generative AI model, issuing a command to generate a video. The generative AI model receives the prompt as input and generates a video optimized for the user's emotional state. Here, the generative AI utilizes computer vision and natural language processing technologies to perform data calculations and automatically construct the video storyline.
[0206] Step 6:
[0207] The server reviews the generated videos, verifying their content and performing quality checks. If inappropriate content is found, an automatic filter is activated. The reviewed videos are sent to the communication terminal. Additional data processing is performed for security and quality control.
[0208] Step 7:
[0209] The user previews the transmitted video on their communication device. The user reviews the video content and edits or adjusts it as needed. The user's feedback serves as input, and they have the option to regenerate the video. Finally, once a satisfactory video is completed, it can be shared with other users.
[0210] (Application Example 2)
[0211] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0212] In recent years, while a wide variety of video content has been available, it has been difficult to provide users with personalized content that is tailored to their emotional state. In this context, there is a need for a system that allows users to receive optimal video content in real time based on their emotions, thereby improving their experience.
[0213] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0214] In this invention, the server includes means for integrating received information materials and selection information and transmitting instructions to a generation processing device; means for the generation processing device to acquire the user's emotional state through emotion analysis and optimize its operation based on the emotional state; and means for the communication device to feed back the emotional state determined by emotion analysis to the user and present recommended videos that correspond to the user's emotions. This makes it possible to provide personalized video content that corresponds to the user's emotions.
[0215] "Communication equipment" is a general term for devices used by users to exchange information with the outside world.
[0216] "Information material" refers to data such as images and videos that users submit, and is processed on the server.
[0217] An "information processing device" refers to a system that integrates received data and transmits necessary instructions to a generation processing device.
[0218] A "generation processing device" refers to a device that generates and optimizes video content using emotion analysis technology based on received information.
[0219] "Emotional analysis" refers to a technology that analyzes a user's facial expressions and reactions in real time to identify their emotional state.
[0220] "Feedback" refers to a mechanism where the system returns the results and information it has analyzed to the user, influencing the user's choices and actions.
[0221] "Recommended videos" refer to video content provided to users that has been pre-selected based on their emotional state.
[0222] The system consists of a communication device, an information processing device, and a generation processing device. The communication device acquires images and videos from the user and transmits them to the information processing device as information material. The information processing device integrates the received information material and sends instructions to the generation processing device. The generation processing device uses emotion analysis technology to acquire the user's emotional state and optimizes its operation based on the acquired emotional data.
[0223] The system uses machine vision technology for emotion analysis. Specifically, it uses a smartphone camera and software such as OpenCV to analyze the user's emotions in real time from their facial expressions. It also uses services such as Microsoft® Azure® Face API to identify emotional states.
[0224] Based on these analysis results, the information processing device proposes personalized video content using a generative AI model and presents the video generated from the communication device to the user. At this time, the communication device can provide the user with feedback on the analyzed emotional state and present recommended videos based on the user's response.
[0225] For example, if a user is feeling stressed, the generation processing unit will identify the emotional state as "stressed" and instruct it to generate "footage of a forest with calming music in the background." In this way, by prompting the generation AI model with a message such as, "The user's current emotional state is relaxed. Please generate a video with calming music and a beautiful natural landscape in the background," it is possible to provide content that is appropriate for the user.
[0226] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0227] Step 1:
[0228] The communication terminal acquires user images and video data and transmits them to the information processing device as information material. The input is the user's images and video data, and the output is the transmission of that data to the information processing device. In this step, the communication terminal uses its camera function to collect images of the user.
[0229] Step 2:
[0230] The information processing device integrates the received information material with the user's selected story and character selection information. The input consists of the information material and selection information, and the output is this integrated data. This integrated data prepares the device for optimal video generation.
[0231] Step 3:
[0232] The information processing device transmits instructions to the generation processing device based on integrated data. The input is integrated data, and the output is the instructions transmitted to the generation processing device. The information processing device organizes the received information and provides it to the generation processing device in a processable format.
[0233] Step 4:
[0234] The generation and processing device analyzes the user's emotions using machine vision technology based on the received instructions. The input is instruction data, and the output is data on the user's emotional state. Here, for example, facial feature points are analyzed using OpenCV, and emotions are determined through the Microsoft Azure Face API.
[0235] Step 5:
[0236] The generation processing device generates videos using a generation AI model based on the emotion analysis results. The input is the emotion analysis results, and the output is video data optimized for the user's emotions. The device generates prompt text and instructs the generation AI model, such as "The user's current emotional state is relaxed. Please generate a video with calming music and a beautiful natural landscape as the background," to generate the video.
[0237] Step 6:
[0238] The generated video is sent to the information processing device and prepared for display to the user. The input is the generated video data, and the output is displayed on the communication terminal. This step also includes a final check to ensure the video does not contain any inappropriate content.
[0239] Step 7:
[0240] The communication terminal presents the generated video to the user and obtains user feedback. The input is emotional data for presenting the video to the user and obtaining feedback, while the output is user response data. This aims to improve the overall system responsiveness and user experience.
[0241] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0242] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0243] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0244] [Second Embodiment]
[0245] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0246] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0247] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0248] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0249] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0250] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0251] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0252] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0253] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0254] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0255] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0256] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0257] This invention is a system centered around a user's communication terminal and server, utilizing generation AI to generate and manage creative and free-form videos. The functions and procedures of each component are described in detail below.
[0258] First, the user uses a communication device such as a smartphone to select images or videos they have taken or saved. This communication device has the function of sending the selected images or videos to a server. In terms of configuration, an application on the device provides the user interface, allowing the user to intuitively select the material they want to turn into a video.
[0259] Next, the user selects a story and characters to serve as a template for video generation via an application on their device. These choices are necessary to form the video's structure, visual elements, and scenario, and various templates are provided in advance within the application. The user can combine these options to create their own unique settings according to their preferences.
[0260] The selected materials and settings are sent from the terminal to the server. The server receives and integrates this input data and instructs the generative AI to generate a new video. The generative AI uses computer vision and natural language processing techniques to automatically create a video depicting the specified story and characters. The system also has a function to automatically detect inappropriate content, ensuring that no elements that may offend the user are generated.
[0261] Once the video generation is complete, it undergoes a verification process on the server and is then notified to the user's communication device. Upon receiving the notification, the user can preview the generated video within the application. This video is presented to the user as comprehensively edited content, including visual information and music or sound effects. The user can then choose to share the video on other platforms such as social media, or publish it within the platform itself.
[0262] (Specific example)
[0263] Suppose a user takes a photo of a beautiful landscape while traveling. This user launches a video generation app and selects the landscape photo. Next, the user chooses the theme "Fantasy Adventure" as the story theme and selects a dragon as the character. Based on this information, a generation AI summoned by the server creates an animation of a dragon flying through the sky and over mountains, accompanied by a grand background music. After checking the quality and appropriateness of the video, the user watches the finished video, is satisfied, and decides to share it with friends on social media. In this way, the present invention provides a form for users to intuitively and easily generate high-quality content.
[0264] The following describes the processing flow.
[0265] Step 1:
[0266] The user launches an application on the communication terminal and selects images or videos from the terminal to be used as material for video generation. After the user's selection, the communication terminal packets the selected media files and prepares to send them to the server.
[0267] Step 2:
[0268] The device sends the image or video data selected by the user to the server. This transmission usually occurs in real time via an internet connection.
[0269] Step 3:
[0270] Users select story templates and character models presented through the application's UI, thereby concretizing their concept for the type of video they want to create.
[0271] Step 4:
[0272] The device sends the user's setting selections, along with story and character information, to the server in a defined data format.
[0273] Step 5:
[0274] The server integrates the image or video data received from the terminal and the story and character selection information. This integrated data serves as the basis for generating a new video.
[0275] Step 6:
[0276] The server inputs this integrated data into the generation AI and issues an instruction to start the video generation process.
[0277] Step 7:
[0278] The generation AI utilizes computer vision and natural language processing technologies to generate a video based on the integrated data. The AI also checks whether inappropriate content is included during the generation process.
[0279] Step 8:
[0280] The server receives the video sent from the generation AI and reviews the video content for final confirmation. If necessary, an instruction for correction can be sent back to the generation AI.
[0281] Step 9:
[0282] The server converts the data of the confirmed video into an appropriate format and prepares to save it in a form accessible to the user.
[0283] Step 10:
[0284] The server notifies the user that the video has been generated and is ready for publication. This notification is sent to the terminal and reflected in the application.
[0285] Step 11:
[0286] When the terminal receives a notification from the server, it provides the user with a preview of the newly generated video. The user can view this preview and select options for publication or sharing.
[0287] (Example 1)
[0288] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0289] In a conventional video generation system, there is a problem that it is difficult for a user to quickly generate a video that is creatively and individually desired. In addition, filtering of inappropriate content is insufficient, which may cause discomfort to the user. Thus, efficient and high-quality video generation is required.
[0290] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0291] In this invention, the server includes means for integrating the received media material and selection information and transmitting an instruction to the generation artificial intelligence, means for the generation artificial intelligence to generate a video based on the received information and detect and filter inappropriate content, and means for executing a process of investigating the video generated by the generation artificial intelligence. Thereby, it becomes possible to quickly generate a safe and high-quality video for the user and improve the usability of the product.
[0292] A "communication terminal" is a device that has a function for a user to acquire images and videos and transmit information to a computing device.
[0293] A "computing device" is a device that has a function for processing the received media material and selection information and transmitting an instruction to the generation artificial intelligence.
[0294] A "generation artificial intelligence" is an artificial intelligence technology that has the ability to generate a video based on the received information and detect and filter inappropriate content.
[0295] "Media material" refers to image and video data acquired by users using communication devices.
[0296] "Story and characters" refers to the visual and contextual elements necessary for video creation, as selected by the user.
[0297] A "prompt message" refers to text input used to provide specific instructions or requests to a generative artificial intelligence.
[0298] "Computational vision" refers to the ability of artificial intelligence to analyze image data and understand it as visual information.
[0299] "Natural language processing technology" refers to the ability of artificial intelligence to understand and process human language.
[0300] The "data storage unit" refers to a storage device that stores the generated video data and provides access to various information exchange platforms.
[0301] An "information exchange platform" refers to a platform that allows users to share videos they have created and communicate with other users.
[0302] This invention is a system for generating and managing creative videos, primarily using a communication terminal, a computing device (server), and generative artificial intelligence.
[0303] The user first selects images and videos they have taken using a communication device. This communication device includes mobile information devices such as smartphones and tablets. Accordingly, an application on the communication device provides a user interface, allowing the user to visually select the materials.
[0304] Next, the user selects the story and characters of the video through the application. This selection is made through a pre-prepared template, and various video formats are made possible by the user's selection. The selected materials and input information are sent to the computing device via a communication protocol. The computing device processes the received information and sends instructions to the generative artificial intelligence.
[0305] The generative artificial intelligence utilizes computer vision and natural language processing technologies to generate a video based on the given prompt text. The prompt text is provided in a form such as "Please generate a video combining a flying dragon and mountain scenery based on an adventure story". The generative artificial intelligence analyzes this prompt text and automatically creates corresponding visual elements and animations.
[0306] When the generation of the video is completed, the computing device checks whether there is any inappropriate content. The system executes this process using frame-by-frame image recognition technology. After that, it is confirmed that a safe and high-quality video has been generated, and a notification is sent to the user's communication terminal.
[0307] Finally, the user can preview the generated video on the communication terminal and, if satisfied with the content, share it on an information exchange platform, that is, SNS or other platforms. This enables the user to easily and intuitively create high-quality content and share it with many people.
[0308] The flow of the specific process in Example 1 will be described using FIG. 11.
[0309] Step 1:
[0310] The user launches the application on the communication terminal and selects images or videos. In this step, since the user selects materials from their own library, the communication terminal provides an interface for acquiring images through touch operations. The input is the images and videos selected by the user, and the output is that these data are selected and ready to be sent to the next step.
[0311] Step 2:
[0312] The user selects the story and characters within the application. The input data consists of selections made from the provided template options. The communication terminal is designed to allow visual selection through the user interface. The output is story and character data reflecting the user's selections.
[0313] Step 3:
[0314] The communication terminal transmits the selected materials and configuration information to the computing device. The input consists of the image selected in step 1 and the story and character information selected in step 2. The output is the transfer of this data to the computing device via a secure communication protocol.
[0315] Step 4:
[0316] The server integrates the received data and invokes the generative artificial intelligence to issue instructions for video generation. The input consists of media materials and selection information sent from the terminal. The computing device processes these integrally, using the generative AI model to create prompts for video generation. The output is the sending of these prompts to the generative artificial intelligence.
[0317] Step 5:
[0318] The generative artificial intelligence generates videos based on prompt text and filters out inappropriate content. The input is prompt text received from a server, and data processing utilizes computational vision and natural language processing techniques. The output is the generated video, which is the final video version after filtering.
[0319] Step 6:
[0320] The server analyzes the generated video to check for any inappropriate elements. The input is video data generated by the generation artificial intelligence. The output is a verified and safe video. This video is then ready to be sent to the user. The server uses image recognition technology to perform this analysis process.
[0321] Step 7:
[0322] A notification of the generated video is sent to the user's communication device. The input is notification data from the server, and the output is the notification to the user and a previewable video display. The user watches the video using the playback function within the application.
[0323] Step 8:
[0324] Users share videos they enjoy on social media and other information exchange platforms. Input is the user's operation of the share button, and output is the posting of the video to the selected platform. The communication terminal provides the application with sharing options to support this.
[0325] (Application Example 1)
[0326] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0327] In today's digital society, there is a need for individual users to easily generate high-quality visual artifacts and effectively share them on public platforms such as social networks. However, conventional visual artifact generation systems struggle to intuitively generate personalized content tailored to the user's needs and themes, and they lack sufficient processes to verify the inappropriateness of the generated content. Therefore, the challenge lies in providing a system that is easy for users to use and delivers personalized, high-quality content.
[0328] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0329] In this invention, the server includes means for integrating received information material and selected information and transmitting commands to a generative intelligence; means for the generative intelligence to generate a visual composite based on the received information and detect inappropriate content; and means for verifying the generated visual composite. This makes it possible to securely generate personalized, high-quality visual composites based on themes and visual experiences selected by individual users, and to easily share them on social media.
[0330] A "communication device" is a device used to acquire and transmit information, and is primarily a terminal operated by the user.
[0331] A "central device" is a centralized management device for receiving and processing information, such as a cloud server.
[0332] A "visual medium" is a collection of visually recognizable information, such as images and videos.
[0333] "Information material" refers to the data that is transmitted from a communication device to a central device, and includes visual media.
[0334] "Story" refers to a series of storylines or themes used in video creation.
[0335] "Featured elements" refer to elements such as characters and objects used within the video.
[0336] "Selection information" refers to information about the story and elements that the user has chosen.
[0337] "Generative intelligence" refers to artificial intelligence that has the ability to automatically generate visual constructs based on received information.
[0338] A "visual synthesizer" is a visual content product, such as images or videos, generated by generative intelligence.
[0339] "Inappropriate content" refers to content that violates public order and morals or that may cause offense to users.
[0340] A "social network" is an online platform for users to publish or share content they have created.
[0341] An "information storage medium" is a system resource or device for storing data such as generated visual composites.
[0342] The invention will now be described in terms of embodiments. This application provides a system that automatically generates highly personalized videos that can be shared on social networks. The entire system consists of a communication device, a central device, and generative intelligence.
[0343] The communication device, or user terminal, allows the user to acquire visual media and select narrative elements. Specifically, smartphones and tablets are envisioned, which provide the functionality to select visual media from camera rolls, etc., and transmit them to the central device. This creates an intuitive and easy-to-use interface for material selection for the user.
[0344] The central server receives information materials and selected information transmitted from communication devices and integrates them. The integrated information is sent to generative intelligence, which uses artificial intelligence technology to analyze the content and generate new visual artifacts. Specifically, it uses a combination of image recognition technology such as the Google Cloud Vision API and natural language processing technology. Through this process, high-quality content is automatically generated that reflects themes and viewing experiences based on the materials.
[0345] Finally, the generated visual composite is reviewed on the server to ensure it does not contain any inappropriate content before being returned to the communication device. Users can then review and edit this visual composite through the application and share it as needed through social networks or other information sharing media.
[0346] For example, if the theme is landscape photos taken during a trip, and the user selects a story and fantasy characters related to "Summer Adventure," the generative AI will use this information to create a visually engaging adventure video. An example of a prompt input to the generative AI model would be, "Create a 3-minute animated video with a beach adventure theme."
[0347] The above describes the modes for carrying out the invention.
[0348] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0349] Step 1:
[0350] The device captures or selects visual media from the user. Images and videos are displayed on the screen through the device's interface, allowing the user to search and select materials by date of capture, tags, etc. The input is the visual media selected by the user, and the output is the data of the selection results.
[0351] Step 2:
[0352] The terminal transmits information about the visual medium selected by the user, along with the story and elements specified by the user, to the central device. Data communication is used for transmission. The input is the selection information of the story and elements set by the user, and the output is the integrated data transmitted to the central device.
[0353] Step 3:
[0354] The server receives visual media and configuration information transmitted from the terminal, integrates them, and prepares them for submission to the generative intelligence. The received data is formatted according to the specified format, and prompt statements are generated. The input is the visual media and selection information data from the terminal, and the output is the prompt statements and integrated data for the generative intelligence.
[0355] Step 4:
[0356] The generative intelligence generates visual artifacts based on prompt text and integrated data sent from the server. Here, natural language processing is used to analyze the narrative, and image recognition techniques are employed to analyze the visual elements. The input consists of prompt text and integrated data from the server, and the output is the generated visual artifact.
[0357] Step 5:
[0358] The server reviews the visual constructs generated by generative intelligence and checks for inappropriate content. An algorithm for detecting inappropriate elements is incorporated. The input is the generated visual construct, and the output is the reviewed visual construct. If there are no problems, it is approved and ready to be sent to the terminal.
[0359] Step 6:
[0360] The terminal receives the final visual composite sent from the server and provides the user with preview and editing functions. After reviewing the video, the user can add text and music as needed. The input is the visual composite from the server, and the output is the final content that the user views.
[0361] Step 7:
[0362] Users choose whether or not to share the generated content on social networks, and through this action, they share the visual composite with other users. Sharing is done through the application, and links or embed codes are generated. The input is the sharing instruction based on the user's action, and the output is the link information of the shared content.
[0363] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0364] This invention relates to a system that recognizes a user's emotions in real time and generates customized videos based on those emotions. Here, we describe a detailed embodiment of an innovative video generation system that combines an emotion engine.
[0365] First, the user operates a communication terminal and launches the video generation application. The user selects the image or video files to be used in the video they want to generate using the communication terminal and sends them to the server.
[0366] Next, the user selects from multiple templates via an interface on their communication terminal to set up the story and characters. The selected information, along with the results of the emotion engine's analysis, is sent to the server.
[0367] Here, the emotion engine comes into play, collecting facial expression data through the user's camera. This engine analyzes the user's facial feature points and can identify emotions in real time. The analyzed emotion data influences decisions about what kind of story development and character behavior to adopt in video production.
[0368] Next, the server integrates the received image or video data, story and character information, and real-time analyzed emotion data. Based on this integrated data, the server issues a command to the generating AI to create a video. The generating AI then generates a video optimized for the user's current emotional state based on the emotion analysis results.
[0369] The videos created by the generation AI are verified by the server and then sent to the communication terminal. Users can preview the generated videos on the communication terminal and adjust them as needed.
[0370] (Specific example)
[0371] For example, consider a scenario where a user wants to generate a video to relieve work stress. The user selects a story with a relaxation theme and gentle characters. The communication device inputs the user's smile and calm expressions into the emotion engine via the camera. After the emotion engine detects the user's relaxed state, the generation AI creates a video with calming music and a beautiful natural landscape as the background. Finally, this video is tailored specifically for the user and previewed on the communication device. The user can then watch this relaxing video and even share it with other users.
[0372] This invention enables users to generate and enjoy content adapted to their individual emotional states, thus allowing for a more personalized experience.
[0373] The following describes the processing flow.
[0374] Step 1:
[0375] The user launches a video generation app installed on their communication device. The app's UI is displayed, and the user selects images or videos saved for the video they want to generate.
[0376] Step 2:
[0377] The device sends the image or video data selected by the user to the server. At this point, the data is packetized and securely transferred to the server.
[0378] Step 3:
[0379] The user selects from story templates and character options provided using the communication terminal interface, thereby determining the concept of the video.
[0380] Step 4:
[0381] The device sends story and character selection information to the server. This data will be used for future video generation.
[0382] Step 5:
[0383] An emotion engine operates on the device and uses the user's camera to collect facial expression data. The emotion engine analyzes this data and recognizes emotions in real time.
[0384] Step 6:
[0385] The device sends the analyzed emotional data, along with story and character information, to the server. In this way, the user's emotional state is reflected in the video generation.
[0386] Step 7:
[0387] The server integrates all received data and sends instructions to the generating AI. Based on the integrated data, the generating AI creates a video adapted to the user's emotions.
[0388] Step 8:
[0389] The generation AI takes into account the results of emotion analysis and creates videos tailored to the specified story and characters. Music and special effects are added to the videos to enhance their quality.
[0390] Step 9:
[0391] The server receives the video sent from the generating AI and reviews its content. After reconfirming that there are no inappropriate elements, it retains the final version of the video.
[0392] Step 10:
[0393] The server sends the final version of the video to the user's device and notifies them that the generation is complete.
[0394] Step 11:
[0395] Users can preview the generated video on their device and review its content. They can make adjustments as needed, and if they are satisfied, they can save or share the video.
[0396] (Example 2)
[0397] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0398] Traditional video generation systems have a problem in automatically generating personalized content based on user emotions. In particular, the process of generating videos that reflect the user's real-time emotional state is complex and often requires manual adjustments. As a result, it has been difficult to quickly and efficiently obtain videos that satisfy users.
[0399] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0400] In this invention, the server includes means for integrating received media material, selection information, and emotion data; means for transmitting instructions to a generating artificial intelligence; and means for reviewing the generated video. This enables the automatic generation of videos optimized for the user's emotional state.
[0401] A "communication terminal" is a device used by a user to acquire images and videos, and to send and receive information with a server.
[0402] A "server" is a central processing unit that receives data from communication terminals, integrates emotional data and selection information, and manages instructions for the generating artificial intelligence.
[0403] "Generative artificial intelligence" is an algorithm that generates videos optimized for the user's emotional state based on information received from a server.
[0404] An "emotion analysis device" is a mechanism that analyzes the feature points of a user's face and identifies emotional data in real time.
[0405] A "prompt statement" is a command statement created by the server to instruct the AI model on the specifications for video generation.
[0406] A "video" is a dynamic visual medium generated based on user emotional data, and is content that appeals to both sight and hearing.
[0407] A "data storage device" is a storage device that stores generated video data and enables users to access and share it.
[0408] This invention is a system that recognizes a user's emotions in real time and generates a customized video based on those emotions. The following describes an embodiment of this system in detail.
[0409] The user launches a video generation application using a communication terminal. This terminal refers to a typical computer or mobile device equipped with a camera, which can acquire image and video files. Through the application, the user selects the media materials to be used for the video and sends them to the server.
[0410] This communication terminal is equipped with an emotion analysis device that collects user facial expression data. This device analyzes the user's facial features and has the function of identifying their emotional state in real time. The analysis results are transmitted from the communication terminal to a server.
[0411] The server integrates the received media materials, the story and character information selected by the user, and the sentiment analysis results. This integrated data is used as a prompt message to send instructions to the generation AI model. This prompt message is a set of commands that specify the video generation specifications and is written in a format that the generation AI model can understand.
[0412] The generative AI model generates videos appropriate to the user's emotional state based on the received prompt text. This AI model operates using machine learning algorithms combined with computer vision and natural language processing techniques. The generated videos are reviewed by a server and then sent to a communication terminal. On this terminal, the user can preview the video and adjust or modify it as needed.
[0413] As a concrete example, a user might want to generate a "relaxing video." The user selects a relaxation-themed template and uses the camera on their communication device to have their facial expressions detected by an emotion analysis device. The server sends a prompt message to the generation AI model: "Generate a video incorporating calming music and natural scenery." The generation AI model generates the video based on these instructions, and it is finally delivered to the communication device so that the user can view it.
[0414] In this way, this system can efficiently generate and deliver personalized video content that responds to the user's emotional state.
[0415] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0416] Step 1:
[0417] The user launches a video generation application using a communication terminal and selects image or video files within it. This serves as the input. The communication terminal acquires the selected media material as digital data and sends it to the server. Here, data format conversion and data transfer based on the communication protocol take place.
[0418] Step 2:
[0419] The user selects story and character templates using an interface on a communication terminal. The selected information becomes the input. The communication terminal sends this selection information to the server as digital data. In this step, the user's selected information is organized and packaged as data.
[0420] Step 3:
[0421] The emotion analysis device in the communication terminal captures the user's facial expressions with a camera and collects emotion data. The captured facial data becomes the input. This data is analyzed in real time, and the emotional state is identified based on the user's facial feature points. The analysis results are sent to the server as emotion data.
[0422] Step 4:
[0423] The server integrates the media material, selection information, and sentiment data received in steps 1 through 3. Based on this input data, it creates prompts for the generative AI model. The integrated data is then processed and formatted to create the video scenario.
[0424] Step 5:
[0425] The server sends a prompt to the generative AI model, issuing a command to generate a video. The generative AI model receives the prompt as input and generates a video optimized for the user's emotional state. Here, the generative AI utilizes computer vision and natural language processing technologies to perform data calculations and automatically construct the video storyline.
[0426] Step 6:
[0427] The server reviews the generated videos, verifying their content and performing quality checks. If inappropriate content is found, an automatic filter is activated. The reviewed videos are sent to the communication terminal. Additional data processing is performed for security and quality control.
[0428] Step 7:
[0429] The user previews the transmitted video on their communication device. The user reviews the video content and edits or adjusts it as needed. The user's feedback serves as input, and they have the option to regenerate the video. Finally, once a satisfactory video is completed, it can be shared with other users.
[0430] (Application Example 2)
[0431] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0432] In recent years, while a wide variety of video content has been available, it has been difficult to provide users with personalized content that is tailored to their emotional state. In this context, there is a need for a system that allows users to receive optimal video content in real time based on their emotions, thereby improving their experience.
[0433] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0434] In this invention, the server includes means for integrating received information materials and selection information and transmitting instructions to a generation processing device; means for the generation processing device to acquire the user's emotional state through emotion analysis and optimize its operation based on the emotional state; and means for the communication device to feed back the emotional state determined by emotion analysis to the user and present recommended videos that correspond to the user's emotions. This makes it possible to provide personalized video content that corresponds to the user's emotions.
[0435] "Communication equipment" is a general term for devices used by users to exchange information with the outside world.
[0436] "Information material" refers to data such as images and videos that users submit, and is processed on the server.
[0437] An "information processing device" refers to a system that integrates received data and transmits necessary instructions to a generation processing device.
[0438] A "generation processing device" refers to a device that generates and optimizes video content using emotion analysis technology based on received information.
[0439] "Emotional analysis" refers to a technology that analyzes a user's facial expressions and reactions in real time to identify their emotional state.
[0440] "Feedback" refers to a mechanism where the system returns the results and information it has analyzed to the user, influencing the user's choices and actions.
[0441] "Recommended videos" refer to video content provided to users that has been pre-selected based on their emotional state.
[0442] The system consists of a communication device, an information processing device, and a generation processing device. The communication device acquires images and videos from the user and transmits them to the information processing device as information material. The information processing device integrates the received information material and sends instructions to the generation processing device. The generation processing device uses emotion analysis technology to acquire the user's emotional state and optimizes its operation based on the acquired emotional data.
[0443] The system utilizes machine vision technology for emotion analysis. Specifically, it uses a smartphone camera and software such as OpenCV to analyze the user's emotions in real time from their facial expressions. It also uses services such as Microsoft Azure Face API to identify emotional states.
[0444] Based on these analysis results, the information processing device proposes personalized video content using a generative AI model and presents the video generated from the communication device to the user. At this time, the communication device can provide the user with feedback on the analyzed emotional state and present recommended videos based on the user's response.
[0445] For example, if a user is feeling stressed, the generation processing unit will identify the emotional state as "stressed" and instruct it to generate "footage of a forest with calming music in the background." In this way, by prompting the generation AI model with a message such as, "The user's current emotional state is relaxed. Please generate a video with calming music and a beautiful natural landscape in the background," it is possible to provide content that is appropriate for the user.
[0446] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0447] Step 1:
[0448] The communication terminal acquires user images and video data and transmits them to the information processing device as information material. The input is the user's images and video data, and the output is the transmission of that data to the information processing device. In this step, the communication terminal uses its camera function to collect images of the user.
[0449] Step 2:
[0450] The information processing device integrates the received information material with the user's selected story and character selection information. The input consists of the information material and selection information, and the output is this integrated data. This integrated data prepares the device for optimal video generation.
[0451] Step 3:
[0452] The information processing device transmits instructions to the generation processing device based on integrated data. The input is integrated data, and the output is the instructions transmitted to the generation processing device. The information processing device organizes the received information and provides it to the generation processing device in a processable format.
[0453] Step 4:
[0454] The generation and processing device analyzes the user's emotions using machine vision technology based on the received instructions. The input is instruction data, and the output is data on the user's emotional state. Here, for example, facial feature points are analyzed using OpenCV, and emotions are determined through the Microsoft Azure Face API.
[0455] Step 5:
[0456] The generation processing device generates videos using a generation AI model based on the emotion analysis results. The input is the emotion analysis results, and the output is video data optimized for the user's emotions. The device generates prompt text and instructs the generation AI model, such as "The user's current emotional state is relaxed. Please generate a video with calming music and a beautiful natural landscape as the background," to generate the video.
[0457] Step 6:
[0458] The generated video is sent to the information processing device and prepared for display to the user. The input is the generated video data, and the output is displayed on the communication terminal. This step also includes a final check to ensure the video does not contain any inappropriate content.
[0459] Step 7:
[0460] The communication terminal presents the generated video to the user and obtains user feedback. The input is emotional data for presenting the video to the user and obtaining feedback, while the output is user response data. This aims to improve the overall system responsiveness and user experience.
[0461] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0462] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0463] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0464] [Third Embodiment]
[0465] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0466] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0467] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0468] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0469] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0470] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0471] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0472] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0473] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0474] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0475] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0476] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0477] This invention is a system centered around a user's communication terminal and server, utilizing generation AI to generate and manage creative and free-form videos. The functions and procedures of each component are described in detail below.
[0478] First, the user uses a communication device such as a smartphone to select images or videos they have taken or saved. This communication device has the function of sending the selected images or videos to a server. In terms of configuration, an application on the device provides the user interface, allowing the user to intuitively select the material they want to turn into a video.
[0479] Next, the user selects a story and characters to serve as a template for video generation via an application on their device. These choices are necessary to form the video's structure, visual elements, and scenario, and various templates are provided in advance within the application. The user can combine these options to create their own unique settings according to their preferences.
[0480] The selected materials and settings are sent from the terminal to the server. The server receives and integrates this input data and instructs the generative AI to generate a new video. The generative AI uses computer vision and natural language processing techniques to automatically create a video depicting the specified story and characters. The system also has a function to automatically detect inappropriate content, ensuring that no elements that may offend the user are generated.
[0481] Once the video generation is complete, it undergoes a verification process on the server and is then notified to the user's communication device. Upon receiving the notification, the user can preview the generated video within the application. This video is presented to the user as comprehensively edited content, including visual information and music or sound effects. The user can then choose to share the video on other platforms such as social media, or publish it within the platform itself.
[0482] (Specific example)
[0483] Suppose a user takes a photo of a beautiful landscape while traveling. This user launches a video generation app and selects the landscape photo. Next, the user chooses the theme "Fantasy Adventure" as the story theme and selects a dragon as the character. Based on this information, a generation AI summoned by the server creates an animation of a dragon flying through the sky and over mountains, accompanied by a grand background music. After checking the quality and appropriateness of the video, the user watches the finished video, is satisfied, and decides to share it with friends on social media. In this way, the present invention provides a form for users to intuitively and easily generate high-quality content.
[0484] The following describes the processing flow.
[0485] Step 1:
[0486] The user launches an application on the communication terminal and selects images or videos from the terminal to be used as material for video generation. After the user's selection, the communication terminal packets the selected media files and prepares to send them to the server.
[0487] Step 2:
[0488] The device sends the image or video data selected by the user to the server. This transmission usually occurs in real time via an internet connection.
[0489] Step 3:
[0490] Users select story templates and character models presented through the application's UI, thereby concretizing their concept for the type of video they want to create.
[0491] Step 4:
[0492] The device sends the user's setting selections, along with story and character information, to the server in a defined data format.
[0493] Step 5:
[0494] The server integrates image or video data received from the terminal along with story and character selection information. This integrated data forms the basis for generating new videos.
[0495] Step 6:
[0496] The server inputs this integrated data into the generating AI and instructs it to begin the video generation process.
[0497] Step 7:
[0498] The generative AI utilizes computer vision and natural language processing technologies to generate videos based on integrated data. Simultaneously, the AI checks for inappropriate content during the generation process.
[0499] Step 8:
[0500] The server receives the video sent from the generating AI and reviews its content for final confirmation. If necessary, it can send correction instructions back to the generating AI.
[0501] Step 9:
[0502] The server converts the verified video data into the appropriate format and prepares it for storage in a form accessible to the user.
[0503] Step 10:
[0504] The server notifies the user that the video has been generated and is ready for publication. This notification is sent to the device and reflected in the application.
[0505] Step 11:
[0506] When the device receives a notification from the server, it provides the user with a preview of the newly generated video. The user can then review this preview and choose whether to publish or share it.
[0507] (Example 1)
[0508] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0509] Conventional video generation systems have the challenge of not being able to quickly generate videos that meet users' creative and individual requests. Furthermore, insufficient filtering of inappropriate content can potentially cause offense to users. Therefore, there is a need for efficient and high-quality video generation.
[0510] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0511] In this invention, the server includes means for integrating received media material and selection information and transmitting instructions to a generating artificial intelligence; means for the generating artificial intelligence to generate a video based on the received information and to detect and filter inappropriate content; and means for performing a process to examine the video generated by the generating artificial intelligence. This makes it possible to quickly generate safe and high-quality videos for users and improve the usability of the product.
[0512] A "communication terminal" is a device that has the function of allowing users to acquire images and videos and transmit that information to a computing device.
[0513] A "computational device" is a device that processes received media materials and selection information and sends instructions to the generating artificial intelligence.
[0514] "Generative artificial intelligence" is an artificial intelligence technology that generates videos based on received information and has the ability to detect and filter inappropriate content.
[0515] "Media material" refers to image and video data acquired by users using communication devices.
[0516] "Story and characters" refers to the visual and contextual elements necessary for video creation, as selected by the user.
[0517] A "prompt message" refers to text input used to provide specific instructions or requests to a generative artificial intelligence.
[0518] "Computational vision" refers to the ability of artificial intelligence to analyze image data and understand it as visual information.
[0519] "Natural language processing technology" refers to the ability of artificial intelligence to understand and process human language.
[0520] The "data storage unit" refers to a storage device that stores the generated video data and provides access to various information exchange platforms.
[0521] An "information exchange platform" refers to a platform that allows users to share videos they have created and communicate with other users.
[0522] This invention is a system for generating and managing creative videos, primarily using a communication terminal, a computing device (server), and generative artificial intelligence.
[0523] The user first selects images and videos they have taken using a communication device. This communication device includes mobile information devices such as smartphones and tablets. Accordingly, an application on the communication device provides a user interface, allowing the user to visually select the materials.
[0524] Next, the user selects the video's story and characters through the application. This selection is done through pre-prepared templates, allowing for a variety of video formats depending on the user's choices. The selected materials and input information are transmitted to a computing device via a communication protocol. The computing device processes the received information and sends instructions to the generating artificial intelligence.
[0525] Generative artificial intelligence utilizes computational vision and natural language processing technologies to generate videos based on a given prompt. The prompt might be provided in the form of, for example, "Generate a video based on an adventure story, combining a flying dragon with a mountain landscape." The generative AI analyzes this prompt and automatically creates the corresponding visual elements and animations.
[0526] Once video generation is complete, the computing device checks for any inappropriate content. The system performs this process using frame-by-frame image recognition technology. After confirming that a safe and high-quality video has been generated, a notification is sent to the user's communication terminal.
[0527] Ultimately, users can preview the generated video on their communication device and, if satisfied with the content, share it on information exchange platforms such as social media or other platforms. This makes it easy and intuitive for users to create high-quality content and share it with many people.
[0528] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0529] Step 1:
[0530] The user launches the application on the communication terminal and selects images or videos. In this step, the communication terminal provides a touch interface to acquire images, allowing the user to select materials from their own library. The input is the images or videos selected by the user, and the output is these data being selected and ready to be sent to the next step.
[0531] Step 2:
[0532] The user selects the story and characters within the application. The input data consists of selections made from the provided template options. The communication terminal is designed to allow visual selection through the user interface. The output is story and character data reflecting the user's selections.
[0533] Step 3:
[0534] The communication terminal transmits the selected materials and configuration information to the computing device. The input consists of the image selected in step 1 and the story and character information selected in step 2. The output is the transfer of this data to the computing device via a secure communication protocol.
[0535] Step 4:
[0536] The server integrates the received data and invokes the generative artificial intelligence to issue instructions for video generation. The input consists of media materials and selection information sent from the terminal. The computing device processes these integrally, using the generative AI model to create prompts for video generation. The output is the sending of these prompts to the generative artificial intelligence.
[0537] Step 5:
[0538] The generative artificial intelligence generates videos based on prompt text and filters out inappropriate content. The input is prompt text received from a server, and data processing utilizes computational vision and natural language processing techniques. The output is the generated video, which is the final video version after filtering.
[0539] Step 6:
[0540] The server analyzes the generated video to check for any inappropriate elements. The input is video data generated by the generation artificial intelligence. The output is a verified and safe video. This video is then ready to be sent to the user. The server uses image recognition technology to perform this analysis process.
[0541] Step 7:
[0542] A notification of the generated video is sent to the user's communication device. The input is notification data from the server, and the output is the notification to the user and a previewable video display. The user watches the video using the playback function within the application.
[0543] Step 8:
[0544] Users share videos they enjoy on social media and other information exchange platforms. Input is the user's operation of the share button, and output is the posting of the video to the selected platform. The communication terminal provides the application with sharing options to support this.
[0545] (Application Example 1)
[0546] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0547] In today's digital society, there is a need for individual users to easily generate high-quality visual artifacts and effectively share them on public platforms such as social networks. However, conventional visual artifact generation systems struggle to intuitively generate personalized content tailored to the user's needs and themes, and they lack sufficient processes to verify the inappropriateness of the generated content. Therefore, the challenge lies in providing a system that is easy for users to use and delivers personalized, high-quality content.
[0548] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0549] In this invention, the server includes means for integrating received information material and selected information and transmitting commands to a generative intelligence; means for the generative intelligence to generate a visual composite based on the received information and detect inappropriate content; and means for verifying the generated visual composite. This makes it possible to securely generate personalized, high-quality visual composites based on themes and visual experiences selected by individual users, and to easily share them on social media.
[0550] A "communication device" is a device used to acquire and transmit information, and is primarily a terminal operated by the user.
[0551] A "central device" is a centralized management device for receiving and processing information, such as a cloud server.
[0552] A "visual medium" is a collection of visually recognizable information, such as images and videos.
[0553] "Information material" refers to the data that is transmitted from a communication device to a central device, and includes visual media.
[0554] "Story" refers to a series of storylines or themes used in video creation.
[0555] "Featured elements" refer to elements such as characters and objects used within the video.
[0556] "Selection information" refers to information about the story and elements that the user has chosen.
[0557] "Generative intelligence" refers to artificial intelligence that has the ability to automatically generate visual constructs based on received information.
[0558] A "visual synthesizer" is a visual content product, such as images or videos, generated by generative intelligence.
[0559] "Inappropriate content" refers to content that violates public order and morals or that may cause offense to users.
[0560] A "social network" is an online platform for users to publish or share content they have created.
[0561] An "information storage medium" is a system resource or device for storing data such as generated visual composites.
[0562] The invention will now be described in terms of embodiments. This application provides a system that automatically generates highly personalized videos that can be shared on social networks. The entire system consists of a communication device, a central device, and generative intelligence.
[0563] The communication device, or user terminal, allows the user to acquire visual media and select narrative elements. Specifically, smartphones and tablets are envisioned, which provide the functionality to select visual media from camera rolls, etc., and transmit them to the central device. This creates an intuitive and easy-to-use interface for material selection for the user.
[0564] The central server receives information materials and selected information transmitted from communication devices and integrates them. The integrated information is sent to generative intelligence, which uses artificial intelligence technology to analyze the content and generate new visual artifacts. Specifically, it uses a combination of image recognition technology such as the Google Cloud Vision API and natural language processing technology. Through this process, high-quality content is automatically generated that reflects themes and viewing experiences based on the materials.
[0565] Finally, the generated visual composite is reviewed on the server to ensure it does not contain any inappropriate content before being returned to the communication device. Users can then review and edit this visual composite through the application and share it as needed through social networks or other information sharing media.
[0566] For example, if the theme is landscape photos taken during a trip, and the user selects a story and fantasy characters related to "Summer Adventure," the generative AI will use this information to create a visually engaging adventure video. An example of a prompt input to the generative AI model would be, "Create a 3-minute animated video with a beach adventure theme."
[0567] The above describes the modes for carrying out the invention.
[0568] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0569] Step 1:
[0570] The device captures or selects visual media from the user. Images and videos are displayed on the screen through the device's interface, allowing the user to search and select materials by date of capture, tags, etc. The input is the visual media selected by the user, and the output is the data of the selection results.
[0571] Step 2:
[0572] The terminal transmits information about the visual medium selected by the user, along with the story and elements specified by the user, to the central device. Data communication is used for transmission. The input is the selection information of the story and elements set by the user, and the output is the integrated data transmitted to the central device.
[0573] Step 3:
[0574] The server receives visual media and configuration information transmitted from the terminal, integrates them, and prepares them for submission to the generative intelligence. The received data is formatted according to the specified format, and prompt statements are generated. The input is the visual media and selection information data from the terminal, and the output is the prompt statements and integrated data for the generative intelligence.
[0575] Step 4:
[0576] The generative intelligence generates visual artifacts based on prompt text and integrated data sent from the server. Here, natural language processing is used to analyze the narrative, and image recognition techniques are employed to analyze the visual elements. The input consists of prompt text and integrated data from the server, and the output is the generated visual artifact.
[0577] Step 5:
[0578] The server reviews the visual constructs generated by generative intelligence and checks for inappropriate content. An algorithm for detecting inappropriate elements is incorporated. The input is the generated visual construct, and the output is the reviewed visual construct. If there are no problems, it is approved and ready to be sent to the terminal.
[0579] Step 6:
[0580] The terminal receives the final visual composite sent from the server and provides the user with preview and editing functions. After reviewing the video, the user can add text and music as needed. The input is the visual composite from the server, and the output is the final content that the user views.
[0581] Step 7:
[0582] Users choose whether or not to share the generated content on social networks, and through this action, they share the visual composite with other users. Sharing is done through the application, and links or embed codes are generated. The input is the sharing instruction based on the user's action, and the output is the link information of the shared content.
[0583] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0584] This invention relates to a system that recognizes a user's emotions in real time and generates customized videos based on those emotions. Here, we describe a detailed embodiment of an innovative video generation system that combines an emotion engine.
[0585] First, the user operates a communication terminal and launches the video generation application. The user selects the image or video files to be used in the video they want to generate using the communication terminal and sends them to the server.
[0586] Next, the user selects from multiple templates via an interface on their communication terminal to set up the story and characters. The selected information, along with the results of the emotion engine's analysis, is sent to the server.
[0587] Here, the emotion engine comes into play, collecting facial expression data through the user's camera. This engine analyzes the user's facial feature points and can identify emotions in real time. The analyzed emotion data influences decisions about what kind of story development and character behavior to adopt in video production.
[0588] Next, the server integrates the received image or video data, story and character information, and real-time analyzed emotion data. Based on this integrated data, the server issues a command to the generating AI to create a video. The generating AI then generates a video optimized for the user's current emotional state based on the emotion analysis results.
[0589] The videos created by the generation AI are verified by the server and then sent to the communication terminal. Users can preview the generated videos on the communication terminal and adjust them as needed.
[0590] (Specific example)
[0591] For example, consider a scenario where a user wants to generate a video to relieve work stress. The user selects a story with a relaxation theme and gentle characters. The communication device inputs the user's smile and calm expressions into the emotion engine via the camera. After the emotion engine detects the user's relaxed state, the generation AI creates a video with calming music and a beautiful natural landscape as the background. Finally, this video is tailored specifically for the user and previewed on the communication device. The user can then watch this relaxing video and even share it with other users.
[0592] This invention enables users to generate and enjoy content adapted to their individual emotional states, thus allowing for a more personalized experience.
[0593] The following describes the processing flow.
[0594] Step 1:
[0595] The user launches a video generation app installed on their communication device. The app's UI is displayed, and the user selects the image or video saved for the video they want to generate.
[0596] Step 2:
[0597] The device sends the image or video data selected by the user to the server. At this point, the data is packetized and securely transferred to the server.
[0598] Step 3:
[0599] The user selects from story templates and character options provided using the communication terminal interface, thereby determining the concept of the video.
[0600] Step 4:
[0601] The device sends story and character selection information to the server. This data will be used for future video generation.
[0602] Step 5:
[0603] An emotion engine operates on the device and uses the user's camera to collect facial expression data. The emotion engine analyzes this data and recognizes emotions in real time.
[0604] Step 6:
[0605] The device sends the analyzed emotional data, along with story and character information, to the server. In this way, the user's emotional state is reflected in the video generation.
[0606] Step 7:
[0607] The server integrates all received data and sends instructions to the generating AI. Based on the integrated data, the generating AI creates a video adapted to the user's emotions.
[0608] Step 8:
[0609] The generation AI takes into account the results of emotion analysis and creates videos that match the specified story and characters. Music and special effects are added to the videos to enhance their quality.
[0610] Step 9:
[0611] The server receives the video sent from the generating AI and reviews its content. After reconfirming that there are no inappropriate elements, it retains the final version of the video.
[0612] Step 10:
[0613] The server sends the final version of the video to the user's device and notifies them that the generation is complete.
[0614] Step 11:
[0615] Users can preview the generated video on their device and review its content. They can make adjustments as needed, and if they are satisfied, they can save or share the video.
[0616] (Example 2)
[0617] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0618] Traditional video generation systems have a problem in automatically generating personalized content based on user emotions. In particular, the process of generating videos that reflect the user's real-time emotional state is complex and often requires manual adjustments. As a result, it has been difficult to quickly and efficiently obtain videos that satisfy users.
[0619] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0620] In this invention, the server includes means for integrating received media material, selection information, and emotion data; means for transmitting instructions to a generating artificial intelligence; and means for reviewing the generated video. This enables the automatic generation of videos optimized for the user's emotional state.
[0621] A "communication terminal" is a device used by a user to acquire images and videos, and to send and receive information with a server.
[0622] A "server" is a central processing unit that receives data from communication terminals, integrates emotional data and selection information, and manages instructions for the generating artificial intelligence.
[0623] "Generative artificial intelligence" is an algorithm that generates videos optimized for the user's emotional state based on information received from a server.
[0624] An "emotion analysis device" is a mechanism that analyzes the feature points of a user's face and identifies emotional data in real time.
[0625] A "prompt statement" is a command statement created by the server to instruct the AI model on the specifications for video generation.
[0626] A "video" is a dynamic visual medium generated based on user emotional data, and is content that appeals to both sight and hearing.
[0627] A "data storage device" is a storage device that stores generated video data and enables users to access and share it.
[0628] This invention is a system that recognizes a user's emotions in real time and generates a customized video based on those emotions. The following describes an embodiment of this system in detail.
[0629] The user launches a video generation application using a communication terminal. This terminal refers to a typical computer or mobile device equipped with a camera, which can acquire image and video files. Through the application, the user selects the media materials to be used for the video and sends them to the server.
[0630] This communication terminal is equipped with an emotion analysis device that collects user facial expression data. This device analyzes the user's facial feature points and has the function of identifying their emotional state in real time. The analysis results are transmitted from the communication terminal to a server.
[0631] The server integrates the received media materials, the story and character information selected by the user, and the sentiment analysis results. This integrated data is used as a prompt message to send instructions to the generation AI model. This prompt message is a set of commands that specify the video generation specifications and is written in a format that the generation AI model can understand.
[0632] The generative AI model generates videos appropriate to the user's emotional state based on the received prompt text. This AI model operates using machine learning algorithms combined with computer vision and natural language processing techniques. The generated videos are reviewed by a server and then sent to a communication terminal. On this terminal, the user can preview the video and adjust or modify it as needed.
[0633] As a concrete example, a user might want to generate a "relaxing video." The user selects a relaxation-themed template and uses the camera on their communication device to have their facial expressions detected by an emotion analysis device. The server sends a prompt message to the generation AI model: "Generate a video incorporating calming music and natural scenery." The generation AI model generates the video based on these instructions, and it is finally delivered to the communication device so that the user can view it.
[0634] In this way, this system can efficiently generate and deliver personalized video content that responds to the user's emotional state.
[0635] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0636] Step 1:
[0637] The user launches a video generation application using a communication terminal and selects image or video files within it. This serves as the input. The communication terminal acquires the selected media material as digital data and sends it to the server. Here, data format conversion and data transfer based on the communication protocol take place.
[0638] Step 2:
[0639] The user selects story and character templates using an interface on a communication terminal. The selected information becomes the input. The communication terminal sends this selection information to the server as digital data. In this step, the user's selected information is organized and packaged as data.
[0640] Step 3:
[0641] The emotion analysis device in the communication terminal captures the user's facial expressions with a camera and collects emotion data. The captured facial data becomes the input. This data is analyzed in real time, and the emotional state is identified based on the user's facial feature points. The analysis results are sent to the server as emotion data.
[0642] Step 4:
[0643] The server integrates the media material, selection information, and sentiment data received in steps 1 through 3. Based on this input data, it creates prompts for the generative AI model. The integrated data is then processed and formatted to create the video scenario.
[0644] Step 5:
[0645] The server sends a prompt to the generative AI model, issuing a command to generate a video. The generative AI model receives the prompt as input and generates a video optimized for the user's emotional state. Here, the generative AI utilizes computer vision and natural language processing technologies to perform data calculations and automatically construct the video storyline.
[0646] Step 6:
[0647] The server reviews the generated videos, verifying their content and performing quality checks. If inappropriate content is found, an automatic filter is activated. The reviewed videos are sent to the communication terminal. Additional data processing is performed for security and quality control.
[0648] Step 7:
[0649] The user previews the transmitted video on their communication device. The user reviews the video content and edits or adjusts it as needed. The user's feedback serves as input, and they have the option to regenerate the video. Finally, once a satisfactory video is completed, it can be shared with other users.
[0650] (Application Example 2)
[0651] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0652] In recent years, while a wide variety of video content has been available, it has been difficult to provide users with personalized content that is tailored to their emotional state. In this context, there is a need for a system that allows users to receive optimal video content in real time based on their emotions, thereby improving their experience.
[0653] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0654] In this invention, the server includes means for integrating received information materials and selection information and transmitting instructions to a generation processing device; means for the generation processing device to acquire the user's emotional state through emotion analysis and optimize its operation based on the emotional state; and means for the communication device to feed back the emotional state determined by emotion analysis to the user and present recommended videos that correspond to the user's emotions. This makes it possible to provide personalized video content that corresponds to the user's emotions.
[0655] "Communication equipment" is a general term for devices used by users to exchange information with the outside world.
[0656] "Information material" refers to data such as images and videos that users submit, and is processed on the server.
[0657] An "information processing device" refers to a system that integrates received data and transmits necessary instructions to a generation processing device.
[0658] A "generation processing device" refers to a device that generates and optimizes video content using emotion analysis technology based on received information.
[0659] "Emotional analysis" refers to a technology that analyzes a user's facial expressions and reactions in real time to identify their emotional state.
[0660] "Feedback" refers to a mechanism where the system returns the results and information it has analyzed to the user, influencing the user's choices and actions.
[0661] "Recommended videos" refer to video content provided to users that has been pre-selected based on their emotional state.
[0662] The system consists of a communication device, an information processing device, and a generation processing device. The communication device acquires images and videos from the user and transmits them to the information processing device as information material. The information processing device integrates the received information material and sends instructions to the generation processing device. The generation processing device uses emotion analysis technology to acquire the user's emotional state and optimizes its operation based on the acquired emotional data.
[0663] The system utilizes machine vision technology for emotion analysis. Specifically, it uses a smartphone camera and software such as OpenCV to analyze the user's emotions in real time from their facial expressions. It also uses services such as Microsoft Azure Face API to identify emotional states.
[0664] Based on these analysis results, the information processing device proposes personalized video content using a generative AI model and presents the video generated from the communication device to the user. At this time, the communication device can provide the user with feedback on the analyzed emotional state and present recommended videos based on the user's response.
[0665] For example, if a user is feeling stressed, the generation processing unit will identify the emotional state as "stressed" and instruct it to generate "footage of a forest with calming music in the background." In this way, by prompting the generation AI model with a message such as, "The user's current emotional state is relaxed. Please generate a video with calming music and a beautiful natural landscape in the background," it is possible to provide content that is appropriate for the user.
[0666] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0667] Step 1:
[0668] The communication terminal acquires user images and video data and transmits them to the information processing device as information material. The input is the user's images and video data, and the output is the transmission of that data to the information processing device. In this step, the communication terminal uses its camera function to collect images of the user.
[0669] Step 2:
[0670] The information processing device integrates the received information material with the user's selected story and character selection information. The input consists of the information material and selection information, and the output is this integrated data. This integrated data prepares the device for optimal video generation.
[0671] Step 3:
[0672] The information processing device transmits instructions to the generation processing device based on integrated data. The input is integrated data, and the output is the instructions transmitted to the generation processing device. The information processing device organizes the received information and provides it to the generation processing device in a processable format.
[0673] Step 4:
[0674] The generation and processing device analyzes the user's emotions using machine vision technology based on the received instructions. The input is instruction data, and the output is data on the user's emotional state. Here, for example, facial feature points are analyzed using OpenCV, and emotions are determined through the Microsoft Azure Face API.
[0675] Step 5:
[0676] The generation processing device generates videos using a generation AI model based on the emotion analysis results. The input is the emotion analysis results, and the output is video data optimized for the user's emotions. The device generates prompt text and instructs the generation AI model, such as "The user's current emotional state is relaxed. Please generate a video with calming music and a beautiful natural landscape as the background," to generate the video.
[0677] Step 6:
[0678] The generated video is sent to the information processing device and prepared for display to the user. The input is the generated video data, and the output is displayed on the communication terminal. This step also includes a final check to ensure the video does not contain any inappropriate content.
[0679] Step 7:
[0680] The communication terminal presents the generated video to the user and obtains user feedback. The input is emotional data for presenting the video to the user and obtaining feedback, while the output is user response data. This aims to improve the overall system responsiveness and user experience.
[0681] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0682] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0683] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0684] [Fourth Embodiment]
[0685] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0686] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0687] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0688] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0689] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0690] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0691] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0692] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0693] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0694] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0695] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0696] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0697] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0698] This invention is a system centered around a user's communication terminal and server, utilizing generation AI to generate and manage creative and free-form videos. The functions and procedures of each component are described in detail below.
[0699] First, the user uses a communication device such as a smartphone to select images or videos they have taken or saved. This communication device has the function of sending the selected images or videos to a server. In terms of configuration, an application on the device provides the user interface, allowing the user to intuitively select the material they want to turn into a video.
[0700] Next, the user selects a story and characters to serve as a template for video generation via an application on their device. These choices are necessary to form the video's structure, visual elements, and scenario, and various templates are provided in advance within the application. The user can combine these options to create their own unique settings according to their preferences.
[0701] The selected materials and settings are sent from the terminal to the server. The server receives and integrates this input data and instructs the generative AI to generate a new video. The generative AI uses computer vision and natural language processing techniques to automatically create a video depicting the specified story and characters. The system also has a function to automatically detect inappropriate content, ensuring that no elements that may offend the user are generated.
[0702] Once the video generation is complete, it undergoes a verification process on the server and is then notified to the user's communication device. Upon receiving the notification, the user can preview the generated video within the application. This video is presented to the user as comprehensively edited content, including visual information and music or sound effects. The user can then choose to share the video on other platforms such as social media, or publish it within the platform itself.
[0703] (Specific example)
[0704] Suppose a user takes a photo of a beautiful landscape while traveling. This user launches a video generation app and selects the landscape photo. Next, the user chooses the theme "Fantasy Adventure" as the story theme and selects a dragon as the character. Based on this information, a generation AI summoned by the server creates an animation of a dragon flying through the sky and over mountains, accompanied by a grand background music. After checking the quality and appropriateness of the video, the user watches the finished video, is satisfied, and decides to share it with friends on social media. In this way, the present invention provides a form for users to intuitively and easily generate high-quality content.
[0705] The following describes the processing flow.
[0706] Step 1:
[0707] The user launches an application on the communication terminal and selects images or videos from the terminal to be used as material for video generation. After the user's selection, the communication terminal packets the selected media files and prepares to send them to the server.
[0708] Step 2:
[0709] The device sends the image or video data selected by the user to the server. This transmission usually occurs in real time via an internet connection.
[0710] Step 3:
[0711] Users select story templates and character models presented through the application's UI, thereby concretizing their concept for the type of video they want to create.
[0712] Step 4:
[0713] The device sends the user's setting selections, along with story and character information, to the server in a defined data format.
[0714] Step 5:
[0715] The server integrates image or video data received from the terminal along with story and character selection information. This integrated data forms the basis for generating new videos.
[0716] Step 6:
[0717] The server inputs this integrated data into the generating AI and instructs it to begin the video generation process.
[0718] Step 7:
[0719] The generative AI utilizes computer vision and natural language processing technologies to generate videos based on integrated data. Simultaneously, the AI checks for inappropriate content during the generation process.
[0720] Step 8:
[0721] The server receives the video sent from the generating AI and reviews its content for final confirmation. If necessary, it can send correction instructions back to the generating AI.
[0722] Step 9:
[0723] The server converts the verified video data into the appropriate format and prepares it for storage in a form accessible to the user.
[0724] Step 10:
[0725] The server notifies the user that the video has been generated and is ready for publication. This notification is sent to the device and reflected in the application.
[0726] Step 11:
[0727] When the device receives a notification from the server, it provides the user with a preview of the newly generated video. The user can then review this preview and choose whether to publish or share it.
[0728] (Example 1)
[0729] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0730] Conventional video generation systems have the challenge of not being able to quickly generate videos that meet users' creative and individual requests. Furthermore, insufficient filtering of inappropriate content can potentially cause offense to users. Therefore, there is a need for efficient and high-quality video generation.
[0731] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0732] In this invention, the server includes means for integrating received media material and selection information and transmitting instructions to a generating artificial intelligence; means for the generating artificial intelligence to generate a video based on the received information and to detect and filter inappropriate content; and means for performing a process to examine the video generated by the generating artificial intelligence. This makes it possible to quickly generate safe and high-quality videos for users and improve the usability of the product.
[0733] A "communication terminal" is a device that has the function of allowing users to acquire images and videos and transmit that information to a computing device.
[0734] A "computational device" is a device that processes received media materials and selection information and sends instructions to the generating artificial intelligence.
[0735] "Generative artificial intelligence" is an artificial intelligence technology that generates videos based on received information and has the ability to detect and filter inappropriate content.
[0736] "Media material" refers to image and video data acquired by users using communication devices.
[0737] "Story and characters" refers to the visual and contextual elements necessary for video creation, as selected by the user.
[0738] A "prompt message" refers to text input used to provide specific instructions or requests to a generative artificial intelligence.
[0739] "Computational vision" refers to the ability of artificial intelligence to analyze image data and understand it as visual information.
[0740] "Natural language processing technology" refers to the ability of artificial intelligence to understand and process human language.
[0741] The "data storage unit" refers to a storage device that stores the generated video data and provides access to various information exchange platforms.
[0742] An "information exchange platform" refers to a platform that allows users to share videos they have created and communicate with other users.
[0743] This invention is a system for generating and managing creative videos, primarily using a communication terminal, a computing device (server), and generative artificial intelligence.
[0744] The user first selects images and videos they have taken using a communication device. This communication device includes mobile information devices such as smartphones and tablets. Accordingly, an application on the communication device provides a user interface, allowing the user to visually select the materials.
[0745] Next, the user selects the video's story and characters through the application. This selection is done through pre-prepared templates, allowing for a variety of video formats depending on the user's choices. The selected materials and input information are transmitted to a computing device via a communication protocol. The computing device processes the received information and sends instructions to the generating artificial intelligence.
[0746] Generative artificial intelligence utilizes computational vision and natural language processing technologies to generate videos based on a given prompt. The prompt might be provided in the form of, for example, "Generate a video based on an adventure story, combining a flying dragon with a mountain landscape." The generative AI analyzes this prompt and automatically creates the corresponding visual elements and animations.
[0747] Once video generation is complete, the computing device checks for any inappropriate content. The system performs this process using frame-by-frame image recognition technology. After confirming that a safe and high-quality video has been generated, a notification is sent to the user's communication terminal.
[0748] Ultimately, users can preview the generated video on their communication device and, if satisfied with the content, share it on information exchange platforms such as social media or other platforms. This makes it easy and intuitive for users to create high-quality content and share it with many people.
[0749] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0750] Step 1:
[0751] The user launches the application on the communication terminal and selects images or videos. In this step, the communication terminal provides a touch interface to acquire images, allowing the user to select materials from their own library. The input is the images or videos selected by the user, and the output is these data being selected and ready to be sent to the next step.
[0752] Step 2:
[0753] The user selects the story and characters within the application. The input data consists of selections made from the provided template options. The communication terminal is designed to allow visual selection through the user interface. The output is story and character data reflecting the user's selections.
[0754] Step 3:
[0755] The communication terminal transmits the selected materials and configuration information to the computing device. The input consists of the image selected in step 1 and the story and character information selected in step 2. The output is the transfer of this data to the computing device via a secure communication protocol.
[0756] Step 4:
[0757] The server integrates the received data and invokes the generative artificial intelligence to issue instructions for video generation. The input consists of media materials and selection information sent from the terminal. The computing device processes these integrally, using the generative AI model to create prompts for video generation. The output is the sending of these prompts to the generative artificial intelligence.
[0758] Step 5:
[0759] The generative artificial intelligence generates videos based on prompt text and filters out inappropriate content. The input is prompt text received from a server, and data processing utilizes computational vision and natural language processing techniques. The output is the generated video, which is the final video version after filtering.
[0760] Step 6:
[0761] The server analyzes the generated video to check for any inappropriate elements. The input is video data generated by the generation artificial intelligence. The output is a verified and safe video. This video is then ready to be sent to the user. The server uses image recognition technology to perform this analysis process.
[0762] Step 7:
[0763] A notification of the generated video is sent to the user's communication device. The input is notification data from the server, and the output is the notification to the user and a previewable video display. The user watches the video using the playback function within the application.
[0764] Step 8:
[0765] Users share videos they enjoy on social media and other information exchange platforms. Input is the user's operation of the share button, and output is the posting of the video to the selected platform. The communication terminal provides the application with sharing options to support this.
[0766] (Application Example 1)
[0767] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0768] In today's digital society, there is a need for individual users to easily generate high-quality visual artifacts and effectively share them on public platforms such as social networks. However, conventional visual artifact generation systems struggle to intuitively generate personalized content tailored to the user's needs and themes, and they lack sufficient processes to verify the inappropriateness of the generated content. Therefore, the challenge lies in providing a system that is easy for users to use and delivers personalized, high-quality content.
[0769] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0770] In this invention, the server includes means for integrating received information material and selected information and transmitting commands to a generative intelligence; means for the generative intelligence to generate a visual composite based on the received information and detect inappropriate content; and means for verifying the generated visual composite. This makes it possible to securely generate personalized, high-quality visual composites based on themes and visual experiences selected by individual users, and to easily share them on social media.
[0771] A "communication device" is a device used to acquire and transmit information, and is primarily a terminal operated by the user.
[0772] A "central device" is a centralized management device for receiving and processing information, such as a cloud server.
[0773] A "visual medium" is a collection of visually recognizable information, such as images and videos.
[0774] "Information material" refers to the data that is transmitted from a communication device to a central device, and includes visual media.
[0775] "Story" refers to a series of storylines or themes used in video creation.
[0776] "Featured elements" refer to elements such as characters and objects used within the video.
[0777] "Selection information" refers to information about the story and elements that the user has chosen.
[0778] "Generative intelligence" refers to artificial intelligence that has the ability to automatically generate visual constructs based on received information.
[0779] A "visual synthesizer" is a visual content product, such as images or videos, generated by generative intelligence.
[0780] "Inappropriate content" refers to content that violates public order and morals or that may cause offense to users.
[0781] A "social network" is an online platform for users to publish or share content they have created.
[0782] An "information storage medium" is a system resource or device for storing data such as generated visual composites.
[0783] The invention will now be described in terms of embodiments. This application provides a system that automatically generates highly personalized videos that can be shared on social networks. The entire system consists of a communication device, a central device, and generative intelligence.
[0784] The communication device, or user terminal, allows the user to acquire visual media and select narrative elements. Specifically, smartphones and tablets are envisioned, which provide the functionality to select visual media from camera rolls, etc., and transmit them to the central device. This creates an intuitive and easy-to-use interface for material selection for the user.
[0785] The central server receives information materials and selected information transmitted from communication devices and integrates them. The integrated information is sent to generative intelligence, which uses artificial intelligence technology to analyze the content and generate new visual artifacts. Specifically, it uses a combination of image recognition technology such as the Google Cloud Vision API and natural language processing technology. Through this process, high-quality content is automatically generated that reflects themes and viewing experiences based on the materials.
[0786] Finally, the generated visual composite is reviewed on the server to ensure it does not contain any inappropriate content before being returned to the communication device. Users can then review and edit this visual composite through the application and share it as needed through social networks or other information sharing media.
[0787] For example, if the theme is landscape photos taken during a trip, and the user selects a story and fantasy characters related to "Summer Adventure," the generative AI will use this information to create a visually engaging adventure video. An example of a prompt input to the generative AI model would be, "Create a 3-minute animated video with a beach adventure theme."
[0788] The above describes the modes for carrying out the invention.
[0789] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0790] Step 1:
[0791] The device captures or selects visual media from the user. Images and videos are displayed on the screen through the device's interface, allowing the user to search and select materials by date of capture, tags, etc. The input is the visual media selected by the user, and the output is the data of the selection results.
[0792] Step 2:
[0793] The terminal transmits information about the visual medium selected by the user, along with the story and elements specified by the user, to the central device. Data communication is used for transmission. The input is the selection information of the story and elements set by the user, and the output is the integrated data transmitted to the central device.
[0794] Step 3:
[0795] The server receives visual media and configuration information transmitted from the terminal, integrates them, and prepares them for submission to the generative intelligence. The received data is formatted according to the specified format, and prompt statements are generated. The input is the visual media and selection information data from the terminal, and the output is the prompt statements and integrated data for the generative intelligence.
[0796] Step 4:
[0797] The generative intelligence generates visual artifacts based on prompt text and integrated data sent from the server. Here, natural language processing is used to analyze the narrative, and image recognition techniques are employed to analyze the visual elements. The input consists of prompt text and integrated data from the server, and the output is the generated visual artifact.
[0798] Step 5:
[0799] The server reviews the visual constructs generated by generative intelligence and checks for inappropriate content. An algorithm for detecting inappropriate elements is incorporated. The input is the generated visual construct, and the output is the reviewed visual construct. If there are no problems, it is approved and ready to be sent to the terminal.
[0800] Step 6:
[0801] The terminal receives the final visual composite sent from the server and provides the user with preview and editing functions. After reviewing the video, the user can add text and music as needed. The input is the visual composite from the server, and the output is the final content that the user views.
[0802] Step 7:
[0803] Users choose whether or not to share the generated content on social networks, and through this action, they share the visual composite with other users. Sharing is done through the application, and links or embed codes are generated. The input is the sharing instruction based on the user's action, and the output is the link information of the shared content.
[0804] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0805] This invention relates to a system that recognizes a user's emotions in real time and generates customized videos based on those emotions. Here, we describe a detailed embodiment of an innovative video generation system that combines an emotion engine.
[0806] First, the user operates a communication terminal and launches the video generation application. The user selects the image or video files to be used in the video they want to generate using the communication terminal and sends them to the server.
[0807] Next, the user selects from multiple templates via an interface on their communication terminal to set up the story and characters. The selected information, along with the results of the emotion engine's analysis, is sent to the server.
[0808] Here, the emotion engine comes into play, collecting facial expression data through the user's camera. This engine analyzes the user's facial feature points and can identify emotions in real time. The analyzed emotion data influences decisions about what kind of story development and character behavior to adopt in video production.
[0809] Next, the server integrates the received image or video data, story and character information, and real-time analyzed emotion data. Based on this integrated data, the server issues a command to the generating AI to create a video. The generating AI then generates a video optimized for the user's current emotional state based on the emotion analysis results.
[0810] The videos created by the generation AI are verified by the server and then sent to the communication terminal. Users can preview the generated videos on the communication terminal and adjust them as needed.
[0811] (Specific example)
[0812] For example, consider a scenario where a user wants to generate a video to relieve work stress. The user selects a story with a relaxation theme and gentle characters. The communication device inputs the user's smile and calm expressions into the emotion engine via the camera. After the emotion engine detects the user's relaxed state, the generation AI creates a video with calming music and a beautiful natural landscape as the background. Finally, this video is tailored specifically for the user and previewed on the communication device. The user can then watch this relaxing video and even share it with other users.
[0813] This invention enables users to generate and enjoy content adapted to their individual emotional states, thus allowing for a more personalized experience.
[0814] The following describes the processing flow.
[0815] Step 1:
[0816] The user launches a video generation app installed on their communication device. The app's UI is displayed, and the user selects the image or video saved for the video they want to generate.
[0817] Step 2:
[0818] The device sends the image or video data selected by the user to the server. At this point, the data is packetized and securely transferred to the server.
[0819] Step 3:
[0820] The user selects from story templates and character options provided using the communication terminal interface, thereby determining the concept of the video.
[0821] Step 4:
[0822] The device sends story and character selection information to the server. This data will be used for future video generation.
[0823] Step 5:
[0824] An emotion engine operates on the device and uses the user's camera to collect facial expression data. The emotion engine analyzes this data and recognizes emotions in real time.
[0825] Step 6:
[0826] The device sends the analyzed emotional data, along with story and character information, to the server. In this way, the user's emotional state is reflected in the video generation.
[0827] Step 7:
[0828] The server integrates all received data and sends instructions to the generating AI. Based on the integrated data, the generating AI creates a video adapted to the user's emotions.
[0829] Step 8:
[0830] The generation AI takes into account the results of emotion analysis and creates videos that match the specified story and characters. Music and special effects are added to the videos to enhance their quality.
[0831] Step 9:
[0832] The server receives the video sent from the generating AI and reviews its content. After reconfirming that there are no inappropriate elements, it retains the final version of the video.
[0833] Step 10:
[0834] The server sends the final version of the video to the user's device and notifies them that the generation is complete.
[0835] Step 11:
[0836] Users can preview the generated video on their device and review its content. They can make adjustments as needed, and if they are satisfied, they can save or share the video.
[0837] (Example 2)
[0838] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0839] Traditional video generation systems have a problem in automatically generating personalized content based on user emotions. In particular, the process of generating videos that reflect the user's real-time emotional state is complex and often requires manual adjustments. As a result, it has been difficult to quickly and efficiently obtain videos that satisfy users.
[0840] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0841] In this invention, the server includes means for integrating received media material, selection information, and emotion data; means for transmitting instructions to a generating artificial intelligence; and means for reviewing the generated video. This enables the automatic generation of videos optimized for the user's emotional state.
[0842] A "communication terminal" is a device used by a user to acquire images and videos, and to send and receive information with a server.
[0843] A "server" is a central processing unit that receives data from communication terminals, integrates emotional data and selection information, and manages instructions for the generating artificial intelligence.
[0844] "Generative artificial intelligence" is an algorithm that generates videos optimized for the user's emotional state based on information received from a server.
[0845] An "emotion analysis device" is a mechanism that analyzes the feature points of a user's face and identifies emotional data in real time.
[0846] A "prompt statement" is a command statement created by the server to instruct the AI model on the specifications for video generation.
[0847] A "video" is a dynamic visual medium generated based on user emotional data, and is content that appeals to both sight and hearing.
[0848] A "data storage device" is a storage device that stores generated video data and enables users to access and share it.
[0849] This invention is a system that recognizes a user's emotions in real time and generates a customized video based on those emotions. The following describes an embodiment of this system in detail.
[0850] The user launches a video generation application using a communication terminal. This terminal refers to a typical computer or mobile device equipped with a camera, which can acquire image and video files. Through the application, the user selects the media materials to be used for the video and sends them to the server.
[0851] This communication terminal is equipped with an emotion analysis device that collects user facial expression data. This device analyzes the user's facial feature points and has the function of identifying their emotional state in real time. The analysis results are transmitted from the communication terminal to a server.
[0852] The server integrates the received media materials, the story and character information selected by the user, and the sentiment analysis results. This integrated data is used as a prompt message to send instructions to the generation AI model. This prompt message is a set of commands that specify the video generation specifications and is written in a format that the generation AI model can understand.
[0853] The generative AI model generates videos appropriate to the user's emotional state based on the received prompt text. This AI model operates using machine learning algorithms combined with computer vision and natural language processing techniques. The generated videos are reviewed by a server and then sent to a communication terminal. On this terminal, the user can preview the video and adjust or modify it as needed.
[0854] As a concrete example, a user might want to generate a "relaxing video." The user selects a relaxation-themed template and uses the camera on their communication device to have their facial expressions detected by an emotion analysis device. The server sends a prompt message to the generation AI model: "Generate a video incorporating calming music and natural scenery." The generation AI model generates the video based on these instructions, and it is finally delivered to the communication device so that the user can view it.
[0855] In this way, this system can efficiently generate and deliver personalized video content that responds to the user's emotional state.
[0856] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0857] Step 1:
[0858] The user launches a video generation application using a communication terminal and selects image or video files within it. This serves as the input. The communication terminal acquires the selected media material as digital data and sends it to the server. Here, data format conversion and data transfer based on the communication protocol take place.
[0859] Step 2:
[0860] The user selects story and character templates using an interface on a communication terminal. The selected information becomes the input. The communication terminal sends this selection information to the server as digital data. In this step, the user's selected information is organized and packaged as data.
[0861] Step 3:
[0862] The emotion analysis device in the communication terminal captures the user's facial expressions with a camera and collects emotion data. The captured facial data becomes the input. This data is analyzed in real time, and the emotional state is identified based on the user's facial feature points. The analysis results are sent to the server as emotion data.
[0863] Step 4:
[0864] The server integrates the media material, selection information, and sentiment data received in steps 1 through 3. Based on this input data, it creates prompts for the generative AI model. The integrated data is then processed and formatted to create the video scenario.
[0865] Step 5:
[0866] The server sends a prompt to the generative AI model, issuing a command to generate a video. The generative AI model receives the prompt as input and generates a video optimized for the user's emotional state. Here, the generative AI utilizes computer vision and natural language processing technologies to perform data calculations and automatically construct the video storyline.
[0867] Step 6:
[0868] The server reviews the generated videos, verifying their content and performing quality checks. If inappropriate content is found, an automatic filter is activated. The reviewed videos are sent to the communication terminal. Additional data processing is performed for security and quality control.
[0869] Step 7:
[0870] The user previews the transmitted video on their communication device. The user reviews the video content and edits or adjusts it as needed. The user's feedback serves as input, and they have the option to regenerate the video. Finally, once a satisfactory video is completed, it can be shared with other users.
[0871] (Application Example 2)
[0872] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0873] In recent years, while a wide variety of video content has been available, it has been difficult to provide users with personalized content that is tailored to their emotional state. In this context, there is a need for a system that allows users to receive optimal video content in real time based on their emotions, thereby improving their experience.
[0874] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0875] In this invention, the server includes means for integrating received information materials and selection information and transmitting instructions to a generation processing device; means for the generation processing device to acquire the user's emotional state through emotion analysis and optimize its operation based on the emotional state; and means for the communication device to feed back the emotional state determined by emotion analysis to the user and present recommended videos that correspond to the user's emotions. This makes it possible to provide personalized video content that corresponds to the user's emotions.
[0876] "Communication equipment" is a general term for devices used by users to exchange information with the outside world.
[0877] "Information material" refers to data such as images and videos that users submit, and is processed on the server.
[0878] An "information processing device" refers to a system that integrates received data and transmits necessary instructions to a generation processing device.
[0879] A "generation processing device" refers to a device that generates and optimizes video content using emotion analysis technology based on received information.
[0880] "Emotional analysis" refers to a technology that analyzes a user's facial expressions and reactions in real time to identify their emotional state.
[0881] "Feedback" refers to a mechanism where the system returns the results and information it has analyzed to the user, influencing the user's choices and actions.
[0882] "Recommended videos" refer to video content provided to users that has been pre-selected based on their emotional state.
[0883] The system consists of a communication device, an information processing device, and a generation processing device. The communication device acquires images and videos from the user and transmits them to the information processing device as information material. The information processing device integrates the received information material and sends instructions to the generation processing device. The generation processing device uses emotion analysis technology to acquire the user's emotional state and optimizes its operation based on the acquired emotional data.
[0884] The system utilizes machine vision technology for emotion analysis. Specifically, it uses a smartphone camera and software such as OpenCV to analyze the user's emotions in real time from their facial expressions. It also uses services such as Microsoft Azure Face API to identify emotional states.
[0885] Based on these analysis results, the information processing device proposes personalized video content using a generative AI model and presents the video generated from the communication device to the user. At this time, the communication device can provide the user with feedback on the analyzed emotional state and present recommended videos based on the user's response.
[0886] For example, if a user is feeling stressed, the generation processing unit will identify the emotional state as "stressed" and instruct it to generate "footage of a forest with calming music in the background." In this way, by prompting the generation AI model with a message such as, "The user's current emotional state is relaxed. Please generate a video with calming music and a beautiful natural landscape in the background," it is possible to provide content that is appropriate for the user.
[0887] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0888] Step 1:
[0889] The communication terminal acquires user images and video data and transmits them to the information processing device as information material. The input is the user's images and video data, and the output is the transmission of that data to the information processing device. In this step, the communication terminal uses its camera function to collect images of the user.
[0890] Step 2:
[0891] The information processing device integrates the received information material with the user's selected story and character selection information. The input consists of the information material and selection information, and the output is this integrated data. This integrated data prepares the device for optimal video generation.
[0892] Step 3:
[0893] The information processing device transmits instructions to the generation processing device based on integrated data. The input is integrated data, and the output is the instructions transmitted to the generation processing device. The information processing device organizes the received information and provides it to the generation processing device in a processable format.
[0894] Step 4:
[0895] The generation and processing device analyzes the user's emotions using machine vision technology based on the received instructions. The input is instruction data, and the output is data on the user's emotional state. Here, for example, facial feature points are analyzed using OpenCV, and emotions are determined through the Microsoft Azure Face API.
[0896] Step 5:
[0897] The generation processing device generates videos using a generation AI model based on the emotion analysis results. The input is the emotion analysis results, and the output is video data optimized for the user's emotions. The device generates prompt text and instructs the generation AI model, such as "The user's current emotional state is relaxed. Please generate a video with calming music and a beautiful natural landscape as the background," to generate the video.
[0898] Step 6:
[0899] The generated video is sent to the information processing device and prepared for display to the user. The input is the generated video data, and the output is displayed on the communication terminal. This step also includes a final check to ensure the video does not contain any inappropriate content.
[0900] Step 7:
[0901] The communication terminal presents the generated video to the user and obtains user feedback. The input is emotional data for presenting the video to the user and obtaining feedback, while the output is user response data. This aims to improve the overall system responsiveness and user experience.
[0902] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0903] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0904] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0905] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0906] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0907] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0908] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0909] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0910] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0911] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0912] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0913] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0914] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0915] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0916] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0917] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0918] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0919] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0920] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0921] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0922] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0923] The following is further disclosed regarding the embodiments described above.
[0924] (Claim 1)
[0925] A means for acquiring images or videos on a communication terminal and transmitting them to a server as media material,
[0926] The aforementioned communication terminal includes means for receiving information on the user's selection of story and characters, and transmitting the selection information to the server,
[0927] The server includes means for integrating the received media materials and selection information and transmitting instructions to the generating artificial intelligence,
[0928] The aforementioned artificial intelligence generation system includes means for generating a video based on received information and detecting inappropriate content,
[0929] The server includes means for reviewing the video generated by the artificial intelligence,
[0930] A system including means for presenting a generated video to a user in the aforementioned communication terminal.
[0931] (Claim 2)
[0932] The system according to claim 1, characterized in that the generating artificial intelligence uses computer vision and natural language processing technologies.
[0933] (Claim 3)
[0934] The system according to claim 1, characterized in that the server stores the generated video data in a database and enables users to share it.
[0935] "Example 1"
[0936] (Claim 1)
[0937] A means for acquiring an image or video in a communication terminal and transmitting it to a computing device as media material,
[0938] The aforementioned communication terminal includes means for receiving information on the user's selection of story and characters, and transmitting the selection information to the computing device.
[0939] The aforementioned computing device includes means for integrating received media materials and selection information and transmitting instructions to the generating artificial intelligence,
[0940] The aforementioned artificial intelligence generation system includes means for generating videos based on received information and for detecting and filtering inappropriate content,
[0941] The computing device includes means for performing a process of investigating the video generated by the artificial intelligence,
[0942] A system including means for presenting a generated video to a user and enabling sharing operations in the aforementioned communication terminal.
[0943] (Claim 2)
[0944] The system according to claim 1, characterized in that the generating artificial intelligence uses computational vision and natural language processing techniques to generate a video based on a prompt sentence.
[0945] (Claim 3)
[0946] The system according to claim 1, characterized in that the computing device stores the generated video data in a data storage unit and facilitates sharing to various information exchange platforms.
[0947] "Application Example 1"
[0948] (Claim 1)
[0949] In a communication device, means for acquiring a visual medium and transmitting it to a central device as information material,
[0950] The communication device includes means for receiving information on the selection of story and elements by the user, and transmitting the selection information to the central device.
[0951] The central device includes means for integrating received information materials and selected information and transmitting commands to the generative intelligence,
[0952] The aforementioned generative intelligence includes means for generating a visual composite based on received information and detecting inappropriate content,
[0953] The central device includes means for confirming the visual composite generated by the generative intelligence,
[0954] The aforementioned communication device includes means for presenting a generated visual composite to the user,
[0955] A means of adding a viewing experience to a visual composite based on a theme selected by the user,
[0956] A means of providing a function to share generated visual composites on social networks.
[0957] A system that includes this.
[0958] (Claim 2)
[0959] The system according to claim 1, characterized in that the generative intelligence uses image recognition technology and natural language processing technology.
[0960] (Claim 3)
[0961] The system according to claim 1, characterized in that the central device stores the generated visual composite in an information storage medium and enables public operation by the user.
[0962] "Example 2 of combining an emotion engine"
[0963] (Claim 1)
[0964] A means for acquiring images or videos on a communication terminal and transmitting them to a server as media material,
[0965] The aforementioned communication terminal includes means for receiving information on the user's selection of story and characters, and transmitting the selection information to the server,
[0966] A means for collecting user emotion data using the emotion analysis device of the communication terminal and transmitting the emotion data to a server,
[0967] The server includes means for integrating received media materials, selection information, and emotion data, and transmitting instructions to a generating artificial intelligence.
[0968] The aforementioned artificial intelligence generation system includes means for generating videos optimized for the user's emotional state based on received information and for detecting inappropriate content,
[0969] The server includes means for reviewing the video generated by the artificial intelligence,
[0970] The aforementioned communication terminal includes means for presenting a generated video to the user and allowing the user to adjust the video as needed.
[0971] A system that includes means to enable sharing videos with other users.
[0972] (Claim 2)
[0973] The system according to claim 1, characterized in that the generating artificial intelligence generates videos based on the results of user emotion analysis, and uses computer vision and natural language processing technologies.
[0974] (Claim 3)
[0975] The system according to claim 1, characterized in that the server stores the generated video data in a data storage device and enables users to share the data.
[0976] "Application example 2 when combining with an emotional engine"
[0977] (Claim 1)
[0978] In a communication device, a means for acquiring video data and transmitting it to an information processing device as information material,
[0979] The aforementioned communication device includes means for receiving information on the selection of stories and characters by the user and transmitting the selection information to the information processing device,
[0980] The information processing device includes means for integrating received information material and selection information and transmitting instructions to a generation processing device,
[0981] The generation processing apparatus includes means for generating an operation based on the received information and detecting inappropriate content,
[0982] The information processing device includes means for confirming the operation generated by the generation processing device,
[0983] The aforementioned communication device includes means for presenting the generated actions to the user,
[0984] The aforementioned generation processing device includes means for acquiring the user's emotional state through emotion analysis and optimizing its operation based on the emotional state,
[0985] The aforementioned communication device includes means for providing feedback to the user on the emotional state determined by emotion analysis and for presenting recommended videos that correspond to the user's emotions.
[0986] (Claim 2)
[0987] The system according to claim 1, characterized in that the generation processing device uses machine vision and language processing technology.
[0988] (Claim 3)
[0989] The system according to claim 1, characterized in that the information processing device stores the generated operation data in an information storage device and enables shared operation by users. [Explanation of Symbols]
[0990] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring images or videos on a communication terminal and transmitting them to a server as media material, The aforementioned communication terminal includes means for receiving information on the user's selection of story and characters, and transmitting the selection information to the server, The server includes means for integrating the received media materials and selection information and transmitting instructions to the generating artificial intelligence, The aforementioned artificial intelligence generation system includes means for generating a video based on received information and detecting inappropriate content, The server includes means for reviewing the video generated by the artificial intelligence, A system including means for presenting a generated video to a user in the aforementioned communication terminal.
2. The system according to claim 1, characterized in that the generating artificial intelligence uses computer vision and natural language processing technologies.
3. The system according to claim 1, characterized in that the server stores the generated video data in a database and enables users to share it.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A