System
The system simplifies high-quality video creation by allowing users to input settings and using AI to integrate character models, movements, and music, addressing the need for specialized knowledge in video production.
Patent Information
- Application Number
- JP2024133456
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Creating high-quality videos requires specialized knowledge and time, making it difficult for users without such expertise to produce professional-quality videos.
A system that allows users to input character settings, plot information, and select sound effects and background music, using physics simulation and AI to automatically generate and integrate these elements into a video.
Enables users to easily create professional-quality videos by configuring detailed settings, with the system automatically integrating character models, movements, sound effects, and background music.
Smart Images

Figure 2026030473000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Video production requires many steps, such as creating a plot, determining the length, deciding on the composition, setting the movements, and adding sound effects and background music, and completing these processes requires specialized knowledge and time. This creates a problem in that users without specialized knowledge cannot easily create high-quality videos. The present invention aims to solve this problem by providing a system that allows users to easily create high-quality videos. [Means for solving the problem]
[0005] The present invention provides a system including a means for receiving character settings specified by a user, a means for generating a character model based on the character settings, a means for receiving plot information specified by a user, a means for generating character movements based on the plot information, a means for selecting sound effects based on the generated movements and environment, a means for selecting background music based on the plot information, and a means for generating a video by integrating the character model, movements, sound effects, and background music. The above-mentioned problems are also solved by providing a system including a means for selecting the movements and sound effects of the generated character model based on a physical simulation, and a system including a means for transmitting the generated video to a user terminal.
[0006] "Character settings" refers to the act of a user inputting characteristic information such as the personality and appearance of a character that will appear in a video, or the information itself.
[0007] "Character Model" refers to a digital representation of a three-dimensional character generated based on a character profile.
[0008] "Plot information" refers to information including details of the scenario and scenes of a video specified by a user.
[0009] "Action" refers to the animation that a character model performs based on plot information.
[0010] "Sound effects" refers to sound effects that are set to match specific actions or environments.
[0011] "Background music" refers to music used to create the atmosphere of a video.
[0012] "Physics simulation" refers to the technology of calculating the behavior and sound effects of characters and environments based on the laws of physics in the real world.
[0013] An "action sequence" refers to a series of specific actions that a character takes to execute a series of movements.
[0014] "Content Generation Engine" means a software engine for generating and integrating characters, actions, sound effects, background music, etc. based on user input.
[0015] The term "system" refers to a combination of a set of hardware and software including all means and elements described in the present invention. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention is a system that allows users to easily generate high-quality videos, and provides a means for automatically integrating user-specified character settings, plot information, sound effects, and background music and outputting them as a single video.
[0038] Program Overview
[0039] Character setting input
[0040] Users can use their devices to enter detailed settings such as the character's personality, appearance, clothing, etc., through an interface, allowing them to specifically design their own character.
[0041] Character Model Generation
[0042] The server receives the character setting information sent from the device, analyzes it, and generates a character model, a three-dimensional digital representation of the character with an appearance and personality based on the setting.
[0043] Plot information input
[0044] The user inputs the details of the scenes (plot information) required for the video via the terminal. For example, they can specifically describe the scenario, such as "a character sitting on a bench in a park."
[0045] motion generation
[0046] The server analyzes the plot information and automatically generates character movements, including setting animations and movement paths so that the characters move naturally.
[0047] Sound effect selection
[0048] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[0049] Background music selection
[0050] Based on the plot information, the server selects background music to create an appropriate atmosphere for the scene, for example, relaxing music for a park scene.
[0051] Video Integration and Generation
[0052] The server generates the video by integrating character models, movements, sound effects, and background music, and then synchronizes all the elements to create a single, high-quality video.
[0053] Specific examples
[0054] Character Settings
[0055] The user sets up a character with a bright personality, red hair, and blue clothes on the device screen. This information is sent from the device to the server.
[0056] Character model generation
[0057] Based on the user's settings, the server generates a character model with red hair and blue clothes. This character has a cheerful personality, which is reflected in their facial expressions and movements.
[0058] Behavior Settings
[0059] When a user inputs plot information such as "a character sits on a bench in a park," the server receives this information and automatically generates a series of actions for the character to walk and sit on the bench.
[0060] Sound effects and background music selection
[0061] Using physics simulation, the server selects sound effects that match the park's ambient sounds and the characters' movements, as well as background music that creates a relaxing atmosphere.
[0062] Video Integration and Generation
[0063] The server then combines the generated character models, movements, sound effects, and background music into a single video file, which users can then download or stream to their devices.
[0064] In this way, the system of the present invention allows the user to automatically generate professional quality videos simply by configuring detailed settings.
[0065] The processing flow will be explained below.
[0066] Step 1:
[0067] The user opens the character setting screen on the device and enters detailed settings for the character, such as personality, appearance, clothing, etc. Once the input is complete, the setting information is sent to the server.
[0068] Step 2:
[0069] The server receives the character setting information sent from the device. The server analyzes this information and generates a character model with the set characteristics. The generated character model is expressed as a digital 3D model.
[0070] Step 3:
[0071] The user inputs the video scenario and scene details (plot information) on the device. For example, they specify specific locations and actions, such as "a character sitting on a bench in a park." Once input is complete, the plot information is sent to the server.
[0072] Step 4:
[0073] The server receives the plot information sent from the device. The server analyzes the plot information and generates movement animations to determine how the character should move in the specified scenario or scene. For example, it calculates and sets the movement of a character walking and sitting on a park bench.
[0074] Step 5:
[0075] After the movement animation is determined, the server selects sound effects appropriate for that movement and the surrounding environment, using physics simulation to determine, for example, the sound of a character walking or the ambient sounds of a park (birds chirping, wind noise, etc.).
[0076] Step 6:
[0077] The server selects background music (BGM) that matches the scene based on the plot information. The server selects music from a database that matches the atmosphere of the plot. For example, it selects calming BGM to match a relaxing scene in a park.
[0078] Step 7:
[0079] The server combines character models, movement animations, sound effects, and background music into a single video, precisely timing and matching each element to produce a high-quality video.
[0080] Step 8:
[0081] The server sends the generated video to the user's device, where the user can download and watch the video.
[0082] Step 9:
[0083] If a user wishes to make corrections to a video, they input the corrections from their device and send them back to the server. The server then regenerates the character model and movement animations based on the corrections, selects new sound effects and background music, and regenerates the corrected video.
[0084] Example 1
[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0086] With conventional video generation systems, it was difficult for users to generate high-quality videos using the character settings and plot information they desired. Furthermore, there was a lack of means to properly integrate character movements, sound effects, and background music, which often resulted in a decline in the quality of the generated videos. There is a need for a system that can solve these problems and enable easier, high-quality video generation.
[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0088] In this invention, the server includes means for receiving character settings specified by a user, means for generating a character model based on the character settings, means for receiving plot information specified by a user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for generating a video by integrating the character model, movements, sound effects, and background music, means for selecting environmental sounds and movement sounds using a physical simulation based on the generated character model and movements, means for transmitting the generated video to a user terminal, means for synchronizing the generated character model, movements, sound effects, and background music using a timeline editor, and means for automatically generating a character model and video using a generative AI model based on the character settings and plot information input by a user. This allows users to easily generate high-quality videos and enables the realization of well-integrated, professional-quality videos.
[0089] "Character configuration" is the process by which a user inputs details about a character, such as personality, appearance, and clothing.
[0090] A "character model" is a three-dimensional digital character generated based on a character setting.
[0091] "Plot information" is detailed information about the video scenario specified by the user.
[0092] "Character behavior" refers to a series of actions that a character performs based on plot information.
[0093] "Sound effects" are audio data used to add realistic sound effects to character movements and environments.
[0094] "Background music" is music selected to create the atmosphere of a scene based on plot information.
[0095] "Method of generating animation" is the process of integrating character models, movements, sound effects, and background music into a single format.
[0096] "Physics simulation" is a simulation technique for providing realistic sound effects based on generated motions and environments.
[0097] A "timeline editor" is a tool used to properly synchronize character models, movements, sound effects, and background music.
[0098] A "generative AI model" is an algorithm that uses AI technology to automatically generate character models and videos.
[0099] "User terminal" refers to the device through which a user inputs character settings and plot information.
[0100] This invention is a system that allows users to easily generate high-quality videos. It provides a means to automatically integrate user-specified character settings, plot information, sound effects, and background music and output them as a single video. The system consists of a server and a user terminal.
[0101] First, the user uses the device to input detailed settings for the character, such as personality, appearance, and clothing. Once the character settings are complete, the device sends this information to the server. For example, if the user sets a character with a cheerful personality, red hair, and blue clothing, this information is sent to the server.
[0102] The server uses a generative AI model based on the received character setting information to generate a three-dimensional digital character model. This character model reflects the character's set appearance and personality. Next, the user inputs the details of the scenes (plot information) required for the video via their device. For example, they input a scenario such as "the character sits on a bench in a park." This information is also sent from the device to the server.
[0103] The server analyzes the plot information and automatically generates character movements, including animations and movement paths that allow the character to move naturally. The server then selects sound effects based on the generated movements and the environment. Physics simulation is used to determine audio data that realistically reflects the character's movements and environmental sounds.
[0104] The server then selects background music based on the plot information. To create a suitable atmosphere for the scene, relaxing music is selected for the park scene. After these elements are gathered, the server generates a video by integrating character models, movements, sound effects, and background music. A timeline editor is used to properly align the timing of all elements to create a high-quality video.
[0105] The generated video is sent from the server to the user's device, where the user can watch it by downloading or streaming. This system allows users to automatically generate professional-quality videos by simply configuring detailed settings.
[0106] Specific examples
[0107] The user configures the character settings on the device screen. For example, they can configure a character with a cheerful personality, red hair, and blue clothes, and send this information from the device to the server. The server generates a character model with red hair and blue clothes based on the configuration information received from the user. This character has a cheerful personality, which is reflected in their facial expressions and movements. Next, the user inputs plot information such as "the character sits on a bench in a park." Based on this information, the server automatically generates a series of actions for the character to walk and sit on the bench. For this action, physical simulation is used to select sound effects that match the ambient sounds of the park and the character's movements. Background music is also selected to create a relaxing atmosphere.
[0108] Once all these elements are in place, the server combines the character movements, sound effects, and background music into a video that users can then download or stream to their device.
[0109] Prompt Sentence Examples
[0110] Character: A cheerful character with red hair and blue clothes.
[0111] Plot Info: A character sits on a bench in a park.
[0112] Sound Effects: Added ambient bird sounds and character walking sounds
[0113] Background music: Relaxing music
[0114] When this prompt is input into a generative AI model, a video that integrates all the elements is automatically generated.
[0115] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0116] Step 1:
[0117] The user inputs character setting information using the terminal. In the input form, the user sets details such as the character's personality, appearance, and clothing. For example, the user can set a "character with a cheerful personality, red hair, and blue clothes." Once this information is entered, the terminal sends the setting information to the server.
[0118] Input: Detailed setting information such as character personality, appearance, clothing, etc.
[0119] Output: Character configuration information sent to the server.
[0120] Step 2:
[0121] The server analyzes the character setting information received from the device, and the generative AI model on the server generates a three-dimensional digital character model based on this information.
[0122] Input: Character configuration information.
[0123] Output: 3D digital character model.
[0124] What it does: The generative AI model references the database and sets the character's hair color to red and their clothing to blue.
[0125] Step 3:
[0126] The user inputs the plot information required for the video via the terminal. In the input form, the user writes down the details of the scene. For example, the user enters a scenario such as "A character sits on a bench in a park." Once this information is entered, the terminal sends the plot information to the server.
[0127] Input: Scene details (plot information).
[0128] Output: Plot information sent to the server.
[0129] Step 4:
[0130] The server analyzes the plot information and automatically generates character movements, including animations and movement paths that allow the characters to move naturally.
[0131] Input: Plot information.
[0132] Output: Character animation data and movement path.
[0133] Specific Action: Based on the instruction "sit on a bench in the park," the plot analysis module generates the action of the character walking and sitting on the bench.
[0134] Step 5:
[0135] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[0136] Input: Character animation data and environment information.
[0137] Output: Suitable sound effects.
[0138] Specific operation: The sound effect selection module selects sounds such as birds singing, footsteps, and bench sounds from the sound library.
[0139] Step 6:
[0140] The server selects background music based on plot information, determining the background music to create an atmosphere that matches the scene.
[0141] Input: Plot information.
[0142] Output: Suitable background music.
[0143] Specific operation: The music selection module selects relaxing music from the music database.
[0144] Step 7:
[0145] The server uses a timeline editor to combine character models, movements, sound effects, and background music into a high-quality video, aligning all elements appropriately to create the final video.
[0146] Input: Character models, animation data, sound effects, background music.
[0147] Output: The finished video file.
[0148] Specific operation: Each element is integrated in the timeline editor, then encoded and generated into a single video.
[0149] Step 8:
[0150] The server sends the generated video to the user's device, where the user can watch the video by downloading or streaming.
[0151] Input: Your finished video file.
[0152] Output: The video sent to the user's device.
[0153] Specific operation: The server compresses the video file and sends it to the user's device. The user then downloads the video and watches it.
[0154] (Application example 1)
[0155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0156] Conventional video generation systems require a lot of manual work from the user and specialized knowledge and skills, making it difficult for anyone to easily generate high-quality videos and share them on content distribution services. Furthermore, integrating character settings, plot information, sound effects, and background music during video generation is complicated, making it difficult to maintain professional quality.
[0157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0158] In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for generating a video by integrating the character model, movements, sound effects, and background music, and means for uploading the generated video to a content distribution service. This allows users to easily generate high-quality videos and share them on the content distribution service.
[0159] "Character settings" are detailed setting information such as the character's personality, appearance, clothing, etc., specified by the user.
[0160] A "character model" is a three-dimensional digital representation generated based on a character's settings.
[0161] "Plot information" is detailed information about the scenes and scenarios required for the video specified by the user.
[0162] "Movement generation" is the process of automatically generating character movements based on plot information.
[0163] "Sound effects" are sound effects that are selected based on the action and environment being generated.
[0164] "Background music" is music that is selected to fit a scene based on plot information.
[0165] "Video generation" is the process of integrating character models, movements, sound effects, and background music to generate a single, high-quality video.
[0166] A "content distribution service" is a platform for sharing and viewing generated videos over the Internet.
[0167] A "user terminal" is an electronic device operated by a user, such as a smartphone or computer.
[0168] "Physics simulation" is a simulation technology that realistically reflects the movements and sound effects of generated character models.
[0169] "Smartphone Application" means a smartphone application that enables users to create, distribute, and guide their own video content.
[0170] The present invention provides a system that enables users to easily generate high-quality videos and share them via content distribution services. Specific embodiments of this system are described below.
[0171] Character setting input
[0172] The device provides users with a means to input detailed settings such as the character's personality, appearance, clothing, etc., allowing them to specifically design their own character.
[0173] Character Model Generation
[0174] The server receives the character setting information sent from the device, analyzes it, and generates a character model, a three-dimensional digital representation with an appearance and personality based on the setting.
[0175] Plot information input
[0176] Users are provided with a means to input details of scenes (plot information) required for the video via their device. For example, they can specifically describe a scenario such as "a character sitting on a bench in a park."
[0177] motion generation
[0178] The server analyzes plot information and automatically generates character movements. This movement generation includes setting animations and movement paths so that characters move naturally. Furthermore, it can use physics simulation to reflect realistic movements.
[0179] Sound effect selection
[0180] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[0181] Background music selection
[0182] Based on the plot information, the server selects background music to create an appropriate atmosphere for the scene, for example, relaxing music for a park scene.
[0183] Video Integration and Generation
[0184] The server then combines the character models, movements, sound effects, and background music to create a video, syncing all the elements together to create a single, high-quality video that is then uploaded to a content distribution service.
[0185] Sending to user terminal
[0186] The generated video is sent to the user's device, and the user can watch the video on their device by downloading or streaming it.
[0187] Hardware and software used
[0188] The implementation of this system uses high-performance cloud servers (e.g., Amazon EC2 or Google Cloud Platform) and the Python library aiohttp, an asynchronous HTTP client, with the following APIs:
[0189] Character Model Generation API
[0190] Plotting behavior generation API
[0191] Sound Effect Selection API
[0192] Background music selection API
[0193] Video Generation API
[0194] Specific examples
[0195] A user uses a smartphone application to enter the following information:
[0196] Character description: "A character with a cheerful personality, red hair, and blue clothes."
[0197] Plot information: "A character sits on a bench in a park."
[0198] Sound effects: park ambient sounds, character footsteps
[0199] Background music: Relaxing music
[0200] Based on this information, the server integrates the character's 3D model, movements, sound effects, and background music to generate a high-quality video, which is then uploaded to a content distribution service.
[0201] Prompt Sentence Examples
[0202] "Please create a scene where the user creates a character with a cheerful personality, red hair, and blue clothes, sitting on a bench in a park. Use the ambient sounds of the park and the character's footsteps as sound effects, and choose a relaxing song for the background music."
[0203] The system allows users to easily generate high-quality videos and share them on content distribution services.
[0204] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0205] Step 1:
[0206] The user inputs character settings using a terminal. Specifically, the user enters detailed information about the character, such as personality, appearance, and clothing, into an input form and sends it to the server. The input data includes the character's name, hair color, clothing color, personality, etc. The server analyzes this data and prepares it.
[0207] Step 2:
[0208] The server receives the character setting information sent by the user and generates a 3D model of the character based on that information. The server uses an AI model to generate a character model with the specified characteristics and stores it in an internal database. The generated character model has the characteristics specified by the user, such as appearance and personality.
[0209] Step 3:
[0210] The user inputs plot information using the device, specifically details of the scenes required for the video (e.g., "A character sits on a bench in a park"), and the device sends this plot information to the server.
[0211] Step 4:
[0212] The server analyzes the received plot information and automatically generates character movements. The server combines AI models and physics simulations to generate natural movements and movement paths. For example, based on the input plot information, it generates a series of movements for a character to walk and sit on a bench.
[0213] Step 5:
[0214] The server selects appropriate sound effects based on the generated actions and environment. Specifically, it selects sound effects that correspond to the character's actions (e.g., walking, sitting, etc.) and environmental sounds (e.g., birds chirping in the park). It then extracts appropriate audio files from the audio database and synchronizes them with the actions.
[0215] Step 6:
[0216] The server selects background music based on the plot information. Specifically, it selects music with a relaxing atmosphere to match the scene. If necessary, it adjusts the tempo and atmosphere of the music to match the plot information.
[0217] Step 7:
[0218] The server then combines the generated character models, movements, sound effects, and background music to generate a video, adjusting the timing of all elements and stitching them together into a single, high-quality video. The video file is generated internally and can later be downloaded or streamed.
[0219] Step 8:
[0220] The server uploads the generated video to a content distribution service, where users can access the video over the Internet and share it with other users.
[0221] Step 9:
[0222] The server then sends the generated video to the user's device, where the user can receive it and save it on their device or watch it via streaming. Communication between the server and the device is encrypted to ensure secure data transfer.
[0223] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0224] This invention is a system that allows users to easily generate high-quality videos, and provides a means to automatically integrate user-specified character settings, plot information, sound effects, and background music and output them as a single video. This system also incorporates an emotion engine that recognizes the user's emotions and modifies or changes the character's movements, facial expressions, background music, and plot information based on those emotions.
[0225] Program Overview
[0226] Character setting input
[0227] Users use the device to input detailed settings such as the character's personality, appearance, clothing, etc., allowing them to design a character to their liking.
[0228] Character Model Generation
[0229] The server receives the character setting information sent from the device, analyzes it, and generates a character model, which is generated as a three-dimensional digital representation.
[0230] Plot information input
[0231] The user inputs the video scenario and scene details (plot information) via the device. For example, the user can specifically describe the scenario, such as "a character sits on a bench in a park."
[0232] motion generation
[0233] The server analyzes the plot information and generates movement animations that determine how the characters will move in a given scenario or scene, including setting up animations and movement paths for the characters to move naturally.
[0234] Sound effect selection
[0235] After the action animation is determined, the server selects sound effects that are appropriate for the action and the surrounding environment.The server uses physics simulation to determine the sound effects that correspond to the environment and action.
[0236] Background music selection
[0237] Based on the plot information, the server selects background music that matches the scene, selecting music from a database that matches the atmosphere of the plot and creating the atmosphere of the scene.
[0238] Using the Emotion Engine
[0239] The server receives data from the device to recognize the user's emotions (e.g., voice, facial expressions, behavior logs, etc.). The emotion engine analyzes this data and recognizes the user's emotional state. For example, if the user is laughing happily, the emotion engine recognizes a positive emotion.
[0240] Change movements and expressions
[0241] Based on the user's recognized emotional state, the server changes the character's behavior and facial expression. For example, if the user looks sad, the character's facial expression and behavior will also change to look sad.
[0242] Change background music
[0243] Based on the recognized emotion, the server selects appropriate background music. For example, if the user is in a relaxed state, background music will be selected to create a relaxing atmosphere throughout the scene.
[0244] Modifying plot information
[0245] The emotion engine may automatically modify plot information depending on the user's emotions. For example, if the user is feeling stressed, the scenario may be changed to a simpler, calmer one.
[0246] Video Integration and Generation
[0247] The server combines all elements—character models, movements, sound effects, background music, and emotion-based changes—to generate a single video, resulting in a high-quality video optimized for the user's emotions.
[0248] Specific examples
[0249] Character Settings
[0250] When a user selects a character with a bright personality, red hair, and blue clothes on their device screen, that information is sent to the server, which then generates a character model based on that information.
[0251] Input of plot information and generation of motion
[0252] Users input plot information, such as "a character sits on a bench in the park," which is sent to a server that analyzes the information and generates natural-looking movements for the character.
[0253] Using emotion engines and video generation
[0254] When a user watches a video, the device sends emotional data to the server. For example, if the user is smiling, the emotion engine recognizes this positive emotion and changes the character's facial expressions and movements to a brighter tone and the background music to match. The server combines these elements to generate a video, which is then sent to the device. The user can then download and watch the video.
[0255] In this way, the system of the present invention allows the user to automatically generate high-quality videos that are optimized to suit the user's emotions simply by making simple, detailed settings.
[0256] The processing flow will be explained below.
[0257] Step 1:
[0258] The user opens the character setting screen on the device and enters detailed settings for the character, such as personality, appearance, clothing, etc. Once the input is complete, the setting information is sent to the server.
[0259] Step 2:
[0260] The server receives the character setting information sent from the device. The server analyzes this information and generates a character model with the set characteristics. The generated character model is expressed as a digital 3D model.
[0261] Step 3:
[0262] The user inputs the video scenario and scene details (plot information) on the device. For example, they specify specific locations and actions, such as "a character sitting on a bench in a park." Once input is complete, the plot information is sent to the server.
[0263] Step 4:
[0264] The server receives the plot information sent from the device. The server analyzes the plot information and generates movement animations to determine how the character should move in the specified scenario or scene. For example, it calculates and sets the movement of a character walking and sitting on a park bench.
[0265] Step 5:
[0266] After the movement animation is determined, the server selects sound effects appropriate for that movement and the surrounding environment, using physics simulation to determine, for example, the sound of a character walking or the ambient sounds of a park (birds chirping, wind noise, etc.).
[0267] Step 6:
[0268] The server selects background music (BGM) that matches the scene based on the plot information. The server selects music from a database that matches the atmosphere of the plot. For example, it selects calming BGM to match a relaxing scene in a park.
[0269] Step 7:
[0270] The device collects the user's emotional data (voice, facial expression, behavior log, etc.) and sends it to the server. For example, the device's camera and microphone can be used to analyze the user's facial expression and tone of voice.
[0271] Step 8:
[0272] The server analyzes the received emotion data using an emotion engine to identify the user's emotional state. For example, if the user is smiling, it recognizes a positive emotion.
[0273] Step 9:
[0274] Based on the user's recognized emotions, the server changes the character's movements and facial expressions. For example, if the user is smiling, the character will also smile.
[0275] Step 10:
[0276] Based on the user's recognized emotion, the server changes the background music that is set. For example, if the user is relaxed, it selects a calm background music that matches the user's emotion.
[0277] Step 11:
[0278] The server combines all the elements—character models, movements, sound effects, background music, and emotion-based modifications—to generate a single video.
[0279] Step 12:
[0280] The server sends the generated video to the user's device, where the user can download and watch the video.
[0281] Step 13:
[0282] If a user wishes to make corrections to a video, they input the corrections from their device and send them back to the server. The server then regenerates the character model and movement animations based on the corrections, selects new sound effects and background music, and regenerates the corrected video.
[0283] Example 2
[0284] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0285] Conventional video creation systems have the problem that it is difficult for users to customize based on their emotions, and the quality of the generated videos is not optimized for the user's emotions and preferences. In addition, adjusting each element (character settings, plot information, sound effects, background music) individually requires a great deal of effort and expertise, making it difficult to easily generate high-quality videos.
[0286] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0287] In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for receiving data for recognizing the user's emotion from the terminal, means for changing the character's movements and facial expressions based on the recognized emotional state of the user, means for selecting appropriate background music based on the recognized emotion, means for automatically correcting the plot information according to the user's emotional state, and means for generating a video by integrating the character model, movements, sound effects, and background music. This enables users to easily generate high-quality videos optimized to their own emotions.
[0288] "Users" are individuals or organizations that operate the system and input character settings and plot information.
[0289] The "server" is a central processing unit that receives and analyzes data sent by users and generates character models, movements, sound effects, background music, etc.
[0290] "Character settings" refers to detailed information about a character's personality, appearance, clothing, etc.
[0291] A "character model" is a three-dimensional digital representation generated based on a character's settings.
[0292] "Plot information" is information about the scenario and scene details of a video, and is specified by the user.
[0293] "Movement animation" refers to the specific movements and actions that characters perform based on plot information.
[0294] "Sound effects" are sound effects that are appropriate for the character's movements and environment, and are selected based on physical models.
[0295] "Background music" is music used to create the overall atmosphere of a video scene.
[0296] The "emotion engine" is a technology that analyzes user emotions and customizes various elements of a video based on those emotions.
[0297] "Emotion recognition data" is data used to recognize a user's emotions, and includes voice, facial expressions, behavioral logs, etc.
[0298] "Video generation means" is a mechanism for integrating character models, movements, sound effects, and background music to generate a single video.
[0299] This invention is a system that allows users to easily generate high-quality videos. Users input character settings, plot information, sound effects, and background music using a terminal, and the system automatically integrates them and outputs them as a single video. The system also uses an emotion engine that recognizes the user's emotions and modifies or changes the character's movements, facial expressions, background music, and plot information based on those emotions.
[0300] This system operates in cooperation with both the server and the terminal. The main processing steps and related hardware and software are explained below.
[0301] Character setting input
[0302] Using the device interface, users input detailed settings for their character, such as personality, appearance, clothing, etc. This information is sent from the device to a server, which then analyzes the character settings information.
[0303] Character Model Generation
[0304] The server generates a character model based on the received character setting information. Specifically, it uses 3D modeling software such as Blender or Unity. This character model is then used for subsequent movement generation and animation.
[0305] Plot information input
[0306] The user inputs the video scenario and scene details (plot information) via the terminal. For example, the user inputs a scenario such as "A character sits on a bench in a park." This information is also sent to the server and stored as plot information.
[0307] motion generation
[0308] The server analyzes the plot information and determines how the characters will move in the specified scenario or scene. Animation software such as Maya or Autodesk MotionBuilder is used for movement animation. The generated movements include animations and movement paths for the characters to move naturally.
[0309] Sound effect selection
[0310] After the movement animation is determined, the server selects sound effects appropriate for the movement and environment, using sound effect generation tools such as FMOD and Wwise with physical simulation.
[0311] Background music selection
[0312] Based on the plot information, the server selects background music that matches the scene, selected from music libraries such as Epidemic Sound and Artlist.
[0313] Using the Emotion Engine
[0314] To recognize the user's emotions, the device uses a camera and microphone to collect the user's facial and voice data and sends it to the server. The server then uses an emotion engine (such as OpenCV or Affectiva) to analyze the user's emotional state. For example, if the user is smiling, it is recognized as a positive emotion.
[0315] Change movements and expressions
[0316] The server changes the character's behavior and facial expression based on the user's recognized emotional state. For example, if the user is sad, the character's facial expression and behavior will also change to a sad one.
[0317] Change background music
[0318] The server selects appropriate background music based on the recognized emotion, for example, if the user is relaxed, calm background music is selected.
[0319] Modifying plot information
[0320] The emotion engine automatically modifies plot information according to the user's emotions, so if the user is feeling stressed, the scenario will be changed to a simpler, calmer one.
[0321] Video Integration and Generation
[0322] The server then combines all elements based on the character model, movements, sound effects, background music, and emotional changes into a single video. The final high-quality video is then generated using video editing software such as Adobe After Effects or Final Cut Pro. The resulting video is then sent to the user's device, where they can download and watch it.
[0323] Examples of prompt statements
[0324] Here are some examples of prompts to input to a generative AI model:
[0325] "Create a scene in a park where a character with red hair and a cheerful personality is sitting on a bench. Not only should the character be smiling and sitting on the bench, but also integrate relaxing background music and natural environmental sounds. Also, dynamically change the character's facial expressions and behavior based on emotion recognition while the user is viewing the scene."
[0326] In this way, the system of the present invention allows the user to automatically generate high-quality videos that are optimized to suit the user's emotions simply by making simple, detailed settings.
[0327] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0328] Step 1:
[0329] Entering character settings
[0330] The user sets the character's personality, appearance, clothing, etc. on the terminal screen.
[0331] Input: Enter character settings information via device (e.g., "cheerful personality, red hair, blue clothes")
[0332] Processing: The terminal formats the entered information and sends it to the server.
[0333] Specific operation: Enter the setting information through the terminal's GUI and press the "Send" button, and the information will be sent to the server via API.
[0334] Output: Character setting information sent to the server
[0335] Step 2:
[0336] Character model generation
[0337] The server receives the character setting information sent from the terminal and analyzes it.
[0338] Input: Character information (e.g., "cheerful personality, red hair, blue clothes")
[0339] Process: The server uses the Blender API to generate the character model.
[0340] Specific operation: Passes configuration information to the Blender API and executes a script to generate a 3D digital model.
[0341] Output: Generated character model (3D digital model)
[0342] Step 3:
[0343] Entering plot information
[0344] The user inputs the video scenario and scene details (plot information) via the terminal.
[0345] Input: Plot information (e.g., "A character sits on a bench in a park")
[0346] Processing: The terminal formats the entered plot information and sends it to the server.
[0347] Specific operation: Plot information is entered using the GUI, and the information is sent to the server when the "Send" button is pressed.
[0348] Output: Plot information sent to the server
[0349] Step 4:
[0350] Generate movement animation
[0351] The server analyzes the plot information and generates animations of the characters' movements.
[0352] Input: Plot information (e.g., "A character sits on a bench in a park")
[0353] Processing: Generate movement animation using Maya or Autodesk MotionBuilder.
[0354] Specific operation: The server passes the analyzed plot information to the Maya API and executes a script that generates a movement animation.
[0355] Output: Generated movement animation
[0356] Step 5:
[0357] Sound Effect Selection
[0358] The server selects the appropriate sound effect according to the action animation.
[0359] Input: Action animation (e.g. "Sit on a bench")
[0360] Processing: Select sound effects based on physical simulation using FMOD and Wwise.
[0361] Specific behavior: The behavior animation data is passed to the FMOD API, and an algorithm is run to select the appropriate sound effect.
[0362] Output: Selected sound effects
[0363] Step 6:
[0364] Background music selection
[0365] The server selects background music that matches the scene based on the plot information.
[0366] Input: Plot information (e.g., "A character sits on a bench in a park")
[0367] Processing: Select music that matches the scene from the Epidemic Sound and Artlist databases.
[0368] What it does: It parses the plot information and runs an algorithm that searches the music library using the corresponding keywords.
[0369] Output: Selected background music
[0370] Step 7:
[0371] Emotion engine recognizes user emotions
[0372] To recognize the user's emotions, the device uses a camera and microphone to collect facial expressions and voice sounds, which are then sent to a server.
[0373] Input: User's facial expression data and voice data
[0374] Processing: The device formats the collected data and sends it to the server.
[0375] Specific operation: Collects data in real time through the camera and microphone and sends it to the server.
[0376] Output: Emotion recognition data sent to the server
[0377] Step 8:
[0378] Emotion analysis using an emotion engine
[0379] The server uses an emotion engine to recognize the user's emotional state.
[0380] Input: Emotion recognition data (e.g., "image of a smiling face")
[0381] Processing: Sentiment analysis is performed using Affectiva and OpenCV.
[0382] Specific operation: Runs emotion recognition algorithm and obtains emotion analysis result (e.g., positive emotion).
[0383] Output: Emotion analysis results
[0384] Step 9:
[0385] Change movements and expressions
[0386] The server changes the character's movements and facial expressions based on the results of emotion analysis.
[0387] Input: Sentiment analysis result (e.g., "positive sentiment")
[0388] Processing: Updates animation data to change the character's movements and expressions.
[0389] Specific actions: Regenerate character movements and facial expressions using Maya API and Blender API.
[0390] Output: Modified character movements and expressions
[0391] Step 10:
[0392] Change background music
[0393] The server selects appropriate background music based on the emotion analysis results.
[0394] Input: Sentiment analysis result (e.g., "positive sentiment")
[0395] Processing: Select new background music from Epidemic Sound and Artlist.
[0396] What it does: Search your music library using new keywords to find the right background music.
[0397] Output: Modified background music
[0398] Step 11:
[0399] Modifying plot information
[0400] The server automatically modifies plot information according to the user's emotional state.
[0401] Input: Sentiment analysis results (e.g., "I feel stressed")
[0402] Processing: Reanalyze the plot information and change it into a simple and calm scenario.
[0403] Specific operation: Runs an algorithm that recompiles plot information and generates a revised scenario.
[0404] Output: Corrected plot information
[0405] Step 12:
[0406] Video Integration and Generation
[0407] The server combines all elements (character models, movements, sound effects, background music, and emotional changes) to generate the video.
[0408] Input: Character models, movements, sound effects, background music, modified plot information
[0409] Processing: Edit the video using Adobe After Effects and Final Cut Pro.
[0410] What it does: Place elements on a timeline in video editing software, add effects and transitions, and generate the final video.
[0411] Output: High-quality generated video
[0412] Step 13:
[0413] Sending the generated video
[0414] The server transmits the generated high-quality video to the terminal.
[0415] Input: Generated video
[0416] Processing: Converts the video format and encodes it to prepare it for sending to the device.
[0417] Specific operation: The video is converted into the appropriate format and sent to the device using the HTTP / HTTPS protocol.
[0418] Output: Video sent to device
[0419] Step 14:
[0420] Download and watch videos
[0421] Users download and watch the videos generated on their devices.
[0422] Input: Video sent from the server
[0423] Processing: Receive, store, and play the video on the device.
[0424] Specific operation: Play the video using the device's media player.
[0425] Output: Viewable video
[0426] Through this series of processing steps, users can easily generate and watch high-quality, emotion-optimized videos.
[0427] (Application example 2)
[0428] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0429] When generating videos, conventional systems have difficulty automatically adjusting content based on the user's emotions. This has led to a demand for technology that can automatically generate high-quality videos optimized for the user's emotions and situation. Furthermore, systems that can utilize emotional data to provide content that is more attuned to the user are also needed. Furthermore, to ensure that the content of the generated videos is natural, the elements selected (such as character movements, sound effects, and background music) must also be optimized based on the user's emotions.
[0430] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for recognizing user emotion data and adjusting the character model, movements, sound effects, and background music based on the emotion, and means for generating a video by integrating the character model, movements, sound effects, and background music. This makes it possible to automatically generate high-quality videos optimized for the user's emotions and situation.
[0431] "Means for receiving character settings specified by the user" refers to a function for receiving detailed information such as the character's personality, appearance, clothing, etc., entered by the user using the terminal.
[0432] The "means for generating a character model based on character settings" is a function for generating a digital three-dimensional character model based on character setting information specified by the user.
[0433] The "means for receiving plot information specified by the user" is a function for receiving detailed information about the scenario and scenes of the video input by the user.
[0434] The "means for generating character actions based on plot information" is a function for analyzing received plot information and generating character actions that conform to the scenario.
[0435] The "means for selecting sound effects based on the generated actions and environment" is a function for selecting the most appropriate sound effects according to the generated character's actions and the environment.
[0436] The "means for selecting background music based on plot information" is a function for analyzing plot information and selecting background music suitable for a scene.
[0437] "Means for recognizing a user's emotional data and adjusting a character model, movements, sound effects, and background music based on that emotion" refers to a function that analyzes a user's emotional data (voice, facial expressions, etc.) and adjusts a character's movements, facial expressions, sound effects, and background music according to that emotional state.
[0438] "A means of integrating character models, movements, sound effects, and background music to generate video" is a function that combines all of these elements to generate a single, integrated, high-quality video.
[0439] This invention is a system that generates high-quality videos based on user emotions. Specifically, the server integrates character settings, plot information, sound effects, and background music to automatically generate videos that match the user's emotions.
[0440] Program Overview
[0441] 1. Enter character settings
[0442] The user uses the device to input detailed settings for the character, such as personality, appearance, clothing, etc. For example, they might set a character with a cheerful personality, red hair, and blue clothing. This information is then sent to the server.
[0443] 2. Character Model Generation
[0444] The server analyzes the character setting information and generates a three-dimensional digital character model.
[0445] 3. Enter plot information
[0446] The user inputs the video scenario and scene details (plot information) via the device. For example, the user can specifically describe the scenario, such as "A character sits on a bench in a park." This information is then sent to the server.
[0447] 4. Behavior generation
[0448] The server analyzes the plot information and generates animations for characters to move naturally, including the character's movement path and body movements.
[0449] 5. Sound effect selection
[0450] After the movement animation is generated, the server selects sound effects that suit the movement and the surrounding environment, such as "birds chirping" for a park scene.
[0451] 6. Background Music Selection
[0452] Based on the plot information, the server selects background music that matches the scene. Music that creates the atmosphere of the scene is selected from a database.
[0453] 7. Use of Emotion Engine
[0454] The server receives data from the device to recognize the user's emotions (voice, facial expressions, behavioral logs, etc.), and the emotion engine analyzes this data. For example, if the user is laughing happily, the emotion engine will recognize a positive emotion.
[0455] 8. Change movement and facial expression
[0456] Based on the user's recognized emotional state, the server changes the character's behavior and facial expressions. For example, if the user looks sad, the character will also look sad.
[0457] 9. Change the background music
[0458] The emotion engine selects appropriate background music based on the emotion it recognizes. For example, if the user is in a relaxed state, music that creates a relaxing atmosphere will be selected.
[0459] 10. Correcting plot information
[0460] The emotion engine may also automatically modify plot information depending on the user's emotions, for example, changing the scenario to a simpler, calmer one if the user is feeling stressed.
[0461] 11. Video Integration and Generation
[0462] The server combines the character models, movements, sound effects, background music, and emotion-based modifications to generate a single video.
[0463] Hardware and software used
[0464] Hardware:
[0465] Smartphone or tablet
[0466] Camera (to capture facial expressions)
[0467] Microphone (to capture audio)
[0468] software:
[0469] EmotionRecognizer (emotion recognition library)
[0470] VideoGenerator (library for generating videos: OpenCV as an example)
[0471] Specific examples
[0472] Next, we will explain a specific example of how this system can be used. A user uses a smartphone to create a character with a cheerful personality, red hair, and blue clothes, and enters plot information such as "the character is sitting on a bench in a park." If the user is smiling, the emotion engine recognizes this positive emotion, brightens the character's facial expressions and movements, and changes the background music to a more cheerful one. As a result, a high-quality video optimized for the user is generated.
[0473] Example prompt sentence:
[0474] "Change the character's facial expression to a smile, set the background music to upbeat music, and generate a high-quality video."
[0475] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0476] Step 1:
[0477] The user inputs character setting information using a terminal. Specifically, detailed information such as the character's personality, appearance, and clothing is entered, and this information is sent from the terminal to the server. The input data could be specific information such as "a character with a cheerful personality, red hair, and blue clothing." The server then stores this input data in a database.
[0478] Step 2:
[0479] The server generates a digital 3D character model based on the received character configuration information. It uses a generative AI model to analyze this information and generate the character model using 3D modeling software (e.g., Blender). The output is a completed 3D character model.
[0480] Step 3:
[0481] The user inputs the video's scenario and scene details (plot information) via the device. This input data includes specific plot information, such as "a character sits on a bench in a park." The device then sends this information to the server, which then stores the received plot information in a database.
[0482] Step 4:
[0483] The server analyzes the plot information and generates character movement animation. For example, based on the plot information "a character sits on a bench in a park," it generates a natural movement path for the character. Here, 3D animation software (e.g., Maya) is used. The output is movement animation data.
[0484] Step 5:
[0485] The server selects sound effects based on the generated action and environment. For example, for a park scene, sound effects such as "birds chirping" and "wind sounds" are selected. A physics simulation library (e.g., Bullet Physics) is used to select appropriate sound effects. The output is the selected sound effect file.
[0486] Step 6:
[0487] The server selects background music based on the plot information. It analyzes the atmosphere and selects music from a database that matches the scene. For example, relaxing music is selected for a tranquil scene in a park. The output is the selected background music file.
[0488] Step 7:
[0489] The server analyzes the user's emotional data (voice, facial expressions, behavioral logs, etc.) received from the device to recognize the user's emotional state. It uses emotion recognition software (e.g., Facial Emotion Recognition). For example, if the user is smiling happily, the server recognizes a positive emotion. The output is the user's emotional state data.
[0490] Step 8:
[0491] The server adjusts the character's behavior, facial expressions, sound effects, and background music based on the user's recognized emotional state. For example, if the user is sad, the server changes the character's facial expression to a sad one and the background music to a calm one. The output is each element adjusted based on the emotion.
[0492] Step 9:
[0493] The server integrates character models, movements, sound effects, and background music to create a video using video editing software (e.g., Adobe Premiere Pro). The output is a final, high-quality video file.
[0494] Step 10:
[0495] The generated video is sent from the server to the user's device, where the user can watch it. The output is a video file that the user can watch.
[0496] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0497] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0498] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0499] [Second embodiment]
[0500] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0501] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0502] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0503] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0504] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0505] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0506] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0507] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0508] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0509] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0510] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0511] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0512] This invention is a system that allows users to easily generate high-quality videos, and provides a means for automatically integrating user-specified character settings, plot information, sound effects, and background music and outputting them as a single video.
[0513] Program Overview
[0514] Character setting input
[0515] Users can use their devices to enter detailed settings such as the character's personality, appearance, clothing, etc., through an interface, allowing them to specifically design their own character.
[0516] Character Model Generation
[0517] The server receives the character setting information sent from the device, analyzes it, and generates a character model, a three-dimensional digital representation of the character with an appearance and personality based on the setting.
[0518] Plot information input
[0519] The user inputs the details of the scenes (plot information) required for the video via the terminal. For example, they can specifically describe the scenario, such as "a character sitting on a bench in a park."
[0520] motion generation
[0521] The server analyzes the plot information and automatically generates character movements, including setting animations and movement paths so that the characters move naturally.
[0522] Sound effect selection
[0523] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[0524] Background music selection
[0525] Based on the plot information, the server selects background music to create an appropriate atmosphere for the scene, for example, relaxing music for a park scene.
[0526] Video Integration and Generation
[0527] The server generates the video by integrating character models, movements, sound effects, and background music, and then synchronizes all the elements to create a single, high-quality video.
[0528] Specific examples
[0529] Character Settings
[0530] The user sets up a character with a bright personality, red hair, and blue clothes on the device screen. This information is sent from the device to the server.
[0531] Character model generation
[0532] Based on the user's settings, the server generates a character model with red hair and blue clothes. This character has a cheerful personality, which is reflected in their facial expressions and movements.
[0533] Behavior Settings
[0534] When a user inputs plot information such as "a character sits on a bench in a park," the server receives this information and automatically generates a series of actions for the character to walk and sit on the bench.
[0535] Sound effects and background music selection
[0536] Using physics simulation, the server selects sound effects that match the park's ambient sounds and the characters' movements, as well as background music that creates a relaxing atmosphere.
[0537] Video Integration and Generation
[0538] The server then combines the generated character models, movements, sound effects, and background music into a single video file, which users can then download or stream to their devices.
[0539] In this way, the system of the present invention allows the user to automatically generate professional quality videos simply by configuring detailed settings.
[0540] The processing flow will be explained below.
[0541] Step 1:
[0542] The user opens the character setting screen on the device and enters detailed settings for the character, such as personality, appearance, clothing, etc. Once the input is complete, the setting information is sent to the server.
[0543] Step 2:
[0544] The server receives the character setting information sent from the device. The server analyzes this information and generates a character model with the set characteristics. The generated character model is expressed as a digital 3D model.
[0545] Step 3:
[0546] The user inputs the video scenario and scene details (plot information) on the device. For example, they specify specific locations and actions, such as "a character sitting on a bench in a park." Once input is complete, the plot information is sent to the server.
[0547] Step 4:
[0548] The server receives the plot information sent from the device. The server analyzes the plot information and generates movement animations to determine how the character should move in the specified scenario or scene. For example, it calculates and sets the movement of a character walking and sitting on a park bench.
[0549] Step 5:
[0550] After the movement animation is determined, the server selects sound effects appropriate for that movement and the surrounding environment, using physics simulation to determine, for example, the sound of a character walking or the ambient sounds of a park (birds chirping, wind noise, etc.).
[0551] Step 6:
[0552] The server selects background music (BGM) that matches the scene based on the plot information. The server selects music from a database that matches the atmosphere of the plot. For example, it selects calming BGM to match a relaxing scene in a park.
[0553] Step 7:
[0554] The server combines character models, movement animations, sound effects, and background music into a single video, precisely timing and matching each element to produce a high-quality video.
[0555] Step 8:
[0556] The server sends the generated video to the user's device, where the user can download and watch the video.
[0557] Step 9:
[0558] If a user wishes to make corrections to a video, they input the corrections from their device and send them back to the server. The server then regenerates the character model and movement animations based on the corrections, selects new sound effects and background music, and regenerates the corrected video.
[0559] Example 1
[0560] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0561] With conventional video generation systems, it was difficult for users to generate high-quality videos using the character settings and plot information they desired. Furthermore, there was a lack of means to properly integrate character movements, sound effects, and background music, which often resulted in a decline in the quality of the generated videos. There is a need for a system that can solve these problems and enable easier, high-quality video generation.
[0562] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0563] In this invention, the server includes means for receiving character settings specified by a user, means for generating a character model based on the character settings, means for receiving plot information specified by a user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for generating a video by integrating the character model, movements, sound effects, and background music, means for selecting environmental sounds and movement sounds using a physical simulation based on the generated character model and movements, means for transmitting the generated video to a user terminal, means for synchronizing the generated character model, movements, sound effects, and background music using a timeline editor, and means for automatically generating a character model and video using a generative AI model based on the character settings and plot information input by a user. This allows users to easily generate high-quality videos and enables the realization of well-integrated, professional-quality videos.
[0564] "Character configuration" is the process by which a user inputs details about a character, such as personality, appearance, and clothing.
[0565] A "character model" is a three-dimensional digital character generated based on a character setting.
[0566] "Plot information" is detailed information about the video scenario specified by the user.
[0567] "Character behavior" refers to a series of actions that a character performs based on plot information.
[0568] "Sound effects" are audio data used to add realistic sound effects to character movements and environments.
[0569] "Background music" is music selected to create the atmosphere of a scene based on plot information.
[0570] "Method of generating animation" is the process of integrating character models, movements, sound effects, and background music into a single format.
[0571] "Physics simulation" is a simulation technique for providing realistic sound effects based on generated motions and environments.
[0572] A "timeline editor" is a tool used to properly synchronize character models, movements, sound effects, and background music.
[0573] A "generative AI model" is an algorithm that uses AI technology to automatically generate character models and videos.
[0574] "User terminal" refers to the device through which a user inputs character settings and plot information.
[0575] This invention is a system that allows users to easily generate high-quality videos. It provides a means to automatically integrate user-specified character settings, plot information, sound effects, and background music and output them as a single video. The system consists of a server and a user terminal.
[0576] First, the user uses the device to input detailed settings for the character, such as personality, appearance, and clothing. Once the character settings are complete, the device sends this information to the server. For example, if the user sets a character with a cheerful personality, red hair, and blue clothing, this information is sent to the server.
[0577] The server uses a generative AI model based on the received character setting information to generate a three-dimensional digital character model. This character model reflects the character's set appearance and personality. Next, the user inputs the details of the scenes (plot information) required for the video via their device. For example, they input a scenario such as "the character sits on a bench in a park." This information is also sent from the device to the server.
[0578] The server analyzes the plot information and automatically generates character movements, including animations and movement paths that allow the character to move naturally. The server then selects sound effects based on the generated movements and the environment. Physics simulation is used to determine audio data that realistically reflects the character's movements and environmental sounds.
[0579] The server then selects background music based on the plot information. To create a suitable atmosphere for the scene, relaxing music is selected for the park scene. After these elements are gathered, the server generates a video by integrating character models, movements, sound effects, and background music. A timeline editor is used to properly align the timing of all elements to create a high-quality video.
[0580] The generated video is sent from the server to the user's device, where the user can watch it by downloading or streaming. This system allows users to automatically generate professional-quality videos by simply configuring detailed settings.
[0581] Specific examples
[0582] The user configures the character settings on the device screen. For example, they can configure a character with a cheerful personality, red hair, and blue clothes, and send this information from the device to the server. The server generates a character model with red hair and blue clothes based on the configuration information received from the user. This character has a cheerful personality, which is reflected in their facial expressions and movements. Next, the user inputs plot information such as "the character sits on a bench in a park." Based on this information, the server automatically generates a series of actions for the character to walk and sit on the bench. For this action, physical simulation is used to select sound effects that match the ambient sounds of the park and the character's movements. Background music is also selected to create a relaxing atmosphere.
[0583] Once all these elements are in place, the server combines the character movements, sound effects, and background music into a video that users can then download or stream to their device.
[0584] Prompt Sentence Examples
[0585] Character: A cheerful character with red hair and blue clothes.
[0586] Plot Info: A character sits on a bench in a park.
[0587] Sound Effects: Added ambient bird sounds and character walking sounds
[0588] Background music: Relaxing music
[0589] When this prompt is input into a generative AI model, a video that integrates all the elements is automatically generated.
[0590] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0591] Step 1:
[0592] The user inputs character setting information using the terminal. In the input form, the user sets details such as the character's personality, appearance, and clothing. For example, the user can set a "character with a cheerful personality, red hair, and blue clothes." Once this information is entered, the terminal sends the setting information to the server.
[0593] Input: Detailed setting information such as character personality, appearance, clothing, etc.
[0594] Output: Character configuration information sent to the server.
[0595] Step 2:
[0596] The server analyzes the character setting information received from the device, and the generative AI model on the server generates a three-dimensional digital character model based on this information.
[0597] Input: Character configuration information.
[0598] Output: 3D digital character model.
[0599] What it does: The generative AI model references the database and sets the character's hair color to red and their clothing to blue.
[0600] Step 3:
[0601] The user inputs the plot information required for the video via the terminal. In the input form, the user writes down the details of the scene. For example, the user enters a scenario such as "A character sits on a bench in a park." Once this information is entered, the terminal sends the plot information to the server.
[0602] Input: Scene details (plot information).
[0603] Output: Plot information sent to the server.
[0604] Step 4:
[0605] The server analyzes the plot information and automatically generates character movements, including animations and movement paths that allow the characters to move naturally.
[0606] Input: Plot information.
[0607] Output: Character animation data and movement path.
[0608] Specific Action: Based on the instruction "sit on a bench in the park," the plot analysis module generates the action of the character walking and sitting on the bench.
[0609] Step 5:
[0610] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[0611] Input: Character animation data and environment information.
[0612] Output: Suitable sound effects.
[0613] Specific operation: The sound effect selection module selects sounds such as birds singing, footsteps, and bench sounds from the sound library.
[0614] Step 6:
[0615] The server selects background music based on plot information, determining the background music to create an atmosphere that matches the scene.
[0616] Input: Plot information.
[0617] Output: Suitable background music.
[0618] Specific operation: The music selection module selects relaxing music from the music database.
[0619] Step 7:
[0620] The server uses a timeline editor to combine character models, movements, sound effects, and background music into a high-quality video, aligning all elements appropriately to create the final video.
[0621] Input: Character models, animation data, sound effects, background music.
[0622] Output: The finished video file.
[0623] Specific operation: Each element is integrated in the timeline editor, then encoded and generated into a single video.
[0624] Step 8:
[0625] The server sends the generated video to the user's device, where the user can watch the video by downloading or streaming.
[0626] Input: Your finished video file.
[0627] Output: The video sent to the user's device.
[0628] Specific operation: The server compresses the video file and sends it to the user's device. The user then downloads the video and watches it.
[0629] (Application example 1)
[0630] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0631] Conventional video generation systems require a lot of manual work from the user and specialized knowledge and skills, making it difficult for anyone to easily generate high-quality videos and share them on content distribution services. Furthermore, integrating character settings, plot information, sound effects, and background music during video generation is complicated, making it difficult to maintain professional quality.
[0632] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0633] In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for generating a video by integrating the character model, movements, sound effects, and background music, and means for uploading the generated video to a content distribution service. This allows users to easily generate high-quality videos and share them on the content distribution service.
[0634] "Character settings" are detailed setting information such as the character's personality, appearance, clothing, etc., specified by the user.
[0635] A "character model" is a three-dimensional digital representation generated based on a character's settings.
[0636] "Plot information" is detailed information about the scenes and scenarios required for the video specified by the user.
[0637] "Movement generation" is the process of automatically generating character movements based on plot information.
[0638] "Sound effects" are sound effects that are selected based on the action and environment being generated.
[0639] "Background music" is music that is selected to fit a scene based on plot information.
[0640] "Video generation" is the process of integrating character models, movements, sound effects, and background music to generate a single, high-quality video.
[0641] A "content distribution service" is a platform for sharing and viewing generated videos over the Internet.
[0642] A "user terminal" is an electronic device operated by a user, such as a smartphone or computer.
[0643] "Physics simulation" is a simulation technology that realistically reflects the movements and sound effects of generated character models.
[0644] "Smartphone Application" means a smartphone application that enables users to create, distribute, and guide their own video content.
[0645] The present invention provides a system that enables users to easily generate high-quality videos and share them via content distribution services. Specific embodiments of this system are described below.
[0646] Character setting input
[0647] The device provides users with a means to input detailed settings such as the character's personality, appearance, clothing, etc., allowing them to specifically design their own character.
[0648] Character Model Generation
[0649] The server receives the character setting information sent from the device, analyzes it, and generates a character model, a three-dimensional digital representation with an appearance and personality based on the setting.
[0650] Plot information input
[0651] Users are provided with a means to input details of scenes (plot information) required for the video via their device. For example, they can specifically describe a scenario such as "a character sitting on a bench in a park."
[0652] motion generation
[0653] The server analyzes plot information and automatically generates character movements. This movement generation includes setting animations and movement paths so that characters move naturally. Furthermore, it can use physics simulation to reflect realistic movements.
[0654] Sound effect selection
[0655] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[0656] Background music selection
[0657] Based on the plot information, the server selects background music to create an appropriate atmosphere for the scene, for example, relaxing music for a park scene.
[0658] Video Integration and Generation
[0659] The server then combines the character models, movements, sound effects, and background music to create a video, syncing all the elements together to create a single, high-quality video that is then uploaded to a content distribution service.
[0660] Sending to user terminal
[0661] The generated video is sent to the user's device, and the user can watch the video on their device by downloading or streaming it.
[0662] Hardware and software used
[0663] The implementation of this system uses high-performance cloud servers (e.g., Amazon EC2 or Google Cloud Platform) and the Python library aiohttp, an asynchronous HTTP client, with the following APIs:
[0664] Character Model Generation API
[0665] Plotting behavior generation API
[0666] Sound Effect Selection API
[0667] Background music selection API
[0668] Video Generation API
[0669] Specific examples
[0670] A user uses a smartphone application to enter the following information:
[0671] Character description: "A character with a cheerful personality, red hair, and blue clothes."
[0672] Plot information: "A character sits on a bench in a park."
[0673] Sound effects: park ambient sounds, character footsteps
[0674] Background music: Relaxing music
[0675] Based on this information, the server integrates the character's 3D model, movements, sound effects, and background music to generate a high-quality video, which is then uploaded to a content distribution service.
[0676] Prompt Sentence Examples
[0677] "Please create a scene where the user creates a character with a cheerful personality, red hair, and blue clothes, sitting on a bench in a park. Use the ambient sounds of the park and the character's footsteps as sound effects, and choose a relaxing song for the background music."
[0678] The system allows users to easily generate high-quality videos and share them on content distribution services.
[0679] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0680] Step 1:
[0681] The user inputs character settings using a terminal. Specifically, the user enters detailed information about the character, such as personality, appearance, and clothing, into an input form and sends it to the server. The input data includes the character's name, hair color, clothing color, personality, etc. The server analyzes this data and prepares it.
[0682] Step 2:
[0683] The server receives the character setting information sent by the user and generates a 3D model of the character based on that information. The server uses an AI model to generate a character model with the specified characteristics and stores it in an internal database. The generated character model has the characteristics specified by the user, such as appearance and personality.
[0684] Step 3:
[0685] The user inputs plot information using the device, specifically details of the scenes required for the video (e.g., "A character sits on a bench in a park"), and the device sends this plot information to the server.
[0686] Step 4:
[0687] The server analyzes the received plot information and automatically generates character movements. The server combines AI models and physics simulations to generate natural movements and movement paths. For example, based on the input plot information, it generates a series of movements for a character to walk and sit on a bench.
[0688] Step 5:
[0689] The server selects appropriate sound effects based on the generated actions and environment. Specifically, it selects sound effects that correspond to the character's actions (e.g., walking, sitting, etc.) and environmental sounds (e.g., birds chirping in the park). It then extracts appropriate audio files from the audio database and synchronizes them with the actions.
[0690] Step 6:
[0691] The server selects background music based on the plot information. Specifically, it selects music with a relaxing atmosphere to match the scene. If necessary, it adjusts the tempo and atmosphere of the music to match the plot information.
[0692] Step 7:
[0693] The server then combines the generated character models, movements, sound effects, and background music to generate a video, adjusting the timing of all elements and stitching them together into a single, high-quality video. The video file is generated internally and can later be downloaded or streamed.
[0694] Step 8:
[0695] The server uploads the generated video to a content distribution service, where users can access the video over the Internet and share it with other users.
[0696] Step 9:
[0697] The server then sends the generated video to the user's device, where the user can receive it and save it on their device or watch it via streaming. Communication between the server and the device is encrypted to ensure secure data transfer.
[0698] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0699] This invention is a system that allows users to easily generate high-quality videos, and provides a means to automatically integrate user-specified character settings, plot information, sound effects, and background music and output them as a single video. This system also incorporates an emotion engine that recognizes the user's emotions and modifies or changes the character's movements, facial expressions, background music, and plot information based on those emotions.
[0700] Program Overview
[0701] Character setting input
[0702] Users use the device to input detailed settings such as the character's personality, appearance, clothing, etc., allowing them to design a character to their liking.
[0703] Character Model Generation
[0704] The server receives the character setting information sent from the device, analyzes it, and generates a character model, which is generated as a three-dimensional digital representation.
[0705] Plot information input
[0706] The user inputs the video scenario and scene details (plot information) via the device. For example, the user can specifically describe the scenario, such as "a character sits on a bench in a park."
[0707] motion generation
[0708] The server analyzes the plot information and generates movement animations that determine how the characters will move in a given scenario or scene, including setting up animations and movement paths for the characters to move naturally.
[0709] Sound effect selection
[0710] After the action animation is determined, the server selects sound effects that are appropriate for the action and the surrounding environment.The server uses physics simulation to determine the sound effects that correspond to the environment and action.
[0711] Background music selection
[0712] Based on the plot information, the server selects background music that matches the scene, selecting music from a database that matches the atmosphere of the plot and creating the atmosphere of the scene.
[0713] Using the Emotion Engine
[0714] The server receives data from the device to recognize the user's emotions (e.g., voice, facial expressions, behavior logs, etc.). The emotion engine analyzes this data and recognizes the user's emotional state. For example, if the user is laughing happily, the emotion engine recognizes a positive emotion.
[0715] Change movements and expressions
[0716] Based on the user's recognized emotional state, the server changes the character's behavior and facial expression. For example, if the user looks sad, the character's facial expression and behavior will also change to look sad.
[0717] Change background music
[0718] Based on the recognized emotion, the server selects appropriate background music. For example, if the user is in a relaxed state, background music will be selected to create a relaxing atmosphere throughout the scene.
[0719] Modifying plot information
[0720] The emotion engine may automatically modify plot information depending on the user's emotions. For example, if the user is feeling stressed, the scenario may be changed to a simpler, calmer one.
[0721] Video Integration and Generation
[0722] The server combines all elements—character models, movements, sound effects, background music, and emotion-based changes—to generate a single video, resulting in a high-quality video optimized for the user's emotions.
[0723] Specific examples
[0724] Character Settings
[0725] When a user selects a character with a bright personality, red hair, and blue clothes on their device screen, that information is sent to the server, which then generates a character model based on that information.
[0726] Input of plot information and generation of motion
[0727] Users input plot information, such as "a character sits on a bench in the park," which is sent to a server that analyzes the information and generates natural-looking movements for the character.
[0728] Using emotion engines and video generation
[0729] When a user watches a video, the device sends emotional data to the server. For example, if the user is smiling, the emotion engine recognizes this positive emotion and changes the character's facial expressions and movements to a brighter tone and the background music to match. The server combines these elements to generate a video, which is then sent to the device. The user can then download and watch the video.
[0730] In this way, the system of the present invention allows the user to automatically generate high-quality videos that are optimized to suit the user's emotions simply by making simple, detailed settings.
[0731] The processing flow will be explained below.
[0732] Step 1:
[0733] The user opens the character setting screen on the device and enters detailed settings for the character, such as personality, appearance, clothing, etc. Once the input is complete, the setting information is sent to the server.
[0734] Step 2:
[0735] The server receives the character setting information sent from the device. The server analyzes this information and generates a character model with the set characteristics. The generated character model is expressed as a digital 3D model.
[0736] Step 3:
[0737] The user inputs the video scenario and scene details (plot information) on the device. For example, they specify specific locations and actions, such as "a character sitting on a bench in a park." Once input is complete, the plot information is sent to the server.
[0738] Step 4:
[0739] The server receives the plot information sent from the device. The server analyzes the plot information and generates movement animations to determine how the character should move in the specified scenario or scene. For example, it calculates and sets the movement of a character walking and sitting on a park bench.
[0740] Step 5:
[0741] After the movement animation is determined, the server selects sound effects appropriate for that movement and the surrounding environment, using physics simulation to determine, for example, the sound of a character walking or the ambient sounds of a park (birds chirping, wind noise, etc.).
[0742] Step 6:
[0743] The server selects background music (BGM) that matches the scene based on the plot information. The server selects music from a database that matches the atmosphere of the plot. For example, it selects calming BGM to match a relaxing scene in a park.
[0744] Step 7:
[0745] The device collects the user's emotional data (voice, facial expression, behavior log, etc.) and sends it to the server. For example, the device's camera and microphone can be used to analyze the user's facial expression and tone of voice.
[0746] Step 8:
[0747] The server analyzes the received emotion data using an emotion engine to identify the user's emotional state. For example, if the user is smiling, it recognizes a positive emotion.
[0748] Step 9:
[0749] Based on the user's recognized emotions, the server changes the character's movements and facial expressions. For example, if the user is smiling, the character will also smile.
[0750] Step 10:
[0751] Based on the user's recognized emotion, the server changes the background music that is set. For example, if the user is relaxed, it selects a calm background music that matches the user's emotion.
[0752] Step 11:
[0753] The server combines all the elements—character models, movements, sound effects, background music, and emotion-based modifications—to generate a single video.
[0754] Step 12:
[0755] The server sends the generated video to the user's device, where the user can download and watch the video.
[0756] Step 13:
[0757] If a user wishes to make corrections to a video, they input the corrections from their device and send them back to the server. The server then regenerates the character model and movement animations based on the corrections, selects new sound effects and background music, and regenerates the corrected video.
[0758] Example 2
[0759] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0760] Conventional video creation systems have the problem that it is difficult for users to customize based on their emotions, and the quality of the generated videos is not optimized for the user's emotions and preferences. In addition, adjusting each element (character settings, plot information, sound effects, background music) individually requires a great deal of effort and expertise, making it difficult to easily generate high-quality videos.
[0761] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0762] In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for receiving data for recognizing the user's emotion from the terminal, means for changing the character's movements and facial expressions based on the recognized emotional state of the user, means for selecting appropriate background music based on the recognized emotion, means for automatically correcting the plot information according to the user's emotional state, and means for generating a video by integrating the character model, movements, sound effects, and background music. This enables users to easily generate high-quality videos optimized to their own emotions.
[0763] "Users" are individuals or organizations that operate the system and input character settings and plot information.
[0764] The "server" is a central processing unit that receives and analyzes data sent by users and generates character models, movements, sound effects, background music, etc.
[0765] "Character settings" refers to detailed information about a character's personality, appearance, clothing, etc.
[0766] A "character model" is a three-dimensional digital representation generated based on a character's settings.
[0767] "Plot information" is information about the scenario and scene details of a video, and is specified by the user.
[0768] "Movement animation" refers to the specific movements and actions that characters perform based on plot information.
[0769] "Sound effects" are sound effects that are appropriate for the character's movements and environment, and are selected based on physical models.
[0770] "Background music" is music used to create the overall atmosphere of a video scene.
[0771] The "emotion engine" is a technology that analyzes user emotions and customizes various elements of a video based on those emotions.
[0772] "Emotion recognition data" is data used to recognize a user's emotions, and includes voice, facial expressions, behavioral logs, etc.
[0773] "Video generation means" is a mechanism for integrating character models, movements, sound effects, and background music to generate a single video.
[0774] This invention is a system that allows users to easily generate high-quality videos. Users input character settings, plot information, sound effects, and background music using a terminal, and the system automatically integrates them and outputs them as a single video. The system also uses an emotion engine that recognizes the user's emotions and modifies or changes the character's movements, facial expressions, background music, and plot information based on those emotions.
[0775] This system operates in cooperation with both the server and the terminal. The main processing steps and related hardware and software are explained below.
[0776] Character setting input
[0777] Using the device interface, users input detailed settings for their character, such as personality, appearance, clothing, etc. This information is sent from the device to a server, which then analyzes the character settings information.
[0778] Character Model Generation
[0779] The server generates a character model based on the received character setting information. Specifically, it uses 3D modeling software such as Blender or Unity. This character model is then used for subsequent movement generation and animation.
[0780] Plot information input
[0781] The user inputs the video scenario and scene details (plot information) via the terminal. For example, the user inputs a scenario such as "A character sits on a bench in a park." This information is also sent to the server and stored as plot information.
[0782] motion generation
[0783] The server analyzes the plot information and determines how the characters will move in the specified scenario or scene. Animation software such as Maya or Autodesk MotionBuilder is used for movement animation. The generated movements include animations and movement paths for the characters to move naturally.
[0784] Sound effect selection
[0785] After the movement animation is determined, the server selects sound effects appropriate for the movement and environment, using sound effect generation tools such as FMOD and Wwise with physical simulation.
[0786] Background music selection
[0787] Based on the plot information, the server selects background music that matches the scene, selected from music libraries such as Epidemic Sound and Artlist.
[0788] Using the Emotion Engine
[0789] To recognize the user's emotions, the device uses a camera and microphone to collect the user's facial and voice data and sends it to the server. The server then uses an emotion engine (such as OpenCV or Affectiva) to analyze the user's emotional state. For example, if the user is smiling, it is recognized as a positive emotion.
[0790] Change movements and expressions
[0791] The server changes the character's behavior and facial expression based on the user's recognized emotional state. For example, if the user is sad, the character's facial expression and behavior will also change to a sad one.
[0792] Change background music
[0793] The server selects appropriate background music based on the recognized emotion, for example, if the user is relaxed, calm background music is selected.
[0794] Modifying plot information
[0795] The emotion engine automatically modifies plot information according to the user's emotions, so if the user is feeling stressed, the scenario will be changed to a simpler, calmer one.
[0796] Video Integration and Generation
[0797] The server then combines all elements based on the character model, movements, sound effects, background music, and emotional changes into a single video. The final high-quality video is then generated using video editing software such as Adobe After Effects or Final Cut Pro. The resulting video is then sent to the user's device, where they can download and watch it.
[0798] Examples of prompt statements
[0799] Here are some examples of prompts to input to a generative AI model:
[0800] "Create a scene in a park where a character with red hair and a cheerful personality is sitting on a bench. Not only should the character be smiling and sitting on the bench, but also integrate relaxing background music and natural environmental sounds. Also, dynamically change the character's facial expressions and behavior based on emotion recognition while the user is viewing the scene."
[0801] In this way, the system of the present invention allows the user to automatically generate high-quality videos that are optimized to suit the user's emotions simply by making simple, detailed settings.
[0802] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0803] Step 1:
[0804] Entering character settings
[0805] The user sets the character's personality, appearance, clothing, etc. on the terminal screen.
[0806] Input: Enter character settings information via device (e.g., "cheerful personality, red hair, blue clothes")
[0807] Processing: The terminal formats the entered information and sends it to the server.
[0808] Specific operation: Enter the setting information through the terminal's GUI and press the "Send" button, and the information will be sent to the server via API.
[0809] Output: Character setting information sent to the server
[0810] Step 2:
[0811] Character model generation
[0812] The server receives the character setting information sent from the terminal and analyzes it.
[0813] Input: Character information (e.g., "cheerful personality, red hair, blue clothes")
[0814] Process: The server uses the Blender API to generate the character model.
[0815] Specific operation: Passes configuration information to the Blender API and executes a script to generate a 3D digital model.
[0816] Output: Generated character model (3D digital model)
[0817] Step 3:
[0818] Entering plot information
[0819] The user inputs the video scenario and scene details (plot information) via the terminal.
[0820] Input: Plot information (e.g., "A character sits on a bench in a park")
[0821] Processing: The terminal formats the entered plot information and sends it to the server.
[0822] Specific operation: Plot information is entered using the GUI, and the information is sent to the server when the "Send" button is pressed.
[0823] Output: Plot information sent to the server
[0824] Step 4:
[0825] Generate movement animation
[0826] The server analyzes the plot information and generates animations of the characters' movements.
[0827] Input: Plot information (e.g., "A character sits on a bench in a park")
[0828] Processing: Generate movement animation using Maya or Autodesk MotionBuilder.
[0829] Specific operation: The server passes the analyzed plot information to the Maya API and executes a script that generates a movement animation.
[0830] Output: Generated movement animation
[0831] Step 5:
[0832] Sound Effect Selection
[0833] The server selects the appropriate sound effect according to the action animation.
[0834] Input: Action animation (e.g. "Sit on a bench")
[0835] Processing: Select sound effects based on physical simulation using FMOD and Wwise.
[0836] Specific behavior: The behavior animation data is passed to the FMOD API, and an algorithm is run to select the appropriate sound effect.
[0837] Output: Selected sound effects
[0838] Step 6:
[0839] Background music selection
[0840] The server selects background music that matches the scene based on the plot information.
[0841] Input: Plot information (e.g., "A character sits on a bench in a park")
[0842] Processing: Select music that matches the scene from the Epidemic Sound and Artlist databases.
[0843] What it does: It parses the plot information and runs an algorithm that searches the music library using the corresponding keywords.
[0844] Output: Selected background music
[0845] Step 7:
[0846] Emotion engine recognizes user emotions
[0847] To recognize the user's emotions, the device uses a camera and microphone to collect facial expressions and voice sounds, which are then sent to a server.
[0848] Input: User's facial expression data and voice data
[0849] Processing: The device formats the collected data and sends it to the server.
[0850] Specific operation: Collects data in real time through the camera and microphone and sends it to the server.
[0851] Output: Emotion recognition data sent to the server
[0852] Step 8:
[0853] Emotion analysis using an emotion engine
[0854] The server uses an emotion engine to recognize the user's emotional state.
[0855] Input: Emotion recognition data (e.g., "image of a smiling face")
[0856] Processing: Sentiment analysis is performed using Affectiva and OpenCV.
[0857] Specific operation: Runs emotion recognition algorithm and obtains emotion analysis result (e.g., positive emotion).
[0858] Output: Emotion analysis results
[0859] Step 9:
[0860] Change movements and expressions
[0861] The server changes the character's movements and facial expressions based on the results of emotion analysis.
[0862] Input: Sentiment analysis result (e.g., "positive sentiment")
[0863] Processing: Updates animation data to change the character's movements and expressions.
[0864] Specific actions: Regenerate character movements and facial expressions using Maya API and Blender API.
[0865] Output: Modified character movements and expressions
[0866] Step 10:
[0867] Change background music
[0868] The server selects appropriate background music based on the emotion analysis results.
[0869] Input: Sentiment analysis result (e.g., "positive sentiment")
[0870] Processing: Select new background music from Epidemic Sound and Artlist.
[0871] What it does: Search your music library using new keywords to find the right background music.
[0872] Output: Modified background music
[0873] Step 11:
[0874] Modifying plot information
[0875] The server automatically modifies plot information according to the user's emotional state.
[0876] Input: Sentiment analysis results (e.g., "I feel stressed")
[0877] Processing: Reanalyze the plot information and change it into a simple and calm scenario.
[0878] Specific operation: Runs an algorithm that recompiles plot information and generates a revised scenario.
[0879] Output: Corrected plot information
[0880] Step 12:
[0881] Video Integration and Generation
[0882] The server combines all elements (character models, movements, sound effects, background music, and emotional changes) to generate the video.
[0883] Input: Character models, movements, sound effects, background music, modified plot information
[0884] Processing: Edit the video using Adobe After Effects and Final Cut Pro.
[0885] What it does: Place elements on a timeline in video editing software, add effects and transitions, and generate the final video.
[0886] Output: High-quality generated video
[0887] Step 13:
[0888] Sending the generated video
[0889] The server transmits the generated high-quality video to the terminal.
[0890] Input: Generated video
[0891] Processing: Converts the video format and encodes it to prepare it for sending to the device.
[0892] Specific operation: The video is converted into the appropriate format and sent to the device using the HTTP / HTTPS protocol.
[0893] Output: Video sent to device
[0894] Step 14:
[0895] Download and watch videos
[0896] Users download and watch the videos generated on their devices.
[0897] Input: Video sent from the server
[0898] Processing: Receive, store, and play the video on the device.
[0899] Specific operation: Play the video using the device's media player.
[0900] Output: Viewable video
[0901] Through this series of processing steps, users can easily generate and watch high-quality, emotion-optimized videos.
[0902] (Application example 2)
[0903] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0904] When generating videos, conventional systems have difficulty automatically adjusting content based on the user's emotions. This has led to a demand for technology that can automatically generate high-quality videos optimized for the user's emotions and situation. Furthermore, systems that can utilize emotional data to provide content that is more attuned to the user are also needed. Furthermore, to ensure that the content of the generated videos is natural, the elements selected (such as character movements, sound effects, and background music) must also be optimized based on the user's emotions.
[0905] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for recognizing user emotion data and adjusting the character model, movements, sound effects, and background music based on the emotion, and means for generating a video by integrating the character model, movements, sound effects, and background music. This makes it possible to automatically generate high-quality videos optimized for the user's emotions and situation.
[0906] "Means for receiving character settings specified by the user" refers to a function for receiving detailed information such as the character's personality, appearance, clothing, etc., entered by the user using the terminal.
[0907] The "means for generating a character model based on character settings" is a function for generating a digital three-dimensional character model based on character setting information specified by the user.
[0908] The "means for receiving plot information specified by the user" is a function for receiving detailed information about the scenario and scenes of the video input by the user.
[0909] The "means for generating character actions based on plot information" is a function for analyzing received plot information and generating character actions that conform to the scenario.
[0910] The "means for selecting sound effects based on the generated actions and environment" is a function for selecting the most appropriate sound effects according to the generated character's actions and the environment.
[0911] The "means for selecting background music based on plot information" is a function for analyzing plot information and selecting background music suitable for a scene.
[0912] "Means for recognizing a user's emotional data and adjusting a character model, movements, sound effects, and background music based on that emotion" refers to a function that analyzes a user's emotional data (voice, facial expressions, etc.) and adjusts a character's movements, facial expressions, sound effects, and background music according to that emotional state.
[0913] "A means of integrating character models, movements, sound effects, and background music to generate video" is a function that combines all of these elements to generate a single, integrated, high-quality video.
[0914] This invention is a system that generates high-quality videos based on user emotions. Specifically, the server integrates character settings, plot information, sound effects, and background music to automatically generate videos that match the user's emotions.
[0915] Program Overview
[0916] 1. Enter character settings
[0917] The user uses the device to input detailed settings for the character, such as personality, appearance, clothing, etc. For example, they might set a character with a cheerful personality, red hair, and blue clothing. This information is then sent to the server.
[0918] 2. Character Model Generation
[0919] The server analyzes the character setting information and generates a three-dimensional digital character model.
[0920] 3. Enter plot information
[0921] The user inputs the video scenario and scene details (plot information) via the device. For example, the user can specifically describe the scenario, such as "A character sits on a bench in a park." This information is then sent to the server.
[0922] 4. Behavior generation
[0923] The server analyzes the plot information and generates animations for characters to move naturally, including the character's movement path and body movements.
[0924] 5. Sound effect selection
[0925] After the movement animation is generated, the server selects sound effects that suit the movement and the surrounding environment, such as "birds chirping" for a park scene.
[0926] 6. Background Music Selection
[0927] Based on the plot information, the server selects background music that matches the scene. Music that creates the atmosphere of the scene is selected from a database.
[0928] 7. Use of Emotion Engine
[0929] The server receives data from the device to recognize the user's emotions (voice, facial expressions, behavioral logs, etc.), and the emotion engine analyzes this data. For example, if the user is laughing happily, the emotion engine will recognize a positive emotion.
[0930] 8. Change movement and facial expression
[0931] Based on the user's recognized emotional state, the server changes the character's behavior and facial expressions. For example, if the user looks sad, the character will also look sad.
[0932] 9. Change the background music
[0933] The emotion engine selects appropriate background music based on the emotion it recognizes. For example, if the user is in a relaxed state, music that creates a relaxing atmosphere will be selected.
[0934] 10. Correcting plot information
[0935] The emotion engine may also automatically modify plot information depending on the user's emotions, for example, changing the scenario to a simpler, calmer one if the user is feeling stressed.
[0936] 11. Video Integration and Generation
[0937] The server combines the character models, movements, sound effects, background music, and emotion-based modifications to generate a single video.
[0938] Hardware and software used
[0939] Hardware:
[0940] Smartphone or tablet
[0941] Camera (to capture facial expressions)
[0942] Microphone (to capture audio)
[0943] software:
[0944] EmotionRecognizer (emotion recognition library)
[0945] VideoGenerator (library for generating videos: OpenCV as an example)
[0946] Specific examples
[0947] Next, we will explain a specific example of how this system can be used. A user uses a smartphone to create a character with a cheerful personality, red hair, and blue clothes, and enters plot information such as "the character is sitting on a bench in a park." If the user is smiling, the emotion engine recognizes this positive emotion, brightens the character's facial expressions and movements, and changes the background music to a more cheerful one. As a result, a high-quality video optimized for the user is generated.
[0948] Example prompt sentence:
[0949] "Change the character's facial expression to a smile, set the background music to upbeat music, and generate a high-quality video."
[0950] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0951] Step 1:
[0952] The user inputs character setting information using a terminal. Specifically, detailed information such as the character's personality, appearance, and clothing is entered, and this information is sent from the terminal to the server. The input data could be specific information such as "a character with a cheerful personality, red hair, and blue clothing." The server then stores this input data in a database.
[0953] Step 2:
[0954] The server generates a digital 3D character model based on the received character configuration information. It uses a generative AI model to analyze this information and generate the character model using 3D modeling software (e.g., Blender). The output is a completed 3D character model.
[0955] Step 3:
[0956] The user inputs the video's scenario and scene details (plot information) via the device. This input data includes specific plot information, such as "a character sits on a bench in a park." The device then sends this information to the server, which then stores the received plot information in a database.
[0957] Step 4:
[0958] The server analyzes the plot information and generates character movement animation. For example, based on the plot information "a character sits on a bench in a park," it generates a natural movement path for the character. Here, 3D animation software (e.g., Maya) is used. The output is movement animation data.
[0959] Step 5:
[0960] The server selects sound effects based on the generated action and environment. For example, for a park scene, sound effects such as "birds chirping" and "wind sounds" are selected. A physics simulation library (e.g., Bullet Physics) is used to select appropriate sound effects. The output is the selected sound effect file.
[0961] Step 6:
[0962] The server selects background music based on the plot information. It analyzes the atmosphere and selects music from a database that matches the scene. For example, relaxing music is selected for a tranquil scene in a park. The output is the selected background music file.
[0963] Step 7:
[0964] The server analyzes the user's emotional data (voice, facial expressions, behavioral logs, etc.) received from the device to recognize the user's emotional state. It uses emotion recognition software (e.g., Facial Emotion Recognition). For example, if the user is smiling happily, the server recognizes a positive emotion. The output is the user's emotional state data.
[0965] Step 8:
[0966] The server adjusts the character's behavior, facial expressions, sound effects, and background music based on the user's recognized emotional state. For example, if the user is sad, the server changes the character's facial expression to a sad one and the background music to a calm one. The output is each element adjusted based on the emotion.
[0967] Step 9:
[0968] The server integrates character models, movements, sound effects, and background music to create a video using video editing software (e.g., Adobe Premiere Pro). The output is a final, high-quality video file.
[0969] Step 10:
[0970] The generated video is sent from the server to the user's device, where the user can watch it. The output is a video file that the user can watch.
[0971] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0972] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0973] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0974] [Third embodiment]
[0975] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0976] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0977] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0978] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0979] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0980] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0981] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0982] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0983] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0984] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0985] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0986] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0987] This invention is a system that allows users to easily generate high-quality videos, and provides a means for automatically integrating user-specified character settings, plot information, sound effects, and background music and outputting them as a single video.
[0988] Program Overview
[0989] Character setting input
[0990] Users can use their devices to enter detailed settings such as the character's personality, appearance, clothing, etc., through an interface, allowing them to specifically design their own character.
[0991] Character Model Generation
[0992] The server receives the character setting information sent from the device, analyzes it, and generates a character model, a three-dimensional digital representation of the character with an appearance and personality based on the setting.
[0993] Plot information input
[0994] The user inputs the details of the scenes (plot information) required for the video via the terminal. For example, they can specifically describe the scenario, such as "a character sitting on a bench in a park."
[0995] motion generation
[0996] The server analyzes the plot information and automatically generates character movements, including setting animations and movement paths so that the characters move naturally.
[0997] Sound effect selection
[0998] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[0999] Background music selection
[1000] Based on the plot information, the server selects background music to create an appropriate atmosphere for the scene, for example, relaxing music for a park scene.
[1001] Video Integration and Generation
[1002] The server generates the video by integrating character models, movements, sound effects, and background music, and then synchronizes all the elements to create a single, high-quality video.
[1003] Specific examples
[1004] Character Settings
[1005] The user sets up a character with a bright personality, red hair, and blue clothes on the device screen. This information is sent from the device to the server.
[1006] Character model generation
[1007] Based on the user's settings, the server generates a character model with red hair and blue clothes. This character has a cheerful personality, which is reflected in their facial expressions and movements.
[1008] Behavior Settings
[1009] When a user inputs plot information such as "a character sits on a bench in a park," the server receives this information and automatically generates a series of actions for the character to walk and sit on the bench.
[1010] Sound effects and background music selection
[1011] Using physics simulation, the server selects sound effects that match the park's ambient sounds and the characters' movements, as well as background music that creates a relaxing atmosphere.
[1012] Video Integration and Generation
[1013] The server then combines the generated character models, movements, sound effects, and background music into a single video file, which users can then download or stream to their devices.
[1014] In this way, the system of the present invention allows the user to automatically generate professional quality videos simply by configuring detailed settings.
[1015] The processing flow will be explained below.
[1016] Step 1:
[1017] The user opens the character setting screen on the device and enters detailed settings for the character, such as personality, appearance, clothing, etc. Once the input is complete, the setting information is sent to the server.
[1018] Step 2:
[1019] The server receives the character setting information sent from the device. The server analyzes this information and generates a character model with the set characteristics. The generated character model is expressed as a digital 3D model.
[1020] Step 3:
[1021] The user inputs the video scenario and scene details (plot information) on the device. For example, they specify specific locations and actions, such as "a character sitting on a bench in a park." Once input is complete, the plot information is sent to the server.
[1022] Step 4:
[1023] The server receives the plot information sent from the device. The server analyzes the plot information and generates movement animations to determine how the character should move in the specified scenario or scene. For example, it calculates and sets the movement of a character walking and sitting on a park bench.
[1024] Step 5:
[1025] After the movement animation is determined, the server selects sound effects appropriate for that movement and the surrounding environment, using physics simulation to determine, for example, the sound of a character walking or the ambient sounds of a park (birds chirping, wind noise, etc.).
[1026] Step 6:
[1027] The server selects background music (BGM) that matches the scene based on the plot information. The server selects music from a database that matches the atmosphere of the plot. For example, it selects calming BGM to match a relaxing scene in a park.
[1028] Step 7:
[1029] The server combines character models, movement animations, sound effects, and background music into a single video, precisely timing and matching each element to produce a high-quality video.
[1030] Step 8:
[1031] The server sends the generated video to the user's device, where the user can download and watch the video.
[1032] Step 9:
[1033] If a user wishes to make corrections to a video, they input the corrections from their device and send them back to the server. The server then regenerates the character model and movement animations based on the corrections, selects new sound effects and background music, and regenerates the corrected video.
[1034] Example 1
[1035] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1036] With conventional video generation systems, it was difficult for users to generate high-quality videos using the character settings and plot information they desired. Furthermore, there was a lack of means to properly integrate character movements, sound effects, and background music, which often resulted in a decline in the quality of the generated videos. There is a need for a system that can solve these problems and enable easier, high-quality video generation.
[1037] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1038] In this invention, the server includes means for receiving character settings specified by a user, means for generating a character model based on the character settings, means for receiving plot information specified by a user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for generating a video by integrating the character model, movements, sound effects, and background music, means for selecting environmental sounds and movement sounds using a physical simulation based on the generated character model and movements, means for transmitting the generated video to a user terminal, means for synchronizing the generated character model, movements, sound effects, and background music using a timeline editor, and means for automatically generating a character model and video using a generative AI model based on the character settings and plot information input by a user. This allows users to easily generate high-quality videos and enables the realization of well-integrated, professional-quality videos.
[1039] "Character configuration" is the process by which a user inputs details about a character, such as personality, appearance, and clothing.
[1040] A "character model" is a three-dimensional digital character generated based on a character setting.
[1041] "Plot information" is detailed information about the video scenario specified by the user.
[1042] "Character behavior" refers to a series of actions that a character performs based on plot information.
[1043] "Sound effects" are audio data used to add realistic sound effects to character movements and environments.
[1044] "Background music" is music selected to create the atmosphere of a scene based on plot information.
[1045] "Method of generating animation" is the process of integrating character models, movements, sound effects, and background music into a single format.
[1046] "Physics simulation" is a simulation technique for providing realistic sound effects based on generated motions and environments.
[1047] A "timeline editor" is a tool used to properly synchronize character models, movements, sound effects, and background music.
[1048] A "generative AI model" is an algorithm that uses AI technology to automatically generate character models and videos.
[1049] "User terminal" refers to the device through which a user inputs character settings and plot information.
[1050] This invention is a system that allows users to easily generate high-quality videos. It provides a means to automatically integrate user-specified character settings, plot information, sound effects, and background music and output them as a single video. The system consists of a server and a user terminal.
[1051] First, the user uses the device to input detailed settings for the character, such as personality, appearance, and clothing. Once the character settings are complete, the device sends this information to the server. For example, if the user sets a character with a cheerful personality, red hair, and blue clothing, this information is sent to the server.
[1052] The server uses a generative AI model based on the received character setting information to generate a three-dimensional digital character model. This character model reflects the character's set appearance and personality. Next, the user inputs the details of the scenes (plot information) required for the video via their device. For example, they input a scenario such as "the character sits on a bench in a park." This information is also sent from the device to the server.
[1053] The server analyzes the plot information and automatically generates character movements, including animations and movement paths that allow the character to move naturally. The server then selects sound effects based on the generated movements and the environment. Physics simulation is used to determine audio data that realistically reflects the character's movements and environmental sounds.
[1054] The server then selects background music based on the plot information. To create a suitable atmosphere for the scene, relaxing music is selected for the park scene. After these elements are gathered, the server generates a video by integrating character models, movements, sound effects, and background music. A timeline editor is used to properly align the timing of all elements to create a high-quality video.
[1055] The generated video is sent from the server to the user's device, where the user can watch it by downloading or streaming. This system allows users to automatically generate professional-quality videos by simply configuring detailed settings.
[1056] Specific examples
[1057] The user configures the character settings on the device screen. For example, they can configure a character with a cheerful personality, red hair, and blue clothes, and send this information from the device to the server. The server generates a character model with red hair and blue clothes based on the configuration information received from the user. This character has a cheerful personality, which is reflected in their facial expressions and movements. Next, the user inputs plot information such as "the character sits on a bench in a park." Based on this information, the server automatically generates a series of actions for the character to walk and sit on the bench. For this action, physical simulation is used to select sound effects that match the ambient sounds of the park and the character's movements. Background music is also selected to create a relaxing atmosphere.
[1058] Once all these elements are in place, the server combines the character movements, sound effects, and background music into a video that users can then download or stream to their device.
[1059] Prompt Sentence Examples
[1060] Character: A cheerful character with red hair and blue clothes.
[1061] Plot Info: A character sits on a bench in a park.
[1062] Sound Effects: Added ambient bird sounds and character walking sounds
[1063] Background music: Relaxing music
[1064] When this prompt is input into a generative AI model, a video that integrates all the elements is automatically generated.
[1065] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1066] Step 1:
[1067] The user inputs character setting information using the terminal. In the input form, the user sets details such as the character's personality, appearance, and clothing. For example, the user can set a "character with a cheerful personality, red hair, and blue clothes." Once this information is entered, the terminal sends the setting information to the server.
[1068] Input: Detailed setting information such as character personality, appearance, clothing, etc.
[1069] Output: Character configuration information sent to the server.
[1070] Step 2:
[1071] The server analyzes the character setting information received from the device, and the generative AI model on the server generates a three-dimensional digital character model based on this information.
[1072] Input: Character configuration information.
[1073] Output: 3D digital character model.
[1074] What it does: The generative AI model references the database and sets the character's hair color to red and their clothing to blue.
[1075] Step 3:
[1076] The user inputs the plot information required for the video via the terminal. In the input form, the user writes down the details of the scene. For example, the user enters a scenario such as "A character sits on a bench in a park." Once this information is entered, the terminal sends the plot information to the server.
[1077] Input: Scene details (plot information).
[1078] Output: Plot information sent to the server.
[1079] Step 4:
[1080] The server analyzes the plot information and automatically generates character movements, including animations and movement paths that allow the characters to move naturally.
[1081] Input: Plot information.
[1082] Output: Character animation data and movement path.
[1083] Specific Action: Based on the instruction "sit on a bench in the park," the plot analysis module generates the action of the character walking and sitting on the bench.
[1084] Step 5:
[1085] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[1086] Input: Character animation data and environment information.
[1087] Output: Suitable sound effects.
[1088] Specific operation: The sound effect selection module selects sounds such as birds singing, footsteps, and bench sounds from the sound library.
[1089] Step 6:
[1090] The server selects background music based on plot information, determining the background music to create an atmosphere that matches the scene.
[1091] Input: Plot information.
[1092] Output: Suitable background music.
[1093] Specific operation: The music selection module selects relaxing music from the music database.
[1094] Step 7:
[1095] The server uses a timeline editor to combine character models, movements, sound effects, and background music into a high-quality video, aligning all elements appropriately to create the final video.
[1096] Input: Character models, animation data, sound effects, background music.
[1097] Output: The finished video file.
[1098] Specific operation: Each element is integrated in the timeline editor, then encoded and generated into a single video.
[1099] Step 8:
[1100] The server sends the generated video to the user's device, where the user can watch the video by downloading or streaming.
[1101] Input: Your finished video file.
[1102] Output: The video sent to the user's device.
[1103] Specific operation: The server compresses the video file and sends it to the user's device. The user then downloads the video and watches it.
[1104] (Application example 1)
[1105] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1106] Conventional video generation systems require a lot of manual work from the user and specialized knowledge and skills, making it difficult for anyone to easily generate high-quality videos and share them on content distribution services. Furthermore, integrating character settings, plot information, sound effects, and background music during video generation is complicated, making it difficult to maintain professional quality.
[1107] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1108] In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for generating a video by integrating the character model, movements, sound effects, and background music, and means for uploading the generated video to a content distribution service. This allows users to easily generate high-quality videos and share them on the content distribution service.
[1109] "Character settings" are detailed setting information such as the character's personality, appearance, clothing, etc., specified by the user.
[1110] A "character model" is a three-dimensional digital representation generated based on a character's settings.
[1111] "Plot information" is detailed information about the scenes and scenarios required for the video specified by the user.
[1112] "Movement generation" is the process of automatically generating character movements based on plot information.
[1113] "Sound effects" are sound effects that are selected based on the action and environment being generated.
[1114] "Background music" is music that is selected to fit a scene based on plot information.
[1115] "Video generation" is the process of integrating character models, movements, sound effects, and background music to generate a single, high-quality video.
[1116] A "content distribution service" is a platform for sharing and viewing generated videos over the Internet.
[1117] A "user terminal" is an electronic device operated by a user, such as a smartphone or computer.
[1118] "Physics simulation" is a simulation technology that realistically reflects the movements and sound effects of generated character models.
[1119] "Smartphone Application" means a smartphone application that enables users to create, distribute, and guide their own video content.
[1120] The present invention provides a system that enables users to easily generate high-quality videos and share them via content distribution services. Specific embodiments of this system are described below.
[1121] Character setting input
[1122] The device provides users with a means to input detailed settings such as the character's personality, appearance, clothing, etc., allowing them to specifically design their own character.
[1123] Character Model Generation
[1124] The server receives the character setting information sent from the device, analyzes it, and generates a character model, a three-dimensional digital representation with an appearance and personality based on the setting.
[1125] Plot information input
[1126] Users are provided with a means to input details of scenes (plot information) required for the video via their device. For example, they can specifically describe a scenario such as "a character sitting on a bench in a park."
[1127] motion generation
[1128] The server analyzes plot information and automatically generates character movements. This movement generation includes setting animations and movement paths so that characters move naturally. Furthermore, it can use physics simulation to reflect realistic movements.
[1129] Sound effect selection
[1130] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[1131] Background music selection
[1132] Based on the plot information, the server selects background music to create an appropriate atmosphere for the scene, for example, relaxing music for a park scene.
[1133] Video Integration and Generation
[1134] The server then combines the character models, movements, sound effects, and background music to create a video, syncing all the elements together to create a single, high-quality video that is then uploaded to a content distribution service.
[1135] Sending to user terminal
[1136] The generated video is sent to the user's device, and the user can watch the video on their device by downloading or streaming it.
[1137] Hardware and software used
[1138] The implementation of this system uses high-performance cloud servers (e.g., Amazon EC2 or Google Cloud Platform) and the Python library aiohttp, an asynchronous HTTP client, with the following APIs:
[1139] Character Model Generation API
[1140] Plotting behavior generation API
[1141] Sound Effect Selection API
[1142] Background music selection API
[1143] Video Generation API
[1144] Specific examples
[1145] A user uses a smartphone application to enter the following information:
[1146] Character description: "A character with a cheerful personality, red hair, and blue clothes."
[1147] Plot information: "A character sits on a bench in a park."
[1148] Sound effects: park ambient sounds, character footsteps
[1149] Background music: Relaxing music
[1150] Based on this information, the server integrates the character's 3D model, movements, sound effects, and background music to generate a high-quality video, which is then uploaded to a content distribution service.
[1151] Prompt Sentence Examples
[1152] "Please create a scene where the user creates a character with a cheerful personality, red hair, and blue clothes, sitting on a bench in a park. Use the ambient sounds of the park and the character's footsteps as sound effects, and choose a relaxing song for the background music."
[1153] The system allows users to easily generate high-quality videos and share them on content distribution services.
[1154] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1155] Step 1:
[1156] The user inputs character settings using a terminal. Specifically, the user enters detailed information about the character, such as personality, appearance, and clothing, into an input form and sends it to the server. The input data includes the character's name, hair color, clothing color, personality, etc. The server analyzes this data and prepares it.
[1157] Step 2:
[1158] The server receives the character setting information sent by the user and generates a 3D model of the character based on that information. The server uses an AI model to generate a character model with the specified characteristics and stores it in an internal database. The generated character model has the characteristics specified by the user, such as appearance and personality.
[1159] Step 3:
[1160] The user inputs plot information using the device, specifically details of the scenes required for the video (e.g., "A character sits on a bench in a park"), and the device sends this plot information to the server.
[1161] Step 4:
[1162] The server analyzes the received plot information and automatically generates character movements. The server combines AI models and physics simulations to generate natural movements and movement paths. For example, based on the input plot information, it generates a series of movements for a character to walk and sit on a bench.
[1163] Step 5:
[1164] The server selects appropriate sound effects based on the generated actions and environment. Specifically, it selects sound effects that correspond to the character's actions (e.g., walking, sitting, etc.) and environmental sounds (e.g., birds chirping in the park). It then extracts appropriate audio files from the audio database and synchronizes them with the actions.
[1165] Step 6:
[1166] The server selects background music based on the plot information. Specifically, it selects music with a relaxing atmosphere to match the scene. If necessary, it adjusts the tempo and atmosphere of the music to match the plot information.
[1167] Step 7:
[1168] The server then combines the generated character models, movements, sound effects, and background music to generate a video, adjusting the timing of all elements and stitching them together into a single, high-quality video. The video file is generated internally and can later be downloaded or streamed.
[1169] Step 8:
[1170] The server uploads the generated video to a content distribution service, where users can access the video over the Internet and share it with other users.
[1171] Step 9:
[1172] The server then sends the generated video to the user's device, where the user can receive it and save it on their device or watch it via streaming. Communication between the server and the device is encrypted to ensure secure data transfer.
[1173] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1174] This invention is a system that allows users to easily generate high-quality videos, and provides a means to automatically integrate user-specified character settings, plot information, sound effects, and background music and output them as a single video. This system also incorporates an emotion engine that recognizes the user's emotions and modifies or changes the character's movements, facial expressions, background music, and plot information based on those emotions.
[1175] Program Overview
[1176] Character setting input
[1177] Users use the device to input detailed settings such as the character's personality, appearance, clothing, etc., allowing them to design a character to their liking.
[1178] Character Model Generation
[1179] The server receives the character setting information sent from the device, analyzes it, and generates a character model, which is generated as a three-dimensional digital representation.
[1180] Plot information input
[1181] The user inputs the video scenario and scene details (plot information) via the device. For example, the user can specifically describe the scenario, such as "a character sits on a bench in a park."
[1182] motion generation
[1183] The server analyzes the plot information and generates movement animations that determine how the characters will move in a given scenario or scene, including setting up animations and movement paths for the characters to move naturally.
[1184] Sound effect selection
[1185] After the action animation is determined, the server selects sound effects that are appropriate for the action and the surrounding environment.The server uses physics simulation to determine the sound effects that correspond to the environment and action.
[1186] Background music selection
[1187] Based on the plot information, the server selects background music that matches the scene, selecting music from a database that matches the atmosphere of the plot and creating the atmosphere of the scene.
[1188] Using the Emotion Engine
[1189] The server receives data from the device to recognize the user's emotions (e.g., voice, facial expressions, behavior logs, etc.). The emotion engine analyzes this data and recognizes the user's emotional state. For example, if the user is laughing happily, the emotion engine recognizes a positive emotion.
[1190] Change movements and expressions
[1191] Based on the user's recognized emotional state, the server changes the character's behavior and facial expression. For example, if the user looks sad, the character's facial expression and behavior will also change to look sad.
[1192] Change background music
[1193] Based on the recognized emotion, the server selects appropriate background music. For example, if the user is in a relaxed state, background music will be selected to create a relaxing atmosphere throughout the scene.
[1194] Modifying plot information
[1195] The emotion engine may automatically modify plot information depending on the user's emotions. For example, if the user is feeling stressed, the scenario may be changed to a simpler, calmer one.
[1196] Video Integration and Generation
[1197] The server combines all elements—character models, movements, sound effects, background music, and emotion-based changes—to generate a single video, resulting in a high-quality video optimized for the user's emotions.
[1198] Specific examples
[1199] Character Settings
[1200] When a user selects a character with a bright personality, red hair, and blue clothes on their device screen, that information is sent to the server, which then generates a character model based on that information.
[1201] Input of plot information and generation of motion
[1202] Users input plot information, such as "a character sits on a bench in the park," which is sent to a server that analyzes the information and generates natural-looking movements for the character.
[1203] Using emotion engines and video generation
[1204] When a user watches a video, the device sends emotional data to the server. For example, if the user is smiling, the emotion engine recognizes this positive emotion and changes the character's facial expressions and movements to a brighter tone and the background music to match. The server combines these elements to generate a video, which is then sent to the device. The user can then download and watch the video.
[1205] In this way, the system of the present invention allows the user to automatically generate high-quality videos that are optimized to suit the user's emotions simply by making simple, detailed settings.
[1206] The processing flow will be explained below.
[1207] Step 1:
[1208] The user opens the character setting screen on the device and enters detailed settings for the character, such as personality, appearance, clothing, etc. Once the input is complete, the setting information is sent to the server.
[1209] Step 2:
[1210] The server receives the character setting information sent from the device. The server analyzes this information and generates a character model with the set characteristics. The generated character model is expressed as a digital 3D model.
[1211] Step 3:
[1212] The user inputs the video scenario and scene details (plot information) on the device. For example, they specify specific locations and actions, such as "a character sitting on a bench in a park." Once input is complete, the plot information is sent to the server.
[1213] Step 4:
[1214] The server receives the plot information sent from the device. The server analyzes the plot information and generates movement animations to determine how the character should move in the specified scenario or scene. For example, it calculates and sets the movement of a character walking and sitting on a park bench.
[1215] Step 5:
[1216] After the movement animation is determined, the server selects sound effects appropriate for that movement and the surrounding environment, using physics simulation to determine, for example, the sound of a character walking or the ambient sounds of a park (birds chirping, wind noise, etc.).
[1217] Step 6:
[1218] The server selects background music (BGM) that matches the scene based on the plot information. The server selects music from a database that matches the atmosphere of the plot. For example, it selects calming BGM to match a relaxing scene in a park.
[1219] Step 7:
[1220] The device collects the user's emotional data (voice, facial expression, behavior log, etc.) and sends it to the server. For example, the device's camera and microphone can be used to analyze the user's facial expression and tone of voice.
[1221] Step 8:
[1222] The server analyzes the received emotion data using an emotion engine to identify the user's emotional state. For example, if the user is smiling, it recognizes a positive emotion.
[1223] Step 9:
[1224] Based on the user's recognized emotions, the server changes the character's movements and facial expressions. For example, if the user is smiling, the character will also smile.
[1225] Step 10:
[1226] Based on the user's recognized emotion, the server changes the background music that is set. For example, if the user is relaxed, it selects a calm background music that matches the user's emotion.
[1227] Step 11:
[1228] The server combines all the elements—character models, movements, sound effects, background music, and emotion-based modifications—to generate a single video.
[1229] Step 12:
[1230] The server sends the generated video to the user's device, where the user can download and watch the video.
[1231] Step 13:
[1232] If a user wishes to make corrections to a video, they input the corrections from their device and send them back to the server. The server then regenerates the character model and movement animations based on the corrections, selects new sound effects and background music, and regenerates the corrected video.
[1233] Example 2
[1234] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1235] Conventional video creation systems have the problem that it is difficult for users to customize based on their emotions, and the quality of the generated videos is not optimized for the user's emotions and preferences. In addition, adjusting each element (character settings, plot information, sound effects, background music) individually requires a great deal of effort and expertise, making it difficult to easily generate high-quality videos.
[1236] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1237] In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for receiving data for recognizing the user's emotion from the terminal, means for changing the character's movements and facial expressions based on the recognized emotional state of the user, means for selecting appropriate background music based on the recognized emotion, means for automatically correcting the plot information according to the user's emotional state, and means for generating a video by integrating the character model, movements, sound effects, and background music. This enables users to easily generate high-quality videos optimized to their own emotions.
[1238] "Users" are individuals or organizations that operate the system and input character settings and plot information.
[1239] The "server" is a central processing unit that receives and analyzes data sent by users and generates character models, movements, sound effects, background music, etc.
[1240] "Character settings" refers to detailed information about a character's personality, appearance, clothing, etc.
[1241] A "character model" is a three-dimensional digital representation generated based on a character's settings.
[1242] "Plot information" is information about the scenario and scene details of a video, and is specified by the user.
[1243] "Movement animation" refers to the specific movements and actions that characters perform based on plot information.
[1244] "Sound effects" are sound effects that are appropriate for the character's movements and environment, and are selected based on physical models.
[1245] "Background music" is music used to create the overall atmosphere of a video scene.
[1246] The "emotion engine" is a technology that analyzes user emotions and customizes various elements of a video based on those emotions.
[1247] "Emotion recognition data" is data used to recognize a user's emotions, and includes voice, facial expressions, behavioral logs, etc.
[1248] "Video generation means" is a mechanism for integrating character models, movements, sound effects, and background music to generate a single video.
[1249] This invention is a system that allows users to easily generate high-quality videos. Users input character settings, plot information, sound effects, and background music using a terminal, and the system automatically integrates them and outputs them as a single video. The system also uses an emotion engine that recognizes the user's emotions and modifies or changes the character's movements, facial expressions, background music, and plot information based on those emotions.
[1250] This system operates in cooperation with both the server and the terminal. The main processing steps and related hardware and software are explained below.
[1251] Character setting input
[1252] Using the device interface, users input detailed settings for their character, such as personality, appearance, clothing, etc. This information is sent from the device to a server, which then analyzes the character settings information.
[1253] Character Model Generation
[1254] The server generates a character model based on the received character setting information. Specifically, it uses 3D modeling software such as Blender or Unity. This character model is then used for subsequent movement generation and animation.
[1255] Plot information input
[1256] The user inputs the video scenario and scene details (plot information) via the terminal. For example, the user inputs a scenario such as "A character sits on a bench in a park." This information is also sent to the server and stored as plot information.
[1257] motion generation
[1258] The server analyzes the plot information and determines how the characters will move in the specified scenario or scene. Animation software such as Maya or Autodesk MotionBuilder is used for movement animation. The generated movements include animations and movement paths for the characters to move naturally.
[1259] Sound effect selection
[1260] After the movement animation is determined, the server selects sound effects appropriate for the movement and environment, using sound effect generation tools such as FMOD and Wwise with physical simulation.
[1261] Background music selection
[1262] Based on the plot information, the server selects background music that matches the scene, selected from music libraries such as Epidemic Sound and Artlist.
[1263] Using the Emotion Engine
[1264] To recognize the user's emotions, the device uses a camera and microphone to collect the user's facial and voice data and sends it to the server. The server then uses an emotion engine (such as OpenCV or Affectiva) to analyze the user's emotional state. For example, if the user is smiling, it is recognized as a positive emotion.
[1265] Change movements and expressions
[1266] The server changes the character's behavior and facial expression based on the user's recognized emotional state. For example, if the user is sad, the character's facial expression and behavior will also change to a sad one.
[1267] Change background music
[1268] The server selects appropriate background music based on the recognized emotion, for example, if the user is relaxed, calm background music is selected.
[1269] Modifying plot information
[1270] The emotion engine automatically modifies plot information according to the user's emotions, so if the user is feeling stressed, the scenario will be changed to a simpler, calmer one.
[1271] Video Integration and Generation
[1272] The server then combines all elements based on the character model, movements, sound effects, background music, and emotional changes into a single video. The final high-quality video is then generated using video editing software such as Adobe After Effects or Final Cut Pro. The resulting video is then sent to the user's device, where they can download and watch it.
[1273] Examples of prompt statements
[1274] Here are some examples of prompts to input to a generative AI model:
[1275] "Create a scene in a park where a character with red hair and a cheerful personality is sitting on a bench. Not only should the character be smiling and sitting on the bench, but also integrate relaxing background music and natural environmental sounds. Also, dynamically change the character's facial expressions and behavior based on emotion recognition while the user is viewing the scene."
[1276] In this way, the system of the present invention allows the user to automatically generate high-quality videos that are optimized to suit the user's emotions simply by making simple, detailed settings.
[1277] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1278] Step 1:
[1279] Entering character settings
[1280] The user sets the character's personality, appearance, clothing, etc. on the terminal screen.
[1281] Input: Enter character settings information via device (e.g., "cheerful personality, red hair, blue clothes")
[1282] Processing: The terminal formats the entered information and sends it to the server.
[1283] Specific operation: Enter the setting information through the terminal's GUI and press the "Send" button, and the information will be sent to the server via API.
[1284] Output: Character setting information sent to the server
[1285] Step 2:
[1286] Character model generation
[1287] The server receives the character setting information sent from the terminal and analyzes it.
[1288] Input: Character information (e.g., "cheerful personality, red hair, blue clothes")
[1289] Process: The server uses the Blender API to generate the character model.
[1290] Specific operation: Passes configuration information to the Blender API and executes a script to generate a 3D digital model.
[1291] Output: Generated character model (3D digital model)
[1292] Step 3:
[1293] Entering plot information
[1294] The user inputs the video scenario and scene details (plot information) via the terminal.
[1295] Input: Plot information (e.g., "A character sits on a bench in a park")
[1296] Processing: The terminal formats the entered plot information and sends it to the server.
[1297] Specific operation: Plot information is entered using the GUI, and the information is sent to the server when the "Send" button is pressed.
[1298] Output: Plot information sent to the server
[1299] Step 4:
[1300] Generate movement animation
[1301] The server analyzes the plot information and generates animations of the characters' movements.
[1302] Input: Plot information (e.g., "A character sits on a bench in a park")
[1303] Processing: Generate movement animation using Maya or Autodesk MotionBuilder.
[1304] Specific operation: The server passes the analyzed plot information to the Maya API and executes a script that generates a movement animation.
[1305] Output: Generated movement animation
[1306] Step 5:
[1307] Sound Effect Selection
[1308] The server selects the appropriate sound effect according to the action animation.
[1309] Input: Action animation (e.g. "Sit on a bench")
[1310] Processing: Select sound effects based on physical simulation using FMOD and Wwise.
[1311] Specific behavior: The behavior animation data is passed to the FMOD API, and an algorithm is run to select the appropriate sound effect.
[1312] Output: Selected sound effects
[1313] Step 6:
[1314] Background music selection
[1315] The server selects background music that matches the scene based on the plot information.
[1316] Input: Plot information (e.g., "A character sits on a bench in a park")
[1317] Processing: Select music that matches the scene from the Epidemic Sound and Artlist databases.
[1318] What it does: It parses the plot information and runs an algorithm that searches the music library using the corresponding keywords.
[1319] Output: Selected background music
[1320] Step 7:
[1321] Emotion engine recognizes user emotions
[1322] To recognize the user's emotions, the device uses a camera and microphone to collect facial expressions and voice sounds, which are then sent to a server.
[1323] Input: User's facial expression data and voice data
[1324] Processing: The device formats the collected data and sends it to the server.
[1325] Specific operation: Collects data in real time through the camera and microphone and sends it to the server.
[1326] Output: Emotion recognition data sent to the server
[1327] Step 8:
[1328] Emotion analysis using an emotion engine
[1329] The server uses an emotion engine to recognize the user's emotional state.
[1330] Input: Emotion recognition data (e.g., "image of a smiling face")
[1331] Processing: Sentiment analysis is performed using Affectiva and OpenCV.
[1332] Specific operation: Runs emotion recognition algorithm and obtains emotion analysis result (e.g., positive emotion).
[1333] Output: Emotion analysis results
[1334] Step 9:
[1335] Change movements and expressions
[1336] The server changes the character's movements and facial expressions based on the results of emotion analysis.
[1337] Input: Sentiment analysis result (e.g., "positive sentiment")
[1338] Processing: Updates animation data to change the character's movements and expressions.
[1339] Specific actions: Regenerate character movements and facial expressions using Maya API and Blender API.
[1340] Output: Modified character movements and expressions
[1341] Step 10:
[1342] Change background music
[1343] The server selects appropriate background music based on the emotion analysis results.
[1344] Input: Sentiment analysis result (e.g., "positive sentiment")
[1345] Processing: Select new background music from Epidemic Sound and Artlist.
[1346] What it does: Search your music library using new keywords to find the right background music.
[1347] Output: Modified background music
[1348] Step 11:
[1349] Modifying plot information
[1350] The server automatically modifies plot information according to the user's emotional state.
[1351] Input: Sentiment analysis results (e.g., "I feel stressed")
[1352] Processing: Reanalyze the plot information and change it into a simple and calm scenario.
[1353] Specific operation: Runs an algorithm that recompiles plot information and generates a revised scenario.
[1354] Output: Corrected plot information
[1355] Step 12:
[1356] Video Integration and Generation
[1357] The server combines all elements (character models, movements, sound effects, background music, and emotional changes) to generate the video.
[1358] Input: Character models, movements, sound effects, background music, modified plot information
[1359] Processing: Edit the video using Adobe After Effects and Final Cut Pro.
[1360] What it does: Place elements on a timeline in video editing software, add effects and transitions, and generate the final video.
[1361] Output: High-quality generated video
[1362] Step 13:
[1363] Sending the generated video
[1364] The server transmits the generated high-quality video to the terminal.
[1365] Input: Generated video
[1366] Processing: Converts the video format and encodes it to prepare it for sending to the device.
[1367] Specific operation: The video is converted into the appropriate format and sent to the device using the HTTP / HTTPS protocol.
[1368] Output: Video sent to device
[1369] Step 14:
[1370] Download and watch videos
[1371] Users download and watch the videos generated on their devices.
[1372] Input: Video sent from the server
[1373] Processing: Receive, store, and play the video on the device.
[1374] Specific operation: Play the video using the device's media player.
[1375] Output: Viewable video
[1376] Through this series of processing steps, users can easily generate and watch high-quality, emotion-optimized videos.
[1377] (Application example 2)
[1378] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1379] When generating videos, conventional systems have difficulty automatically adjusting content based on the user's emotions. This has led to a demand for technology that can automatically generate high-quality videos optimized for the user's emotions and situation. Furthermore, systems that can utilize emotional data to provide content that is more attuned to the user are also needed. Furthermore, to ensure that the content of the generated videos is natural, the elements selected (such as character movements, sound effects, and background music) must also be optimized based on the user's emotions.
[1380] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for recognizing user emotion data and adjusting the character model, movements, sound effects, and background music based on the emotion, and means for generating a video by integrating the character model, movements, sound effects, and background music. This makes it possible to automatically generate high-quality videos optimized for the user's emotions and situation.
[1381] "Means for receiving character settings specified by the user" refers to a function for receiving detailed information such as the character's personality, appearance, clothing, etc., entered by the user using the terminal.
[1382] The "means for generating a character model based on character settings" is a function for generating a digital three-dimensional character model based on character setting information specified by the user.
[1383] The "means for receiving plot information specified by the user" is a function for receiving detailed information about the scenario and scenes of the video input by the user.
[1384] The "means for generating character actions based on plot information" is a function for analyzing received plot information and generating character actions that conform to the scenario.
[1385] The "means for selecting sound effects based on the generated actions and environment" is a function for selecting the most appropriate sound effects according to the generated character's actions and the environment.
[1386] The "means for selecting background music based on plot information" is a function for analyzing plot information and selecting background music suitable for a scene.
[1387] "Means for recognizing a user's emotional data and adjusting a character model, movements, sound effects, and background music based on that emotion" refers to a function that analyzes a user's emotional data (voice, facial expressions, etc.) and adjusts a character's movements, facial expressions, sound effects, and background music according to that emotional state.
[1388] "A means of integrating character models, movements, sound effects, and background music to generate video" is a function that combines all of these elements to generate a single, integrated, high-quality video.
[1389] This invention is a system that generates high-quality videos based on user emotions. Specifically, the server integrates character settings, plot information, sound effects, and background music to automatically generate videos that match the user's emotions.
[1390] Program Overview
[1391] 1. Enter character settings
[1392] The user uses the device to input detailed settings for the character, such as personality, appearance, clothing, etc. For example, they might set a character with a cheerful personality, red hair, and blue clothing. This information is then sent to the server.
[1393] 2. Character Model Generation
[1394] The server analyzes the character setting information and generates a three-dimensional digital character model.
[1395] 3. Enter plot information
[1396] The user inputs the video scenario and scene details (plot information) via the device. For example, the user can specifically describe the scenario, such as "A character sits on a bench in a park." This information is then sent to the server.
[1397] 4. Behavior generation
[1398] The server analyzes the plot information and generates animations for characters to move naturally, including the character's movement path and body movements.
[1399] 5. Sound effect selection
[1400] After the movement animation is generated, the server selects sound effects that suit the movement and the surrounding environment, such as "birds chirping" for a park scene.
[1401] 6. Background Music Selection
[1402] Based on the plot information, the server selects background music that matches the scene. Music that creates the atmosphere of the scene is selected from a database.
[1403] 7. Use of Emotion Engine
[1404] The server receives data from the device to recognize the user's emotions (voice, facial expressions, behavioral logs, etc.), and the emotion engine analyzes this data. For example, if the user is laughing happily, the emotion engine will recognize a positive emotion.
[1405] 8. Change movement and facial expression
[1406] Based on the user's recognized emotional state, the server changes the character's behavior and facial expressions. For example, if the user looks sad, the character will also look sad.
[1407] 9. Change the background music
[1408] The emotion engine selects appropriate background music based on the emotion it recognizes. For example, if the user is in a relaxed state, music that creates a relaxing atmosphere will be selected.
[1409] 10. Correcting plot information
[1410] The emotion engine may also automatically modify plot information depending on the user's emotions, for example, changing the scenario to a simpler, calmer one if the user is feeling stressed.
[1411] 11. Video Integration and Generation
[1412] The server combines the character models, movements, sound effects, background music, and emotion-based modifications to generate a single video.
[1413] Hardware and software used
[1414] Hardware:
[1415] Smartphone or tablet
[1416] Camera (to capture facial expressions)
[1417] Microphone (to capture audio)
[1418] software:
[1419] EmotionRecognizer (emotion recognition library)
[1420] VideoGenerator (library for generating videos: OpenCV as an example)
[1421] Specific examples
[1422] Next, we will explain a specific example of how this system can be used. A user uses a smartphone to create a character with a cheerful personality, red hair, and blue clothes, and enters plot information such as "the character is sitting on a bench in a park." If the user is smiling, the emotion engine recognizes this positive emotion, brightens the character's facial expressions and movements, and changes the background music to a more cheerful one. As a result, a high-quality video optimized for the user is generated.
[1423] Example prompt sentence:
[1424] "Change the character's facial expression to a smile, set the background music to upbeat music, and generate a high-quality video."
[1425] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1426] Step 1:
[1427] The user inputs character setting information using a terminal. Specifically, detailed information such as the character's personality, appearance, and clothing is entered, and this information is sent from the terminal to the server. The input data could be specific information such as "a character with a cheerful personality, red hair, and blue clothing." The server then stores this input data in a database.
[1428] Step 2:
[1429] The server generates a digital 3D character model based on the received character configuration information. It uses a generative AI model to analyze this information and generate the character model using 3D modeling software (e.g., Blender). The output is a completed 3D character model.
[1430] Step 3:
[1431] The user inputs the video's scenario and scene details (plot information) via the device. This input data includes specific plot information, such as "a character sits on a bench in a park." The device then sends this information to the server, which then stores the received plot information in a database.
[1432] Step 4:
[1433] The server analyzes the plot information and generates character movement animation. For example, based on the plot information "a character sits on a bench in a park," it generates a natural movement path for the character. Here, 3D animation software (e.g., Maya) is used. The output is movement animation data.
[1434] Step 5:
[1435] The server selects sound effects based on the generated action and environment. For example, for a park scene, sound effects such as "birds chirping" and "wind sounds" are selected. A physics simulation library (e.g., Bullet Physics) is used to select appropriate sound effects. The output is the selected sound effect file.
[1436] Step 6:
[1437] The server selects background music based on the plot information. It analyzes the atmosphere and selects music from a database that matches the scene. For example, relaxing music is selected for a tranquil scene in a park. The output is the selected background music file.
[1438] Step 7:
[1439] The server analyzes the user's emotional data (voice, facial expressions, behavioral logs, etc.) received from the device to recognize the user's emotional state. It uses emotion recognition software (e.g., Facial Emotion Recognition). For example, if the user is smiling happily, the server recognizes a positive emotion. The output is the user's emotional state data.
[1440] Step 8:
[1441] The server adjusts the character's behavior, facial expressions, sound effects, and background music based on the user's recognized emotional state. For example, if the user is sad, the server changes the character's facial expression to a sad one and the background music to a calm one. The output is each element adjusted based on the emotion.
[1442] Step 9:
[1443] The server integrates character models, movements, sound effects, and background music to create a video using video editing software (e.g., Adobe Premiere Pro). The output is a final, high-quality video file.
[1444] Step 10:
[1445] The generated video is sent from the server to the user's device, where the user can watch it. The output is a video file that the user can watch.
[1446] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1447] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1448] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1449] [Fourth embodiment]
[1450] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1451] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1452] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1453] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1454] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1455] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1456] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1457] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1458] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1459] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1460] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1461] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1462] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1463] This invention is a system that allows users to easily generate high-quality videos, and provides a means for automatically integrating user-specified character settings, plot information, sound effects, and background music and outputting them as a single video.
[1464] Program Overview
[1465] Character setting input
[1466] Users can use their devices to enter detailed settings such as the character's personality, appearance, clothing, etc., through an interface, allowing them to specifically design their own character.
[1467] Character Model Generation
[1468] The server receives the character setting information sent from the device, analyzes it, and generates a character model, a three-dimensional digital representation of the character with an appearance and personality based on the setting.
[1469] Plot information input
[1470] The user inputs the details of the scenes (plot information) required for the video via the terminal. For example, they can specifically describe the scenario, such as "a character sitting on a bench in a park."
[1471] motion generation
[1472] The server analyzes the plot information and automatically generates character movements, including setting animations and movement paths so that the characters move naturally.
[1473] Sound effect selection
[1474] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[1475] Background music selection
[1476] Based on the plot information, the server selects background music to create an appropriate atmosphere for the scene, for example, relaxing music for a park scene.
[1477] Video Integration and Generation
[1478] The server generates the video by integrating character models, movements, sound effects, and background music, and then synchronizes all the elements to create a single, high-quality video.
[1479] Specific examples
[1480] Character Settings
[1481] The user sets up a character with a bright personality, red hair, and blue clothes on the device screen. This information is sent from the device to the server.
[1482] Character model generation
[1483] Based on the user's settings, the server generates a character model with red hair and blue clothes. This character has a cheerful personality, which is reflected in their facial expressions and movements.
[1484] Behavior Settings
[1485] When a user inputs plot information such as "a character sits on a bench in a park," the server receives this information and automatically generates a series of actions for the character to walk and sit on the bench.
[1486] Sound effects and background music selection
[1487] Using physics simulation, the server selects sound effects that match the park's ambient sounds and the characters' movements, as well as background music that creates a relaxing atmosphere.
[1488] Video Integration and Generation
[1489] The server then combines the generated character models, movements, sound effects, and background music into a single video file, which users can then download or stream to their devices.
[1490] In this way, the system of the present invention allows the user to automatically generate professional quality videos simply by configuring detailed settings.
[1491] The processing flow will be explained below.
[1492] Step 1:
[1493] The user opens the character setting screen on the device and enters detailed settings for the character, such as personality, appearance, clothing, etc. Once the input is complete, the setting information is sent to the server.
[1494] Step 2:
[1495] The server receives the character setting information sent from the device. The server analyzes this information and generates a character model with the set characteristics. The generated character model is expressed as a digital 3D model.
[1496] Step 3:
[1497] The user inputs the video scenario and scene details (plot information) on the device. For example, they specify specific locations and actions, such as "a character sitting on a bench in a park." Once input is complete, the plot information is sent to the server.
[1498] Step 4:
[1499] The server receives the plot information sent from the device. The server analyzes the plot information and generates movement animations to determine how the character should move in the specified scenario or scene. For example, it calculates and sets the movement of a character walking and sitting on a park bench.
[1500] Step 5:
[1501] After the movement animation is determined, the server selects sound effects appropriate for that movement and the surrounding environment, using physics simulation to determine, for example, the sound of a character walking or the ambient sounds of a park (birds chirping, wind noise, etc.).
[1502] Step 6:
[1503] The server selects background music (BGM) that matches the scene based on the plot information. The server selects music from a database that matches the atmosphere of the plot. For example, it selects calming BGM to match a relaxing scene in a park.
[1504] Step 7:
[1505] The server combines character models, movement animations, sound effects, and background music into a single video, precisely timing and matching each element to produce a high-quality video.
[1506] Step 8:
[1507] The server sends the generated video to the user's device, where the user can download and watch the video.
[1508] Step 9:
[1509] If a user wishes to make corrections to a video, they input the corrections from their device and send them back to the server. The server then regenerates the character model and movement animations based on the corrections, selects new sound effects and background music, and regenerates the corrected video.
[1510] Example 1
[1511] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1512] With conventional video generation systems, it was difficult for users to generate high-quality videos using the character settings and plot information they desired. Furthermore, there was a lack of means to properly integrate character movements, sound effects, and background music, which often resulted in a decline in the quality of the generated videos. There is a need for a system that can solve these problems and enable easier, high-quality video generation.
[1513] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1514] In this invention, the server includes means for receiving character settings specified by a user, means for generating a character model based on the character settings, means for receiving plot information specified by a user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for generating a video by integrating the character model, movements, sound effects, and background music, means for selecting environmental sounds and movement sounds using a physical simulation based on the generated character model and movements, means for transmitting the generated video to a user terminal, means for synchronizing the generated character model, movements, sound effects, and background music using a timeline editor, and means for automatically generating a character model and video using a generative AI model based on the character settings and plot information input by a user. This allows users to easily generate high-quality videos and enables the realization of well-integrated, professional-quality videos.
[1515] "Character configuration" is the process by which a user inputs details about a character, such as personality, appearance, and clothing.
[1516] A "character model" is a three-dimensional digital character generated based on a character setting.
[1517] "Plot information" is detailed information about the video scenario specified by the user.
[1518] "Character behavior" refers to a series of actions that a character performs based on plot information.
[1519] "Sound effects" are audio data used to add realistic sound effects to character movements and environments.
[1520] "Background music" is music selected to create the atmosphere of a scene based on plot information.
[1521] "Method of generating animation" is the process of integrating character models, movements, sound effects, and background music into a single format.
[1522] "Physics simulation" is a simulation technique for providing realistic sound effects based on generated motions and environments.
[1523] A "timeline editor" is a tool used to properly synchronize character models, movements, sound effects, and background music.
[1524] A "generative AI model" is an algorithm that uses AI technology to automatically generate character models and videos.
[1525] "User terminal" refers to the device through which a user inputs character settings and plot information.
[1526] This invention is a system that allows users to easily generate high-quality videos. It provides a means to automatically integrate user-specified character settings, plot information, sound effects, and background music and output them as a single video. The system consists of a server and a user terminal.
[1527] First, the user uses the device to input detailed settings for the character, such as personality, appearance, and clothing. Once the character settings are complete, the device sends this information to the server. For example, if the user sets a character with a cheerful personality, red hair, and blue clothing, this information is sent to the server.
[1528] The server uses a generative AI model based on the received character setting information to generate a three-dimensional digital character model. This character model reflects the character's set appearance and personality. Next, the user inputs the details of the scenes (plot information) required for the video via their device. For example, they input a scenario such as "the character sits on a bench in a park." This information is also sent from the device to the server.
[1529] The server analyzes the plot information and automatically generates character movements, including animations and movement paths that allow the character to move naturally. The server then selects sound effects based on the generated movements and the environment. Physics simulation is used to determine audio data that realistically reflects the character's movements and environmental sounds.
[1530] The server then selects background music based on the plot information. To create a suitable atmosphere for the scene, relaxing music is selected for the park scene. After these elements are gathered, the server generates a video by integrating character models, movements, sound effects, and background music. A timeline editor is used to properly align the timing of all elements to create a high-quality video.
[1531] The generated video is sent from the server to the user's device, where the user can watch it by downloading or streaming. This system allows users to automatically generate professional-quality videos by simply configuring detailed settings.
[1532] Specific examples
[1533] The user configures the character settings on the device screen. For example, they can configure a character with a cheerful personality, red hair, and blue clothes, and send this information from the device to the server. The server generates a character model with red hair and blue clothes based on the configuration information received from the user. This character has a cheerful personality, which is reflected in their facial expressions and movements. Next, the user inputs plot information such as "the character sits on a bench in a park." Based on this information, the server automatically generates a series of actions for the character to walk and sit on the bench. For this action, physical simulation is used to select sound effects that match the ambient sounds of the park and the character's movements. Background music is also selected to create a relaxing atmosphere.
[1534] Once all these elements are in place, the server combines the character movements, sound effects, and background music into a video that users can then download or stream to their device.
[1535] Prompt Sentence Examples
[1536] Character: A cheerful character with red hair and blue clothes.
[1537] Plot Info: A character sits on a bench in a park.
[1538] Sound Effects: Added ambient bird sounds and character walking sounds
[1539] Background music: Relaxing music
[1540] When this prompt is input into a generative AI model, a video that integrates all the elements is automatically generated.
[1541] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1542] Step 1:
[1543] The user inputs character setting information using the terminal. In the input form, the user sets details such as the character's personality, appearance, and clothing. For example, the user can set a "character with a cheerful personality, red hair, and blue clothes." Once this information is entered, the terminal sends the setting information to the server.
[1544] Input: Detailed setting information such as character personality, appearance, clothing, etc.
[1545] Output: Character configuration information sent to the server.
[1546] Step 2:
[1547] The server analyzes the character setting information received from the device, and the generative AI model on the server generates a three-dimensional digital character model based on this information.
[1548] Input: Character configuration information.
[1549] Output: 3D digital character model.
[1550] What it does: The generative AI model references the database and sets the character's hair color to red and their clothing to blue.
[1551] Step 3:
[1552] The user inputs the plot information required for the video via the terminal. In the input form, the user writes down the details of the scene. For example, the user enters a scenario such as "A character sits on a bench in a park." Once this information is entered, the terminal sends the plot information to the server.
[1553] Input: Scene details (plot information).
[1554] Output: Plot information sent to the server.
[1555] Step 4:
[1556] The server analyzes the plot information and automatically generates character movements, including animations and movement paths that allow the characters to move naturally.
[1557] Input: Plot information.
[1558] Output: Character animation data and movement path.
[1559] Specific Action: Based on the instruction "sit on a bench in the park," the plot analysis module generates the action of the character walking and sitting on the bench.
[1560] Step 5:
[1561] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[1562] Input: Character animation data and environment information.
[1563] Output: Suitable sound effects.
[1564] Specific operation: The sound effect selection module selects sounds such as birds singing, footsteps, and bench sounds from the sound library.
[1565] Step 6:
[1566] The server selects background music based on plot information, determining the background music to create an atmosphere that matches the scene.
[1567] Input: Plot information.
[1568] Output: Suitable background music.
[1569] Specific operation: The music selection module selects relaxing music from the music database.
[1570] Step 7:
[1571] The server uses a timeline editor to combine character models, movements, sound effects, and background music into a high-quality video, aligning all elements appropriately to create the final video.
[1572] Input: Character models, animation data, sound effects, background music.
[1573] Output: The finished video file.
[1574] Specific operation: Each element is integrated in the timeline editor, then encoded and generated into a single video.
[1575] Step 8:
[1576] The server sends the generated video to the user's device, where the user can watch the video by downloading or streaming.
[1577] Input: Your finished video file.
[1578] Output: The video sent to the user's device.
[1579] Specific operation: The server compresses the video file and sends it to the user's device. The user then downloads the video and watches it.
[1580] (Application example 1)
[1581] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1582] Conventional video generation systems require a lot of manual work from the user and specialized knowledge and skills, making it difficult for anyone to easily generate high-quality videos and share them on content distribution services. Furthermore, integrating character settings, plot information, sound effects, and background music during video generation is complicated, making it difficult to maintain professional quality.
[1583] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1584] In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for generating a video by integrating the character model, movements, sound effects, and background music, and means for uploading the generated video to a content distribution service. This allows users to easily generate high-quality videos and share them on the content distribution service.
[1585] "Character settings" are detailed setting information such as the character's personality, appearance, clothing, etc., specified by the user.
[1586] A "character model" is a three-dimensional digital representation generated based on a character's settings.
[1587] "Plot information" is detailed information about the scenes and scenarios required for the video specified by the user.
[1588] "Movement generation" is the process of automatically generating character movements based on plot information.
[1589] "Sound effects" are sound effects that are selected based on the action and environment being generated.
[1590] "Background music" is music that is selected to fit a scene based on plot information.
[1591] "Video generation" is the process of integrating character models, movements, sound effects, and background music to generate a single, high-quality video.
[1592] A "content distribution service" is a platform for sharing and viewing generated videos over the Internet.
[1593] A "user terminal" is an electronic device operated by a user, such as a smartphone or computer.
[1594] "Physics simulation" is a simulation technology that realistically reflects the movements and sound effects of generated character models.
[1595] "Smartphone Application" means a smartphone application that enables users to create, distribute, and guide their own video content.
[1596] The present invention provides a system that enables users to easily generate high-quality videos and share them via content distribution services. Specific embodiments of this system are described below.
[1597] Character setting input
[1598] The device provides users with a means to input detailed settings such as the character's personality, appearance, clothing, etc., allowing them to specifically design their own character.
[1599] Character Model Generation
[1600] The server receives the character setting information sent from the device, analyzes it, and generates a character model, a three-dimensional digital representation with an appearance and personality based on the setting.
[1601] Plot information input
[1602] Users are provided with a means to input details of scenes (plot information) required for the video via their device. For example, they can specifically describe a scenario such as "a character sitting on a bench in a park."
[1603] motion generation
[1604] The server analyzes plot information and automatically generates character movements. This movement generation includes setting animations and movement paths so that characters move naturally. Furthermore, it can use physics simulation to reflect realistic movements.
[1605] Sound effect selection
[1606] The server selects sound effects based on the generated movements and environment, using physics simulation to determine the audio data to realistically reflect the character's movements and environmental sounds.
[1607] Background music selection
[1608] Based on the plot information, the server selects background music to create an appropriate atmosphere for the scene, for example, relaxing music for a park scene.
[1609] Video Integration and Generation
[1610] The server then combines the character models, movements, sound effects, and background music to create a video, syncing all the elements together to create a single, high-quality video that is then uploaded to a content distribution service.
[1611] Sending to user terminal
[1612] The generated video is sent to the user's device, and the user can watch the video on their device by downloading or streaming it.
[1613] Hardware and software used
[1614] The implementation of this system uses high-performance cloud servers (e.g., Amazon EC2 or Google Cloud Platform) and the Python library aiohttp, an asynchronous HTTP client, with the following APIs:
[1615] Character Model Generation API
[1616] Plotting behavior generation API
[1617] Sound Effect Selection API
[1618] Background music selection API
[1619] Video Generation API
[1620] Specific examples
[1621] A user uses a smartphone application to enter the following information:
[1622] Character description: "A character with a cheerful personality, red hair, and blue clothes."
[1623] Plot information: "A character sits on a bench in a park."
[1624] Sound effects: park ambient sounds, character footsteps
[1625] Background music: Relaxing music
[1626] Based on this information, the server integrates the character's 3D model, movements, sound effects, and background music to generate a high-quality video, which is then uploaded to a content distribution service.
[1627] Prompt Sentence Examples
[1628] "Please create a scene where the user creates a character with a cheerful personality, red hair, and blue clothes, sitting on a bench in a park. Use the ambient sounds of the park and the character's footsteps as sound effects, and choose a relaxing song for the background music."
[1629] The system allows users to easily generate high-quality videos and share them on content distribution services.
[1630] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1631] Step 1:
[1632] The user inputs character settings using a terminal. Specifically, the user enters detailed information about the character, such as personality, appearance, and clothing, into an input form and sends it to the server. The input data includes the character's name, hair color, clothing color, personality, etc. The server analyzes this data and prepares it.
[1633] Step 2:
[1634] The server receives the character setting information sent by the user and generates a 3D model of the character based on that information. The server uses an AI model to generate a character model with the specified characteristics and stores it in an internal database. The generated character model has the characteristics specified by the user, such as appearance and personality.
[1635] Step 3:
[1636] The user inputs plot information using the device, specifically details of the scenes required for the video (e.g., "A character sits on a bench in a park"), and the device sends this plot information to the server.
[1637] Step 4:
[1638] The server analyzes the received plot information and automatically generates character movements. The server combines AI models and physics simulations to generate natural movements and movement paths. For example, based on the input plot information, it generates a series of movements for a character to walk and sit on a bench.
[1639] Step 5:
[1640] The server selects appropriate sound effects based on the generated actions and environment. Specifically, it selects sound effects that correspond to the character's actions (e.g., walking, sitting, etc.) and environmental sounds (e.g., birds chirping in the park). It then extracts appropriate audio files from the audio database and synchronizes them with the actions.
[1641] Step 6:
[1642] The server selects background music based on the plot information. Specifically, it selects music with a relaxing atmosphere to match the scene. If necessary, it adjusts the tempo and atmosphere of the music to match the plot information.
[1643] Step 7:
[1644] The server then combines the generated character models, movements, sound effects, and background music to generate a video, adjusting the timing of all elements and stitching them together into a single, high-quality video. The video file is generated internally and can later be downloaded or streamed.
[1645] Step 8:
[1646] The server uploads the generated video to a content distribution service, where users can access the video over the Internet and share it with other users.
[1647] Step 9:
[1648] The server then sends the generated video to the user's device, where the user can receive it and save it on their device or watch it via streaming. Communication between the server and the device is encrypted to ensure secure data transfer.
[1649] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1650] This invention is a system that allows users to easily generate high-quality videos, and provides a means to automatically integrate user-specified character settings, plot information, sound effects, and background music and output them as a single video. This system also incorporates an emotion engine that recognizes the user's emotions and modifies or changes the character's movements, facial expressions, background music, and plot information based on those emotions.
[1651] Program Overview
[1652] Character setting input
[1653] Users use the device to input detailed settings such as the character's personality, appearance, clothing, etc., allowing them to design a character to their liking.
[1654] Character Model Generation
[1655] The server receives the character setting information sent from the device, analyzes it, and generates a character model, which is generated as a three-dimensional digital representation.
[1656] Plot information input
[1657] The user inputs the video scenario and scene details (plot information) via the device. For example, the user can specifically describe the scenario, such as "a character sits on a bench in a park."
[1658] motion generation
[1659] The server analyzes the plot information and generates movement animations that determine how the characters will move in a given scenario or scene, including setting up animations and movement paths for the characters to move naturally.
[1660] Sound effect selection
[1661] After the action animation is determined, the server selects sound effects that are appropriate for the action and the surrounding environment.The server uses physics simulation to determine the sound effects that correspond to the environment and action.
[1662] Background music selection
[1663] Based on the plot information, the server selects background music that matches the scene, selecting music from a database that matches the atmosphere of the plot and creating the atmosphere of the scene.
[1664] Using the Emotion Engine
[1665] The server receives data from the device to recognize the user's emotions (e.g., voice, facial expressions, behavior logs, etc.). The emotion engine analyzes this data and recognizes the user's emotional state. For example, if the user is laughing happily, the emotion engine recognizes a positive emotion.
[1666] Change movements and expressions
[1667] Based on the user's recognized emotional state, the server changes the character's behavior and facial expression. For example, if the user looks sad, the character's facial expression and behavior will also change to look sad.
[1668] Change background music
[1669] Based on the recognized emotion, the server selects appropriate background music. For example, if the user is in a relaxed state, background music will be selected to create a relaxing atmosphere throughout the scene.
[1670] Modifying plot information
[1671] The emotion engine may automatically modify plot information depending on the user's emotions. For example, if the user is feeling stressed, the scenario may be changed to a simpler, calmer one.
[1672] Video Integration and Generation
[1673] The server combines all elements—character models, movements, sound effects, background music, and emotion-based changes—to generate a single video, resulting in a high-quality video optimized for the user's emotions.
[1674] Specific examples
[1675] Character Settings
[1676] When a user selects a character with a bright personality, red hair, and blue clothes on their device screen, that information is sent to the server, which then generates a character model based on that information.
[1677] Input of plot information and generation of motion
[1678] Users input plot information, such as "a character sits on a bench in the park," which is sent to a server that analyzes the information and generates natural-looking movements for the character.
[1679] Using emotion engines and video generation
[1680] When a user watches a video, the device sends emotional data to the server. For example, if the user is smiling, the emotion engine recognizes this positive emotion and changes the character's facial expressions and movements to a brighter tone and the background music to match. The server combines these elements to generate a video, which is then sent to the device. The user can then download and watch the video.
[1681] In this way, the system of the present invention allows the user to automatically generate high-quality videos that are optimized to suit the user's emotions simply by making simple, detailed settings.
[1682] The processing flow will be explained below.
[1683] Step 1:
[1684] The user opens the character setting screen on the device and enters detailed settings for the character, such as personality, appearance, clothing, etc. Once the input is complete, the setting information is sent to the server.
[1685] Step 2:
[1686] The server receives the character setting information sent from the device. The server analyzes this information and generates a character model with the set characteristics. The generated character model is expressed as a digital 3D model.
[1687] Step 3:
[1688] The user inputs the video scenario and scene details (plot information) on the device. For example, they specify specific locations and actions, such as "a character sitting on a bench in a park." Once input is complete, the plot information is sent to the server.
[1689] Step 4:
[1690] The server receives the plot information sent from the device. The server analyzes the plot information and generates movement animations to determine how the character should move in the specified scenario or scene. For example, it calculates and sets the movement of a character walking and sitting on a park bench.
[1691] Step 5:
[1692] After the movement animation is determined, the server selects sound effects appropriate for that movement and the surrounding environment, using physics simulation to determine, for example, the sound of a character walking or the ambient sounds of a park (birds chirping, wind noise, etc.).
[1693] Step 6:
[1694] The server selects background music (BGM) that matches the scene based on the plot information. The server selects music from a database that matches the atmosphere of the plot. For example, it selects calming BGM to match a relaxing scene in a park.
[1695] Step 7:
[1696] The device collects the user's emotional data (voice, facial expression, behavior log, etc.) and sends it to the server. For example, the device's camera and microphone can be used to analyze the user's facial expression and tone of voice.
[1697] Step 8:
[1698] The server analyzes the received emotion data using an emotion engine to identify the user's emotional state. For example, if the user is smiling, it recognizes a positive emotion.
[1699] Step 9:
[1700] Based on the user's recognized emotions, the server changes the character's movements and facial expressions. For example, if the user is smiling, the character will also smile.
[1701] Step 10:
[1702] Based on the user's recognized emotion, the server changes the background music that is set. For example, if the user is relaxed, it selects a calm background music that matches the user's emotion.
[1703] Step 11:
[1704] The server combines all the elements—character models, movements, sound effects, background music, and emotion-based modifications—to generate a single video.
[1705] Step 12:
[1706] The server sends the generated video to the user's device, where the user can download and watch the video.
[1707] Step 13:
[1708] If a user wishes to make corrections to a video, they input the corrections from their device and send them back to the server. The server then regenerates the character model and movement animations based on the corrections, selects new sound effects and background music, and regenerates the corrected video.
[1709] Example 2
[1710] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1711] Conventional video creation systems have the problem that it is difficult for users to customize based on their emotions, and the quality of the generated videos is not optimized for the user's emotions and preferences. In addition, adjusting each element (character settings, plot information, sound effects, background music) individually requires a great deal of effort and expertise, making it difficult to easily generate high-quality videos.
[1712] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1713] In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for receiving data for recognizing the user's emotion from the terminal, means for changing the character's movements and facial expressions based on the recognized emotional state of the user, means for selecting appropriate background music based on the recognized emotion, means for automatically correcting the plot information according to the user's emotional state, and means for generating a video by integrating the character model, movements, sound effects, and background music. This enables users to easily generate high-quality videos optimized to their own emotions.
[1714] "Users" are individuals or organizations that operate the system and input character settings and plot information.
[1715] The "server" is a central processing unit that receives and analyzes data sent by users and generates character models, movements, sound effects, background music, etc.
[1716] "Character settings" refers to detailed information about a character's personality, appearance, clothing, etc.
[1717] A "character model" is a three-dimensional digital representation generated based on a character's settings.
[1718] "Plot information" is information about the scenario and scene details of a video, and is specified by the user.
[1719] "Movement animation" refers to the specific movements and actions that characters perform based on plot information.
[1720] "Sound effects" are sound effects that are appropriate for the character's movements and environment, and are selected based on physical models.
[1721] "Background music" is music used to create the overall atmosphere of a video scene.
[1722] The "emotion engine" is a technology that analyzes user emotions and customizes various elements of a video based on those emotions.
[1723] "Emotion recognition data" is data used to recognize a user's emotions, and includes voice, facial expressions, behavioral logs, etc.
[1724] "Video generation means" is a mechanism for integrating character models, movements, sound effects, and background music to generate a single video.
[1725] This invention is a system that allows users to easily generate high-quality videos. Users input character settings, plot information, sound effects, and background music using a terminal, and the system automatically integrates them and outputs them as a single video. The system also uses an emotion engine that recognizes the user's emotions and modifies or changes the character's movements, facial expressions, background music, and plot information based on those emotions.
[1726] This system operates in cooperation with both the server and the terminal. The main processing steps and related hardware and software are explained below.
[1727] Character setting input
[1728] Using the device interface, users input detailed settings for their character, such as personality, appearance, clothing, etc. This information is sent from the device to a server, which then analyzes the character settings information.
[1729] Character Model Generation
[1730] The server generates a character model based on the received character setting information. Specifically, it uses 3D modeling software such as Blender or Unity. This character model is then used for subsequent movement generation and animation.
[1731] Plot information input
[1732] The user inputs the video scenario and scene details (plot information) via the terminal. For example, the user inputs a scenario such as "A character sits on a bench in a park." This information is also sent to the server and stored as plot information.
[1733] motion generation
[1734] The server analyzes the plot information and determines how the characters will move in the specified scenario or scene. Animation software such as Maya or Autodesk MotionBuilder is used for movement animation. The generated movements include animations and movement paths for the characters to move naturally.
[1735] Sound effect selection
[1736] After the movement animation is determined, the server selects sound effects appropriate for the movement and environment, using sound effect generation tools such as FMOD and Wwise with physical simulation.
[1737] Background music selection
[1738] Based on the plot information, the server selects background music that matches the scene, selected from music libraries such as Epidemic Sound and Artlist.
[1739] Using the Emotion Engine
[1740] To recognize the user's emotions, the device uses a camera and microphone to collect the user's facial and voice data and sends it to the server. The server then uses an emotion engine (such as OpenCV or Affectiva) to analyze the user's emotional state. For example, if the user is smiling, it is recognized as a positive emotion.
[1741] Change movements and expressions
[1742] The server changes the character's behavior and facial expression based on the user's recognized emotional state. For example, if the user is sad, the character's facial expression and behavior will also change to a sad one.
[1743] Change background music
[1744] The server selects appropriate background music based on the recognized emotion, for example, if the user is relaxed, calm background music is selected.
[1745] Modifying plot information
[1746] The emotion engine automatically modifies plot information according to the user's emotions, so if the user is feeling stressed, the scenario will be changed to a simpler, calmer one.
[1747] Video Integration and Generation
[1748] The server then combines all elements based on the character model, movements, sound effects, background music, and emotional changes into a single video. The final high-quality video is then generated using video editing software such as Adobe After Effects or Final Cut Pro. The resulting video is then sent to the user's device, where they can download and watch it.
[1749] Examples of prompt statements
[1750] Here are some examples of prompts to input to a generative AI model:
[1751] "Create a scene in a park where a character with red hair and a cheerful personality is sitting on a bench. Not only should the character be smiling and sitting on the bench, but also integrate relaxing background music and natural environmental sounds. Also, dynamically change the character's facial expressions and behavior based on emotion recognition while the user is viewing the scene."
[1752] In this way, the system of the present invention allows the user to automatically generate high-quality videos that are optimized to suit the user's emotions simply by making simple, detailed settings.
[1753] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1754] Step 1:
[1755] Entering character settings
[1756] The user sets the character's personality, appearance, clothing, etc. on the terminal screen.
[1757] Input: Enter character settings information via device (e.g., "cheerful personality, red hair, blue clothes")
[1758] Processing: The terminal formats the entered information and sends it to the server.
[1759] Specific operation: Enter the setting information through the terminal's GUI and press the "Send" button, and the information will be sent to the server via API.
[1760] Output: Character setting information sent to the server
[1761] Step 2:
[1762] Character model generation
[1763] The server receives the character setting information sent from the terminal and analyzes it.
[1764] Input: Character information (e.g., "cheerful personality, red hair, blue clothes")
[1765] Process: The server uses the Blender API to generate the character model.
[1766] Specific operation: Passes configuration information to the Blender API and executes a script to generate a 3D digital model.
[1767] Output: Generated character model (3D digital model)
[1768] Step 3:
[1769] Entering plot information
[1770] The user inputs the video scenario and scene details (plot information) via the terminal.
[1771] Input: Plot information (e.g., "A character sits on a bench in a park")
[1772] Processing: The terminal formats the entered plot information and sends it to the server.
[1773] Specific operation: Plot information is entered using the GUI, and the information is sent to the server when the "Send" button is pressed.
[1774] Output: Plot information sent to the server
[1775] Step 4:
[1776] Generate movement animation
[1777] The server analyzes the plot information and generates animations of the characters' movements.
[1778] Input: Plot information (e.g., "A character sits on a bench in a park")
[1779] Processing: Generate movement animation using Maya or Autodesk MotionBuilder.
[1780] Specific operation: The server passes the analyzed plot information to the Maya API and executes a script that generates a movement animation.
[1781] Output: Generated movement animation
[1782] Step 5:
[1783] Sound Effect Selection
[1784] The server selects the appropriate sound effect according to the action animation.
[1785] Input: Action animation (e.g. "Sit on a bench")
[1786] Processing: Select sound effects based on physical simulation using FMOD and Wwise.
[1787] Specific behavior: The behavior animation data is passed to the FMOD API, and an algorithm is run to select the appropriate sound effect.
[1788] Output: Selected sound effects
[1789] Step 6:
[1790] Background music selection
[1791] The server selects background music that matches the scene based on the plot information.
[1792] Input: Plot information (e.g., "A character sits on a bench in a park")
[1793] Processing: Select music that matches the scene from the Epidemic Sound and Artlist databases.
[1794] What it does: It parses the plot information and runs an algorithm that searches the music library using the corresponding keywords.
[1795] Output: Selected background music
[1796] Step 7:
[1797] Emotion engine recognizes user emotions
[1798] To recognize the user's emotions, the device uses a camera and microphone to collect facial expressions and voice sounds, which are then sent to a server.
[1799] Input: User's facial expression data and voice data
[1800] Processing: The device formats the collected data and sends it to the server.
[1801] Specific operation: Collects data in real time through the camera and microphone and sends it to the server.
[1802] Output: Emotion recognition data sent to the server
[1803] Step 8:
[1804] Emotion analysis using an emotion engine
[1805] The server uses an emotion engine to recognize the user's emotional state.
[1806] Input: Emotion recognition data (e.g., "image of a smiling face")
[1807] Processing: Sentiment analysis is performed using Affectiva and OpenCV.
[1808] Specific operation: Runs emotion recognition algorithm and obtains emotion analysis result (e.g., positive emotion).
[1809] Output: Emotion analysis results
[1810] Step 9:
[1811] Change movements and expressions
[1812] The server changes the character's movements and facial expressions based on the results of emotion analysis.
[1813] Input: Sentiment analysis result (e.g., "positive sentiment")
[1814] Processing: Updates animation data to change the character's movements and expressions.
[1815] Specific actions: Regenerate character movements and facial expressions using Maya API and Blender API.
[1816] Output: Modified character movements and expressions
[1817] Step 10:
[1818] Change background music
[1819] The server selects appropriate background music based on the emotion analysis results.
[1820] Input: Sentiment analysis result (e.g., "positive sentiment")
[1821] Processing: Select new background music from Epidemic Sound and Artlist.
[1822] What it does: Search your music library using new keywords to find the right background music.
[1823] Output: Modified background music
[1824] Step 11:
[1825] Modifying plot information
[1826] The server automatically modifies plot information according to the user's emotional state.
[1827] Input: Sentiment analysis results (e.g., "I feel stressed")
[1828] Processing: Reanalyze the plot information and change it into a simple and calm scenario.
[1829] Specific operation: Runs an algorithm that recompiles plot information and generates a revised scenario.
[1830] Output: Corrected plot information
[1831] Step 12:
[1832] Video Integration and Generation
[1833] The server combines all elements (character models, movements, sound effects, background music, and emotional changes) to generate the video.
[1834] Input: Character models, movements, sound effects, background music, modified plot information
[1835] Processing: Edit the video using Adobe After Effects and Final Cut Pro.
[1836] What it does: Place elements on a timeline in video editing software, add effects and transitions, and generate the final video.
[1837] Output: High-quality generated video
[1838] Step 13:
[1839] Sending the generated video
[1840] The server transmits the generated high-quality video to the terminal.
[1841] Input: Generated video
[1842] Processing: Converts the video format and encodes it to prepare it for sending to the device.
[1843] Specific operation: The video is converted into the appropriate format and sent to the device using the HTTP / HTTPS protocol.
[1844] Output: Video sent to device
[1845] Step 14:
[1846] Download and watch videos
[1847] Users download and watch the videos generated on their devices.
[1848] Input: Video sent from the server
[1849] Processing: Receive, store, and play the video on the device.
[1850] Specific operation: Play the video using the device's media player.
[1851] Output: Viewable video
[1852] Through this series of processing steps, users can easily generate and watch high-quality, emotion-optimized videos.
[1853] (Application example 2)
[1854] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1855] When generating videos, conventional systems have difficulty automatically adjusting content based on the user's emotions. This has led to a demand for technology that can automatically generate high-quality videos optimized for the user's emotions and situation. Furthermore, systems that can utilize emotional data to provide content that is more attuned to the user are also needed. Furthermore, to ensure that the content of the generated videos is natural, the elements selected (such as character movements, sound effects, and background music) must also be optimized based on the user's emotions.
[1856] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings specified by the user, means for generating a character model based on the character settings, means for receiving plot information specified by the user, means for generating character movements based on the plot information, means for selecting sound effects based on the generated movements and environment, means for selecting background music based on the plot information, means for recognizing user emotion data and adjusting the character model, movements, sound effects, and background music based on the emotion, and means for generating a video by integrating the character model, movements, sound effects, and background music. This makes it possible to automatically generate high-quality videos optimized for the user's emotions and situation.
[1857] "Means for receiving character settings specified by the user" refers to a function for receiving detailed information such as the character's personality, appearance, clothing, etc., entered by the user using the terminal.
[1858] The "means for generating a character model based on character settings" is a function for generating a digital three-dimensional character model based on character setting information specified by the user.
[1859] The "means for receiving plot information specified by the user" is a function for receiving detailed information about the scenario and scenes of the video input by the user.
[1860] The "means for generating character actions based on plot information" is a function for analyzing received plot information and generating character actions that conform to the scenario.
[1861] The "means for selecting sound effects based on the generated actions and environment" is a function for selecting the most appropriate sound effects according to the generated character's actions and the environment.
[1862] The "means for selecting background music based on plot information" is a function for analyzing plot information and selecting background music suitable for a scene.
[1863] "Means for recognizing a user's emotional data and adjusting a character model, movements, sound effects, and background music based on that emotion" refers to a function that analyzes a user's emotional data (voice, facial expressions, etc.) and adjusts a character's movements, facial expressions, sound effects, and background music according to that emotional state.
[1864] "A means of integrating character models, movements, sound effects, and background music to generate video" is a function that combines all of these elements to generate a single, integrated, high-quality video.
[1865] This invention is a system that generates high-quality videos based on user emotions. Specifically, the server integrates character settings, plot information, sound effects, and background music to automatically generate videos that match the user's emotions.
[1866] Program Overview
[1867] 1. Enter character settings
[1868] The user uses the device to input detailed settings for the character, such as personality, appearance, clothing, etc. For example, they might set a character with a cheerful personality, red hair, and blue clothing. This information is then sent to the server.
[1869] 2. Character Model Generation
[1870] The server analyzes the character setting information and generates a three-dimensional digital character model.
[1871] 3. Enter plot information
[1872] The user inputs the video scenario and scene details (plot information) via the device. For example, the user can specifically describe the scenario, such as "A character sits on a bench in a park." This information is then sent to the server.
[1873] 4. Behavior generation
[1874] The server analyzes the plot information and generates animations for characters to move naturally, including the character's movement path and body movements.
[1875] 5. Sound effect selection
[1876] After the movement animation is generated, the server selects sound effects that suit the movement and the surrounding environment, such as "birds chirping" for a park scene.
[1877] 6. Background Music Selection
[1878] Based on the plot information, the server selects background music that matches the scene. Music that creates the atmosphere of the scene is selected from a database.
[1879] 7. Use of Emotion Engine
[1880] The server receives data from the device to recognize the user's emotions (voice, facial expressions, behavioral logs, etc.), and the emotion engine analyzes this data. For example, if the user is laughing happily, the emotion engine will recognize a positive emotion.
[1881] 8. Change movement and facial expression
[1882] Based on the user's recognized emotional state, the server changes the character's behavior and facial expressions. For example, if the user looks sad, the character will also look sad.
[1883] 9. Change the background music
[1884] The emotion engine selects appropriate background music based on the emotion it recognizes. For example, if the user is in a relaxed state, music that creates a relaxing atmosphere will be selected.
[1885] 10. Correcting plot information
[1886] The emotion engine may also automatically modify plot information depending on the user's emotions, for example, changing the scenario to a simpler, calmer one if the user is feeling stressed.
[1887] 11. Video Integration and Generation
[1888] The server combines the character models, movements, sound effects, background music, and emotion-based modifications to generate a single video.
[1889] Hardware and software used
[1890] Hardware:
[1891] Smartphone or tablet
[1892] Camera (to capture facial expressions)
[1893] Microphone (to capture audio)
[1894] software:
[1895] EmotionRecognizer (emotion recognition library)
[1896] VideoGenerator (library for generating videos: OpenCV as an example)
[1897] Specific examples
[1898] Next, we will explain a specific example of how this system can be used. A user uses a smartphone to create a character with a cheerful personality, red hair, and blue clothes, and enters plot information such as "the character is sitting on a bench in a park." If the user is smiling, the emotion engine recognizes this positive emotion, brightens the character's facial expressions and movements, and changes the background music to a more cheerful one. As a result, a high-quality video optimized for the user is generated.
[1899] Example prompt sentence:
[1900] "Change the character's facial expression to a smile, set the background music to upbeat music, and generate a high-quality video."
[1901] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1902] Step 1:
[1903] The user inputs character setting information using a terminal. Specifically, detailed information such as the character's personality, appearance, and clothing is entered, and this information is sent from the terminal to the server. The input data could be specific information such as "a character with a cheerful personality, red hair, and blue clothing." The server then stores this input data in a database.
[1904] Step 2:
[1905] The server generates a digital 3D character model based on the received character configuration information. It uses a generative AI model to analyze this information and generate the character model using 3D modeling software (e.g., Blender). The output is a completed 3D character model.
[1906] Step 3:
[1907] The user inputs the video's scenario and scene details (plot information) via the device. This input data includes specific plot information, such as "a character sits on a bench in a park." The device then sends this information to the server, which then stores the received plot information in a database.
[1908] Step 4:
[1909] The server analyzes the plot information and generates character movement animation. For example, based on the plot information "a character sits on a bench in a park," it generates a natural movement path for the character. Here, 3D animation software (e.g., Maya) is used. The output is movement animation data.
[1910] Step 5:
[1911] The server selects sound effects based on the generated action and environment. For example, for a park scene, sound effects such as "birds chirping" and "wind sounds" are selected. A physics simulation library (e.g., Bullet Physics) is used to select appropriate sound effects. The output is the selected sound effect file.
[1912] Step 6:
[1913] The server selects background music based on the plot information. It analyzes the atmosphere and selects music from a database that matches the scene. For example, relaxing music is selected for a tranquil scene in a park. The output is the selected background music file.
[1914] Step 7:
[1915] The server analyzes the user's emotional data (voice, facial expressions, behavioral logs, etc.) received from the device to recognize the user's emotional state. It uses emotion recognition software (e.g., Facial Emotion Recognition). For example, if the user is smiling happily, the server recognizes a positive emotion. The output is the user's emotional state data.
[1916] Step 8:
[1917] The server adjusts the character's behavior, facial expressions, sound effects, and background music based on the user's recognized emotional state. For example, if the user is sad, the server changes the character's facial expression to a sad one and the background music to a calm one. The output is each element adjusted based on the emotion.
[1918] Step 9:
[1919] The server integrates character models, movements, sound effects, and background music to create a video using video editing software (e.g., Adobe Premiere Pro). The output is a final, high-quality video file.
[1920] Step 10:
[1921] The generated video is sent from the server to the user's device, where the user can watch it. The output is a video file that the user can watch.
[1922] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1923] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1924] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1925] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1926] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1927] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1928] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1929] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1930] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1931] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1932] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1933] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1934] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1935] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1936] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1937] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1938] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1939] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1940] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1941] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1942] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1943] The following is further disclosed regarding the above embodiment.
[1944] (Claim 1)
[1945] A means for receiving user-specified character settings;
[1946] A means for generating a character model based on a character setting;
[1947] a means for receiving user-specified plot information;
[1948] means for generating character actions based on plot information;
[1949] means for selecting sound effects based on the generated motion and environment;
[1950] means for selecting background music based on plot information;
[1951] A system that includes a means for integrating character models, movements, sound effects, and background music to generate video.
[1952] (Claim 2)
[1953] and means for selecting motion and sound effects for the generated character model based on a physical simulation.
[1954] (Claim 3)
[1955] 10. The system of claim 1, further comprising means for transmitting the generated video to a user terminal.
[1956]
[1957] "Example 1"
[1958] (Claim 1)
[1959] A means for receiving user-specified character settings;
[1960] A means for generating a character model based on a character setting;
[1961] a means for receiving user-specified plot information;
[1962] means for generating character actions based on plot information;
[1963] means for selecting sound effects based on the generated motion and environment;
[1964] means for selecting background music based on plot information;
[1965] A means to generate videos by integrating character models, movements, sound effects, and background music;
[1966] A means for selecting environmental sounds and action sounds using physical simulation based on the generated character model and action;
[1967] The system includes means for transmitting the generated video to a user terminal.
[1968] (Claim 2)
[1969] The system of claim 1, wherein the generated character model is synchronized with the movement, sound effects, and background music using a timeline editor.
[1970] (Claim 3)
[1971] The system of claim 1, which automatically generates character models and animations using a generative AI model based on chara...
Claims
1. A means for receiving user-specified character settings; A means for generating a character model based on a character setting; a means for receiving user-specified plot information; means for generating character actions based on plot information; means for selecting sound effects based on the generated motion and environment; means for selecting background music based on plot information; A system that includes a means for integrating character models, movements, sound effects, and background music to generate video.
2. and means for selecting motion and sound effects for the generated character model based on a physical simulation.
3. The system of claim 1 further comprising means for transmitting the generated video to a user terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A