system
The system allows users to create highly customizable visual content using a user-owned terminal and cloud server, addressing the need for specialized knowledge in conventional methods by automating key element extraction and animation style application, supporting diverse languages and interactive features.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Conventional visual content generation methods require specialized knowledge and skills, making them difficult for average users to use, and struggle with generating content that supports diverse languages and cultures.
A system that includes a user-owned terminal and a cloud server, utilizing a generative AI model to analyze user input, automatically extract key elements, apply animation styles, generate characters and backgrounds, and add interactive elements, allowing users to create highly customizable visual content without specialized knowledge.
Enables users to easily generate highly customizable visual content tailored to their needs, supporting multiple languages and interactive features, without requiring specialized knowledge.
Smart Images

Figure 2026070185000001_ABST
Abstract
Description
Technical Field
[0004] , , ,
[0005] , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
[0006] "Content data" refers to information such as videos, text, and audio that can be converted into animation.
[0007] "Key elements" refer to keywords and scenes that are particularly necessary for animation, extracted from the content data.
[0008] "Settings options" provide users with choices for selecting animation styles and themes.
[0009] "Animation style" refers to a specific design or method of expression that determines the look and feel of an animation.
[0010] "Characters and backgrounds" refer to the elements that make up the characters and the environment of a scene in an animation.
[0011] "Interactive elements" refer to features or elements that allow users to interact with or react to an animation.
[0012] "Translation" refers to the process of converting the content of an original text into a different language.
[0013] "Voice narration" refers to the audio element that verbally explains the content of an animation in language.
[0014] "Rendering" refers to the process of converting digital data into final visual and auditory content.
[0015] "User" refers to a person who uses this animation generation system to create content according to their own purposes.
Brief Explanation of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Embodiments for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] The animation generation system of the present invention begins with the uploading of content data from the user's device, followed by processing on a server, and finally providing the user with the final animation. The embodiments thereof are described below in natural language.
[0038] First, users upload the content data they wish to animate using their device. This includes video files, audio data, and text documents. Once the upload is complete, users select the animation's visual style, video type, time period, story theme, language settings, and character details through the device's interface.
[0039] Once the content data and user selections are sent to the server, the server analyzes the received data and automatically extracts key elements from the video and text. The server then determines an appropriate animation style based on the user's selected settings and applies it to the animation. For example, if a Ghibli-style visual style is selected, the style will apply a specific color scheme and line style.
[0040] Based on the user's selection, the server generates character clothing and background scenery appropriate to the chosen time period. If the selected period is the Showa era, the background will depict clothing and architecture characteristic of that era. Furthermore, if the story theme is suspense, a scenario designed to create tension will be generated.
[0041] Furthermore, the server can add interactive elements specifically tailored for educational or business use to the animation. Scenes displaying questions in a quiz format, or pop-ups explaining specific situations, can be inserted. Audio narration and subtitles are then generated in the selected language. If the user selects English, the Japanese content will be translated and presented with English subtitles and narration.
[0042] Finally, the server renders the animation file and sends it to the user's terminal. The user can preview this animation and modify or regenerate it as needed. This allows users to create highly customizable animations tailored to their specific needs, even without specialized knowledge.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The user uploads the content data they want to convert to animation using their device. During this process, the system checks whether the file format (video, text, etc.) is permitted and prepares the data accordingly.
[0046] Step 2:
[0047] Users configure animation styles through the device's interface. This includes image type (e.g., Ghibli style), video type (e.g., illustrated animation), time period (e.g., Showa era), story theme (e.g., suspense), language setting (e.g., English), and character details (e.g., hairstyle, clothing).
[0048] Step 3:
[0049] The device sends the user-uploaded content data and selected configuration options to the server. This data package includes content files and configuration information.
[0050] Step 4:
[0051] The server analyzes the received content data and extracts important elements from videos and text. For videos, a scene segmentation algorithm is used to identify important scenes. For text, keyword extraction is performed.
[0052] Step 5:
[0053] The server applies animation styles based on the user's selected settings. It applies specific filters and color palettes depending on the selected image type, determining the animation's style.
[0054] Step 6:
[0055] The server generates characters and backgrounds according to the configured options. If the setting is the Showa era, the character's clothing and background design will be automatically generated in a Showa style.
[0056] Step 7:
[0057] The server adds interactive elements to animations tailored to educational or business purposes. For example, it can display quizzes within a scene or insert pop-ups that explain specific content.
[0058] Step 8:
[0059] The server generates voice narration and subtitles based on the language settings. It uses speech synthesis technology to create multilingual narration files and translates the subtitle text.
[0060] Step 9:
[0061] The server renders the final animation based on all settings. During this process, all visual elements, as well as audio and subtitles, are synchronized.
[0062] Step 10:
[0063] The server sends the completed animation file to the user's terminal. The user can preview this animation and request readjustments or regeneration from the system as needed.
[0064] (Example 1)
[0065] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0066] In today's world, there is a demand for the efficient generation of highly customizable visual content that meets the diverse needs of users. However, conventional visual content generation methods require specialized knowledge and skills, making them difficult for the average user to use. Furthermore, generating content that supports diverse languages and cultures also presents significant challenges. Solving these problems is highly desirable.
[0067] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0068] In this invention, the server includes means for receiving information and extracting important components based on said information; means for applying expression techniques corresponding to visual and viewing types based on a plurality of setting requests selected by the user; and means for creating characters and environments based on the selected settings. This makes it possible for users to easily generate highly customizable visual content that meets their needs without requiring specialized knowledge.
[0069] "Information" is a collection of data or signals that is provided in a form that is understandable or processable to the recipient.
[0070] A "component" is an important element or part extracted from information or data, and is a fundamental element necessary for a system to perform a specific function.
[0071] A "user" is an individual or organization that operates the system or uses its services.
[0072] A "configuration request" is an instruction or selection that a user inputs into the system, used to control the system's operation and output results.
[0073] "Visual type" refers to the format or style in which visual content is presented, and includes, for example, 2D and 3D animation.
[0074] "Viewing type" refers to the format or method of playing visual content, and typically includes movies, animations, and video clips.
[0075] "Expression technology" refers to technical methods for generating or modifying visual or auditory information, and includes image processing, 3DCG generation, and speech synthesis.
[0076] "Characters" refer to characters or people represented within visual content, and constitute part of the story or scene.
[0077] "Environment" refers to the background and elements of a scene in visual content, indicating the place and situation in which the characters exist.
[0078] This invention relates to a system for generating visual content. The system consists primarily of a user-owned terminal and a cloud server. Specific embodiments thereof are described below.
[0079] Users first upload information about the visual content they wish to generate using their device. This information includes video files, audio data, text documents, etc. The uploaded data is input through the device's user-friendly interface.
[0080] Next, the device provides an interface for the user to enter configuration requests. Here, the user can configure detailed settings regarding the visual type (e.g., 2D or 3D animation), viewing type (e.g., movie, video clip), and characters and settings. This includes selecting a Ghibli-style visual style and time period.
[0081] The data and selections set by the user are sent to the server. The server uses a generative AI model to analyze the received information. This analysis automatically extracts important components. Based on the analysis results, the server determines the appropriate representation technique. This includes applying color tones and line styles specific to the set style.
[0082] Furthermore, the server can add interactive elements for learning and business use to the visual content. For example, it can insert scenes displaying quiz-style questions or pop-ups that supplement information. In addition, it can generate translations and voice narration in various languages according to the specified language settings.
[0083] Finally, the server renders the completed visual content and sends the result to the user's device. The user can preview the content on their device and make adjustments or regenerate it as needed.
[0084] (Specific example)
[0085] If a user wants to create visual content depicting the adventures of their pet dog, they upload a video from their device and select a Ghibli-style and comedy theme. This information is processed on the server, a high-quality animation is generated, and delivered to the user's device.
[0086] (Example of a prompt message)
[0087] "Please create this video in a Ghibli-esque style, as an adventure animation set in modern-day Tokyo. The theme is comedy."
[0088] This system allows users to easily generate highly customizable visual content tailored to their needs, even without specialized knowledge.
[0089] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0090] Step 1:
[0091] The user uploads information about the visual content they wish to generate using their device. Inputs include video files, audio data, and text documents. The device temporarily stores the received information and prepares it for transmission to the server. The output is a set of content data ready for transmission.
[0092] Step 2:
[0093] The terminal presents a menu through its interface for the user to select settings. Here, the user determines the visual type (e.g., 2D or 3D), viewing style, style (e.g., Ghibli style), time period, and details about the characters and environment. The user's selections are received as input, and the terminal compiles them into a single configuration file. The configuration file is generated as output.
[0094] Step 3:
[0095] The terminal sends the prepared content data and configuration files to the server. The server receives this data and performs data analysis as input. Using a generative AI model, it extracts important components from the data and sets the initial visual style and narrative structure. The analysis results are obtained as output.
[0096] Step 4:
[0097] The server determines the visual style and generates characters and environments based on the analysis results and user settings. The generating AI model constructs a specific screen using a particular color scheme and line style. The analysis results and configuration files are used as input, and an initial version of the visual content is generated as output.
[0098] Step 5:
[0099] The server adds interactive elements to the visual content as needed, for educational or business purposes. This includes interactive quiz formats and explanatory pop-ups. User-specific requests are referenced as input, and interactive content is generated as output.
[0100] Step 6:
[0101] The server generates translations and audio narration in various languages based on the selected language settings. Natural language processing technology is used, with the content's language data as input. The output consists of translated subtitle files and audio files.
[0102] Step 7:
[0103] Ultimately, the server renders the completed visual content in high quality and sends it to the user's device. Animation, audio, and subtitles are integrated as input, and the complete version of the visual content is generated as output. The user can preview this on their device and make adjustments or regenerate it as needed.
[0104] (Application Example 1)
[0105] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0106] Traditional animation generation systems struggled to provide interactive experiences that fully reflected the individual needs of users. Especially in educational and business settings, there was a demand for advanced interactions utilizing eye-tracking and voice input. Furthermore, there was a lack of technology to deliver real-time customized content based on diverse user-defined options.
[0107] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0108] In this invention, the server includes means for receiving content data and extracting important elements based on the content data; means for applying animation styles corresponding to image and video types based on multiple setting options selected by the user; and means for enabling interaction with the user using eye tracking and voice input. This allows the user to experience customized interactive animations in real time and enables operation using eye tracking and voice in educational and business applications.
[0109] "Content data" refers to informational materials provided by users, including data in formats such as video, audio, and text, which form the basis for animation generation.
[0110] "Methods for extracting important elements" refers to technologies that analyze content data to automatically select key points in story development and visual expression.
[0111] "Means of applying animation styles" refers to techniques for reflecting specific visual representations and dynamic configurations in generated animations based on user settings.
[0112] "Means for generating characters and backgrounds" refers to technology for designing digital characters and backgrounds according to the era and theme selected by the user, based on setting options.
[0113] "Methods for incorporating interactive elements" refer to techniques that introduce features into animations that allow users to make selections and perform actions within the animation.
[0114] "Means for generating translations and audio narration" refers to technologies that enable the creation of subtitles and audio descriptions for animations in multiple languages.
[0115] "Means that enable interaction using eye-tracking and voice input" refers to technology that detects the user's gaze and voice, responds to them as input, and enables two-way communication with animations.
[0116] "Means of delivering rendered animations" refers to technologies that send completed animations to users, making them viewable on their devices.
[0117] To implement this invention, the user must first upload content data that they wish to animate using an information and communication terminal such as a smartphone. This content data includes video, audio, and text. The user selects several options, such as visual style, language settings, and interactive elements, through the terminal's interface.
[0118] The server receives data and configuration options from this user. The server uses a generative AI model as its software. The generative AI model extracts key elements from the content data and applies a specified animation style. This also includes technologies that allow the user to interactively manipulate the animation using eye tracking and voice input. Specific hardware uses such as smart glasses, and an application that works in conjunction with them is installed on a smartphone.
[0119] The generated animation is rendered, including translation and voice narration. The server delivers the completed animation to the user's device or smart glasses. The user can watch the animation and use their gaze and voice to answer quizzes and make selections.
[0120] One concrete example is an interactive animation created for history lessons in educational settings. This animation depicts the Sengoku period (Warring States period) of Japan, allowing viewers to select character actions using their gaze and answer quizzes via voice input. An example of a prompt for a generative AI model would be: "Generate an interactive animation set in Sengoku-era Japan. The visual style should be Ukiyo-e (Japanese woodblock print), and English narration should be included. Questions should be displayed in a quiz format based on user selections, and answers should be selectable via eye-tracking."
[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0122] Step 1:
[0123] The user uploads content data to be animated using an information and communication terminal. Input data accepted includes video files, audio data, and text documents. Output data is sent from the terminal to the server. Specifically, file selection and upload operations are performed through the user interface.
[0124] Step 2:
[0125] On the device, the user selects setting options such as visual style, video type, and language settings. The data received as input is the user's selection information. The setting options are sent to the server as output. Specifically, the options for each choice are displayed as pull-down menus or radio buttons for selection.
[0126] Step 3:
[0127] The server receives content data and user configuration options. The input data received consists of content data and configuration options sent from the terminal. The output provides the basic data necessary for the next processing step. Specifically, the server stores the received data and prepares to begin analysis.
[0128] Step 4:
[0129] The server uses a generative AI model to extract key elements from content data. The input data is content data. The output is key points extracted from text and visuals. Specifically, the generative AI model analyzes the data and uses a particular algorithm to extract the main themes and storylines.
[0130] Step 5:
[0131] The server applies animation styles based on the user's selected settings to generate characters and backgrounds. The input data received consists of the settings options and key elements. The output is styled animated content. Specifically, the AI incorporates visual effects and designs into the animation according to the selected style.
[0132] Step 6:
[0133] The server configures the interaction using eye tracking and voice input. The data received as input includes hardware information and user settings. The output is the animation in which the interactive elements are reflected. Specifically, the server analyzes data from the eye tracking device and microphone and configures the program to reflect the results in the animation's operation.
[0134] Step 7:
[0135] The server generates translations and voice narration in multiple languages. The input data received includes the content's text information and the user's language settings. The output consists of subtitles and voice narration in the specified language. Specifically, translation software and speech synthesis technology are used to generate the text as narration and subtitles.
[0136] Step 8:
[0137] The server renders the final generated animation and delivers it to the user's device or smart glasses. The data received as input is the animation data to be rendered. The animation is delivered as output in a format that the user can view. Specifically, the file is sent via a data transfer protocol and played back on the device.
[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0139] The animation generation system of the present invention incorporates an emotion engine to recognize the user's emotional state in real time and adjust the animation based on that state. The embodiments thereof are described below in natural language.
[0140] First, users use their devices to upload the content data they wish to animate to the system. This content data includes videos, audio, text, and other formats. Next, through the device's interface, users select the animation's visual style, video type, time period, story theme, language settings, and even detailed character settings.
[0141] This data and settings are sent from the terminal to the server. The server analyzes the content data, extracts key elements, and applies animation styles based on the user's selections. The server also generates characters and backgrounds according to the time period and theme, and adds interactive elements to the animation.
[0142] The emotion engine analyzes the user's emotional state through the camera and biosensors built into the device. Based on the data from the emotion engine, the server adjusts the animation scenes, character expressions, and voice tone. This functionality allows the animation to react to the user's emotions, providing a more personalized experience.
[0143] For example, if the user expresses happiness, the characters in the animation will smile, and their voice tone will become brighter. Conversely, if the user is sad, the characters' expressions will become slightly calmer, and their voice tone will become gentler.
[0144] Finally, the server sends the rendered animation, based on all settings and adjustments, to the user's device. The user can preview this animation and verify that the emotion-responsive customizations have been properly implemented. This allows users to easily create more interactive and personalized animations.
[0145] The following describes the processing flow.
[0146] Step 1:
[0147] The user uploads the content data (video, text, etc.) they want to animate to the system using their device. The user then selects the animation image type, video type, time period, story theme, language setting, and character settings through the interface.
[0148] Step 2:
[0149] The device packages the uploaded content data and the user's selected settings information and sends it to the server.
[0150] Step 3:
[0151] The server analyzes the received content data and extracts important elements. In the case of video, a scene segmentation algorithm is used to identify meaningful scenes.
[0152] Step 4:
[0153] The server applies animation styles based on user selections. This includes filtering selected image types and applying color palettes.
[0154] Step 5:
[0155] The server automatically generates character clothing and backgrounds according to the time period and story theme. For example, if the selected era is the Showa period, the design of the clothing and background will change to suit that style.
[0156] Step 6:
[0157] The emotion engine uses the device's built-in camera or biosensors to recognize the user's current emotions. This data is sent to a server.
[0158] Step 7:
[0159] The server adjusts the animation scenes, character expressions, and voice tone based on the emotion engine data. If the user perceives the content as enjoyable, the character will smile and a cheerful voice tone will be used.
[0160] Step 8:
[0161] The server incorporates interactive elements into the animation, inserting quizzes, pop-ups, and other features tailored to educational and business applications.
[0162] Step 9:
[0163] The server generates voice narration and translated subtitles based on the language selected by the user.
[0164] Step 10:
[0165] The server renders the final animation and sends this completed file to the user's terminal. The user can then review the animation and request modifications or regeneration as needed.
[0166] (Example 2)
[0167] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0168] Conventional motion expression generation systems have difficulty achieving interactive expressions that reflect user emotions in real time, and also have limited support for multiple languages and customization of characters and backgrounds in terms of historical setting. As a result, they have been unable to generate motion expressions that match the individual user's situation and intentions, leading to a decline in the quality of the user experience.
[0169] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0170] In this invention, the server includes means for receiving content information and extracting key elements based on said content information; means for applying motion expression styles corresponding to image and video formats based on multiple setting options selected by the user; means for generating characters and backgrounds based on the selected settings; means for incorporating interactive elements into the motion expression; and means for recognizing the user's emotions using an emotion analysis function and adjusting the scenes, character expressions, and sound tones of the motion expression based on the recognized emotions. This makes it possible to generate interactive and personalized motion expressions that respond to the user's emotions and settings.
[0171] "Content information" refers to media data provided by users, such as text, audio, and video.
[0172] "Key elements" refer to important information and features contained in content information that are extracted in order to understand and process the content.
[0173] "Setting options" refer to the choices that users can make to determine the animation's visual style, video type, time period, theme, language setting, and character details.
[0174] "Action style" refers to the style that defines the appearance and movement of an animation.
[0175] "Characters appearing in the animation" refers to characters such as people and animals that are depicted within the generated animation.
[0176] "Background" refers to the environment or scenery in which the characters exist within an animation.
[0177] "Interactive elements" refer to elements that enable two-way interaction with the user and add interactivity to animations.
[0178] "Emotional analysis function" refers to technology used to recognize and analyze a user's emotional state.
[0179] "Action expression" refers to the entire animation generated based on the user's emotions and settings.
[0180] "Acoustic tone" refers to the pitch and quality of the sounds used in animation.
[0181] This invention is a system that generates personalized animations in real time according to the user's emotions and settings. The system operates primarily via a terminal and a server. The terminal provides an interface for the user to upload content information and specify subsequent setting options. The terminal then analyzes the user's emotions using a built-in camera and biosensors. This hardware is used to acquire the user's real-time emotional data.
[0182] The server extracts key elements based on the received content information and applies behavioral expression styles according to the setting choices. This process utilizes natural language processing (NLP) and computer vision technologies to understand and analyze the content information. The server also uses generative AI models to design characters and backgrounds, and to adjust scenes, facial expressions, and sound tones based on data obtained from emotion analysis functions. Specific software includes, for example, software commonly used in 3D modeling.
[0183] For example, if a user wants to create an animation featuring their pet, a possible prompt might be "Generate a fun, dynamic animation themed around the adventures of a pet dog." Based on this prompt, the AI model generates characters and backgrounds to depict an adventure story starring a pet dog, and the animation's actions are customized based on the user's settings and emotions.
[0184] This system allows users to easily create personalized behavioral expressions and enjoy highly interactive experiences, especially those that respond to emotions.
[0185] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0186] Step 1:
[0187] Users upload content information (e.g., videos, audio, text) to the system using their device. They select the necessary content via the device's file selection function. The input is the media data selected by the user, and the output is that data sent to the server.
[0188] Step 2:
[0189] The server analyzes the received content information and extracts key elements. Natural language processing and computer vision technologies are used for the analysis. The input is media data uploaded by the user, and the output is a list of important elements from the data. Specifically, this includes object recognition within videos and keyword extraction from text.
[0190] Step 3:
[0191] The user inputs animation settings from their device. These include visual style, video format, time period, story theme, language settings, and character details. The input represents the user's desired settings, and this information is sent to the server as output. Settings are configured using dropdown menus and sliders on the device's interface.
[0192] Step 4:
[0193] The server selects a behavioral expression style based on the user's settings. It integrates the received configuration information and applies the appropriate style using a generative AI model. The input is the extracted content and the user's configuration information, and the output is the applied animation style.
[0194] Step 5:
[0195] The server generates characters and backgrounds according to the time period and theme. A generation AI model is used here to create prompt statements. This allows the AI to design the appearance of the characters and their backgrounds. The input consists of user settings and content information elements, and the output is the generated characters and backgrounds. Specifically, a prompt statement might be something like, "Generate a fun, animated video about a pet dog's adventure."
[0196] Step 6:
[0197] The device uses its built-in camera and biosensors to analyze the user's emotions in real time. This data is sent to a server and used for emotion analysis. The input is emotion data obtained from the user, and the output is the analyzed emotion information. The server adjusts the animation according to the emotional state.
[0198] Step 7:
[0199] The server adjusts the animation scenes, character expressions, and sound tones based on the results of the emotion analysis. This enables personalized behavioral expressions that respond to the user's emotions. The input is the analysis results, and the output is the adjusted animation.
[0200] Step 8:
[0201] Ultimately, the server renders the animation based on all the settings and adjustments. The input is the adjusted animation data, and the output is a rendered, high-quality animation file.
[0202] Step 9:
[0203] The server sends the rendered animation to the terminal, and the user previews it on the terminal. At this stage, the user verifies that the animation is represented as intended. The input is the rendered animation file, and the output is what is displayed to the user.
[0204] (Application Example 2)
[0205] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0206] There is a challenge in providing personalized visual representations that respond to user emotions in real time. In particular, when content is fixed, it is difficult to respond immediately to changes in the viewer's emotions and provide an individualized experience. Furthermore, smoothly implementing changes in voice tone and facial expressions that match the viewer's emotions while supporting multiple languages is also a technical challenge.
[0207] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0208] In this invention, the server includes means for receiving content data and extracting important elements based on the content data; means for applying visual styles corresponding to image and video types based on a plurality of setting options selected by the user; and means for detecting the user's emotional state through emotion analysis and adjusting the facial expressions and voice tone of the characters accordingly. This makes it possible to generate and deliver personalized visual representations that match the user's emotional state in real time.
[0209] "Content data" refers to information used as material for generating visual representations, and includes formats such as video, audio, and text.
[0210] A "means for extracting important elements" refers to a system that has the function of identifying and analyzing information necessary for visual expression from content data.
[0211] "Settings options" include multiple specification and style choices that users can select when customizing the visual presentation.
[0212] "Visual style" refers to the overall characteristics that form the distinctive features of design and expression in visual representation.
[0213] "Means for generating characters and backgrounds" refers to a system that has a process for creating characters and backgrounds within a visual representation based on set specifications.
[0214] "Interactive elements" refer to components that allow users to interact with visual representations.
[0215] "Means for generating translations and musical narration" refers to a system that supports multiple languages and has a process for creating audio descriptions necessary for visual representations, either mechanically or manually.
[0216] "Emotional analysis" is a technology that analyzes biometric data and facial expression data to understand the user's psychological state.
[0217] A "means of delivering visual representations" refers to a system that has a process for delivering the final created visual representation to the user.
[0218] This invention is a system that generates personalized visual representations in response to the user's emotional state. The server receives content data and extracts important elements from it. Specific software used for data analysis includes "OpenCV" and "EmotionAPI." The user selects setting options to determine the style of the visual representation via a terminal. Based on this, the server uses tools such as "Unity" to apply the visual style and generate characters and backgrounds.
[0219] The emotion analysis process is achieved by using the camera and biosensors installed in the device to analyze data such as the user's facial expressions and heart rate in real time. Based on the data obtained from this analysis, the facial expressions and voice tones of the characters in the visual representation are adjusted. The server then delivers the final rendered visual representation to the user's device.
[0220] For example, if a user smiles while watching an emotionally moving scene in a visual presentation, the server can detect this and adjust the characters' expressions to reflect positive emotions, and change the music to a more upbeat tone. This system also has translation capabilities for different languages, generates musical narration, and provides a smooth viewing experience.
[0221] An example of a prompt to input into the generation AI model would be, "Generate an animation for when the user is laughing. Make the character brighter and the voice tone more energetic." In this way, it is possible to provide interactive visual representations that are tailored to the individual user's emotions in real time.
[0222] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0223] Step 1:
[0224] Users upload content data they wish to animate using their device. This input data includes video, audio, and text data. The device sends this data to the server, which receives it. Based on the received data, the server extracts key elements using OpenCV or the Emotion API. This process involves data analysis and sampling.
[0225] Step 2:
[0226] The user selects setting options to determine the style of visual representation through the terminal interface. Inputs include visual style, video type, time period, story theme, language settings, and character details. This data is sent from the terminal to the server, which uses Unity to apply the visual style based on these options and generate characters and backgrounds. Specifically, the program calls the core graphics module and begins rendering according to the settings.
[0227] Step 3:
[0228] The system uses the camera and biosensors built into the device to analyze the user's emotional state in real time. Inputs include the user's facial expression data and heart rate. The server converts this into emotional data based on the EmotionAPI and obtains the analysis results. Based on this data, the system adjusts the facial expressions and voice tones of the characters in the visual representation. Specifically, the character's animation module is dynamically adjusted based on the input emotional data.
[0229] Step 4:
[0230] The server integrates all settings and analysis results and finally renders the visual representation. This includes generated characters, backgrounds, interactive elements, and translated musical narration. This output data is delivered to the terminal as a visual representation. The visual representation is updated and provided to the user in real time. The user confirms that the customization has been done correctly by viewing it.
[0231] Step 5:
[0232] In this application example, the visual representation changes according to the user's emotions. For example, if the emotion analysis detects that the user is smiling, the characters will have cheerful expressions, and the music will change to a more energetic tone. This allows the user to experience personalized visual representations. The prompts used for the generative AI model would be input statements such as, "Generate an animation for when the user is smiling. Make the characters brighter and the voice tone more energetic."
[0233] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0234] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0235] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0236] [Second Embodiment]
[0237] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0238] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0239] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0240] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0241] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0242] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0243] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0244] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0245] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0246] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0247] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0248] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0249] The animation generation system of the present invention begins with the uploading of content data from a user's device, followed by processing on a server, and finally providing the user with the animation. The embodiments thereof are described below in natural language.
[0250] First, users upload the content data they wish to animate using their device. This includes video files, audio data, and text documents. Once the upload is complete, users select the animation's visual style, video type, time period, story theme, language settings, and character details through the device's interface.
[0251] Once the content data and user selections are sent to the server, the server analyzes the received data and automatically extracts key elements from the video and text. The server then determines an appropriate animation style based on the user's selected settings and applies it to the animation. For example, if a Ghibli-style visual style is selected, the style will apply a specific color scheme and line style.
[0252] Based on the user's selection, the server generates character clothing and background scenery appropriate to the chosen time period. If the selected period is the Showa era, the background will depict clothing and architecture characteristic of that era. Furthermore, if the story theme is suspense, a scenario designed to create tension will be generated.
[0253] Furthermore, the server can add interactive elements specifically tailored for educational or business use to the animation. Scenes displaying questions in a quiz format, or pop-ups explaining specific situations, can be inserted. Audio narration and subtitles are then generated in the selected language. If the user selects English, the Japanese content will be translated and presented with English subtitles and narration.
[0254] Finally, the server renders the animation file and sends it to the user's terminal. The user can preview this animation and modify or regenerate it as needed. This allows users to create highly customizable animations tailored to their specific needs, even without specialized knowledge.
[0255] The following describes the processing flow.
[0256] Step 1:
[0257] The user uploads the content data they want to convert to animation using their device. During this process, the system checks whether the file format (video, text, etc.) is permitted and prepares the data accordingly.
[0258] Step 2:
[0259] Users configure animation styles through the device's interface. This includes image type (e.g., Ghibli style), video type (e.g., illustrated animation), time period (e.g., Showa era), story theme (e.g., suspense), language setting (e.g., English), and character details (e.g., hairstyle, clothing).
[0260] Step 3:
[0261] The device sends the user-uploaded content data and selected configuration options to the server. This data package includes content files and configuration information.
[0262] Step 4:
[0263] The server analyzes the received content data and extracts important elements from videos and text. For videos, a scene segmentation algorithm is used to identify important scenes. For text, keyword extraction is performed.
[0264] Step 5:
[0265] The server applies animation styles based on the user's selected settings. It applies specific filters and color palettes depending on the selected image type, determining the animation's style.
[0266] Step 6:
[0267] The server generates characters and backgrounds according to the configured options. If the setting is the Showa era, the character's clothing and background design will be automatically generated in a Showa style.
[0268] Step 7:
[0269] The server adds interactive elements to animations tailored to educational or business purposes. For example, it can display quizzes within a scene or insert pop-ups that explain specific content.
[0270] Step 8:
[0271] The server generates voice narration and subtitles based on the language settings. It uses speech synthesis technology to create multilingual narration files and translates the subtitle text.
[0272] Step 9:
[0273] The server renders the final animation based on all settings. During this process, all visual elements, as well as audio and subtitles, are synchronized.
[0274] Step 10:
[0275] The server sends the completed animation file to the user's terminal. The user can preview this animation and request readjustments or regeneration from the system as needed.
[0276] (Example 1)
[0277] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0278] In modern times, there is a demand for efficiently generating highly customizable visual content that meets the needs of various users. However, ordinary visual content generation methods require specialized knowledge and skills and are difficult for general users to handle easily. Also, generating content corresponding to various languages and cultures has become a high barrier. It is desired to solve these problems.
[0279] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0280] In this invention, the server includes means for receiving information and extracting important components based on the information, means for applying expression techniques corresponding to visual types and viewing types based on a plurality of setting requests selected by the user, and means for creating characters and environments based on the selected settings. As a result, users can easily generate highly customizable visual content according to their needs without requiring specialized knowledge.
[0281] "Information" is a collection of data or signals and is provided in a form that can be understood or processed by the recipient.
[0282] "Component" is an important element or part extracted from information or data and is a basic element required for the system to execute a specific function.
[0283] "User" is an individual or organization that operates the system or uses its services.
[0284] "Setting request" is an instruction or option input by the user to the system and is used to control the operation and output result of the system.
[0285] "Visual type" refers to the form or style in which visual content is expressed and includes, for example, 2D or 3D animations.
[0286] The "viewing type" refers to the playback format or method of visual content, and usually includes movies, animations, video clips, etc.
[0287] The "expression technology" is a technical method for generating or modifying visual or acoustic information, and includes image processing, 3D CG generation, voice synthesis, etc.
[0288] The "characters" are the characters or people represented in visual content and constitute part of the story or scene.
[0289] The "environment" is the background or scene component of visual content and indicates the place or situation where the characters exist.
[0290] This invention is related to a visual content generation system. The system is mainly composed of a terminal owned by the user and a cloud server. The specific embodiments will be described below.
[0291] First, the user uses the terminal to upload information about the visual content to be generated. This information includes video files, audio data, text documents, etc. to be used. The uploaded data is input through the user-friendly interface of the terminal.
[0292] Next, the terminal provides an interface for the user to input setting requirements. Here, the user makes detailed settings regarding the viewing type (e.g., 2D or 3D animation), the viewing type (e.g., movie, video clip), and the characters and environment. This includes the selection of a visual style like that of Studio Ghibli and the background of the era.
[0293] The data and selections set by the user are sent to the server. The server uses a generative AI model to analyze the received information. This analysis automatically extracts important components. Based on the analysis results, the server determines the appropriate representation technique. This includes applying color tones and line styles specific to the set style.
[0294] Furthermore, the server can add interactive elements for learning and business use to the visual content. For example, it can insert scenes displaying quiz-style questions or pop-ups that supplement information. In addition, it can generate translations and voice narration in various languages according to the specified language settings.
[0295] Finally, the server renders the completed visual content and sends the result to the user's device. The user can preview the content on their device and make adjustments or regenerate it as needed.
[0296] (Specific example)
[0297] If a user wants to create visual content depicting the adventures of their pet dog, they upload a video from their device and select a Ghibli-style and comedy theme. This information is processed on the server, a high-quality animation is generated, and delivered to the user's device.
[0298] (Example of a prompt message)
[0299] "Please create this video in a Ghibli-esque style, as an adventure animation set in modern-day Tokyo. The theme is comedy."
[0300] This system allows users to easily generate highly customizable visual content tailored to their needs, even without specialized knowledge.
[0301] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0302] Step 1:
[0303] The user uploads information about the visual content that they hope to generate using the terminal. The input includes video files, audio data, text documents, etc. The terminal temporarily stores the received information and prepares to send it to the server. As output, a series of content data that is ready for data transmission is obtained.
[0304] Step 2:
[0305] The terminal presents a menu for the user to select settings through the interface. Here, the user determines the visual type (e.g., 2D or 3D), viewing type, style (e.g., Studio Ghibli style), background era, details of the characters and environment. The user's selection is received as input, and the terminal combines this into one configuration file. As output, a configuration file is generated.
[0306] Step 3:
[0307] The terminal sends the prepared content data and the configuration file to the server. The server receives this and performs data analysis as input. It extracts important components from the data using a generation AI model and makes initial settings for the visual style and story structure. As output, an analysis result is obtained.
[0308] Step 4:
[0309] Based on the analysis result and the user's settings, the server determines the visual style and generates the characters and environment. The generation AI model constructs a specific screen using a specific color tone and line style. The analysis result and the configuration file are used as input, and an initial version of the visual content is generated as output.
[0310] Step 5:
[0311] The server adds interactive elements to the visual content as needed, for educational or business purposes. This includes interactive quiz formats and explanatory pop-ups. User-specific requests are referenced as input, and interactive content is generated as output.
[0312] Step 6:
[0313] The server generates translations and audio narration in various languages based on the selected language settings. Natural language processing technology is used, with the content's language data as input. The output consists of translated subtitle files and audio files.
[0314] Step 7:
[0315] Ultimately, the server renders the completed visual content in high quality and sends it to the user's device. Animation, audio, and subtitles are integrated as input, and the complete version of the visual content is generated as output. The user can preview this on their device and make adjustments or regenerate it as needed.
[0316] (Application Example 1)
[0317] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0318] Traditional animation generation systems struggled to provide interactive experiences that fully reflected the individual needs of users. Especially in educational and business settings, there was a demand for advanced interactions utilizing eye-tracking and voice input. Furthermore, there was a lack of technology to deliver real-time customized content based on diverse user-defined options.
[0319] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0320] In this invention, the server includes means for receiving content data and extracting important elements based on the content data; means for applying animation styles corresponding to image and video types based on multiple setting options selected by the user; and means for enabling interaction with the user using eye tracking and voice input. This allows the user to experience customized interactive animations in real time and enables operation using eye tracking and voice in educational and business applications.
[0321] "Content data" refers to informational materials provided by users, including data in formats such as video, audio, and text, which form the basis for animation generation.
[0322] "Methods for extracting important elements" refers to technologies that analyze content data to automatically select key points in story development and visual expression.
[0323] "Means of applying animation styles" refers to techniques for reflecting specific visual representations and dynamic configurations in generated animations based on user settings.
[0324] "Means for generating characters and backgrounds" refers to technology for designing digital characters and backgrounds according to the era and theme selected by the user, based on setting options.
[0325] "Methods for incorporating interactive elements" refer to techniques that introduce features into animations that allow users to make selections and perform actions within the animation.
[0326] "Means for generating translations and audio narration" refers to technologies that enable the creation of subtitles and audio descriptions for animations in multiple languages.
[0327] "Means that enable interaction using eye-tracking and voice input" refers to technology that detects the user's gaze and voice, responds to them as input, and enables two-way communication with animations.
[0328] "Means of delivering rendered animations" refers to technologies that send completed animations to users, making them viewable on their devices.
[0329] To implement this invention, the user must first upload content data that they wish to animate using an information and communication terminal such as a smartphone. This content data includes video, audio, and text. The user selects several options, such as visual style, language settings, and interactive elements, through the terminal's interface.
[0330] The server receives data and configuration options from this user. The server uses a generative AI model as its software. The generative AI model extracts key elements from the content data and applies a specified animation style. This also includes technologies that allow the user to interactively manipulate the animation using eye tracking and voice input. Specific hardware uses such as smart glasses, and an application that works with them is installed on a smartphone.
[0331] The generated animation is rendered, including translation and voice narration. The server delivers the completed animation to the user's device or smart glasses. The user can watch the animation and use their gaze and voice to answer quizzes and make selections.
[0332] One concrete example is an interactive animation created for history lessons in educational settings. This animation depicts the Sengoku period (Warring States period) of Japan, allowing viewers to select character actions using their gaze and answer quizzes via voice input. An example of a prompt for a generative AI model would be: "Generate an interactive animation set in Sengoku-era Japan. The visual style should be Ukiyo-e (Japanese woodblock print), and English narration should be included. Questions should be displayed in a quiz format based on user selections, and answers should be selectable via eye-tracking."
[0333] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0334] Step 1:
[0335] The user uploads content data to be animated using an information and communication terminal. Input data accepted includes video files, audio data, and text documents. Output data is sent from the terminal to the server. Specifically, file selection and upload operations are performed through the user interface.
[0336] Step 2:
[0337] On the device, the user selects setting options such as visual style, video type, and language settings. The data received as input is the user's selection information. The setting options are sent to the server as output. Specifically, the options for each choice are displayed as pull-down menus or radio buttons for selection.
[0338] Step 3:
[0339] The server receives content data and user configuration options. The input data received consists of content data and configuration options sent from the terminal. The output provides the basic data necessary for the next processing step. Specifically, the server stores the received data and prepares to begin analysis.
[0340] Step 4:
[0341] The server uses a generative AI model to extract key elements from content data. The input data is content data. The output is key points extracted from text and visuals. Specifically, the generative AI model analyzes the data and uses a particular algorithm to extract the main themes and storylines.
[0342] Step 5:
[0343] The server applies animation styles based on the user's selected settings to generate characters and backgrounds. The input data received consists of the settings options and key elements. The output is styled animated content. Specifically, the AI incorporates visual effects and designs into the animation according to the selected style.
[0344] Step 6:
[0345] The server configures the interaction using eye tracking and voice input. The data received as input includes hardware information and user settings. The output is the animation in which the interactive elements are reflected. Specifically, the server analyzes data from the eye tracking device and microphone and configures the program to reflect the results in the animation's operation.
[0346] Step 7:
[0347] The server generates translations and voice narration in multiple languages. The input data received includes the content's text information and the user's language settings. The output consists of subtitles and voice narration in the specified language. Specifically, translation software and speech synthesis technology are used to generate the text as narration and subtitles.
[0348] Step 8:
[0349] The server renders the final generated animation and delivers it to the user's device or smart glasses. The data received as input is the animation data to be rendered. The animation is delivered as output in a format that the user can view. Specifically, the file is sent via a data transfer protocol and played back on the device.
[0350] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0351] The animation generation system of the present invention incorporates an emotion engine to recognize the user's emotional state in real time and adjust the animation based on that state. An embodiment thereof is described below in natural language.
[0352] First, users use their devices to upload the content data they wish to animate to the system. This content data includes videos, audio, text, and other formats. Next, through the device's interface, users select the animation's visual style, video type, time period, story theme, language settings, and even detailed character settings.
[0353] This data and settings are sent from the terminal to the server. The server analyzes the content data, extracts key elements, and applies animation styles based on the user's selections. The server also generates characters and backgrounds according to the time period and theme, and adds interactive elements to the animation.
[0354] The emotion engine analyzes the user's emotional state through the camera and biosensors built into the device. Based on the data from the emotion engine, the server adjusts the animation scenes, character expressions, and voice tone. This functionality allows the animation to react to the user's emotions, providing a more personalized experience.
[0355] For example, if the user expresses happiness, the characters in the animation will smile, and their voice tone will become brighter. Conversely, if the user is sad, the characters' expressions will become slightly calmer, and their voice tone will become gentler.
[0356] Finally, the server sends the rendered animation, based on all settings and adjustments, to the user's device. The user can preview this animation and verify that the emotion-responsive customizations have been properly implemented. This allows users to easily create more interactive and personalized animations.
[0357] The following describes the processing flow.
[0358] Step 1:
[0359] The user uploads the content data (video, text, etc.) they want to animate to the system using their device. The user then selects the animation image type, video type, time period, story theme, language setting, and character settings through the interface.
[0360] Step 2:
[0361] The device packages the uploaded content data and the user's selected settings information and sends it to the server.
[0362] Step 3:
[0363] The server analyzes the received content data and extracts important elements. In the case of video, a scene segmentation algorithm is used to identify meaningful scenes.
[0364] Step 4:
[0365] The server applies animation styles based on user selections. This includes filtering selected image types and applying color palettes.
[0366] Step 5:
[0367] The server automatically generates character clothing and backgrounds according to the time period and story theme. For example, if the selected era is the Showa period, the design of the clothing and background will change to suit that style.
[0368] Step 6:
[0369] The emotion engine uses the device's built-in camera or biosensors to recognize the user's current emotions. This data is sent to a server.
[0370] Step 7:
[0371] The server adjusts the animation scenes, character expressions, and voice tone based on the emotion engine data. If the user perceives the content as enjoyable, the character will smile and a cheerful voice tone will be used.
[0372] Step 8:
[0373] The server incorporates interactive elements into the animation, inserting quizzes, pop-ups, and other features tailored to educational and business applications.
[0374] Step 9:
[0375] The server generates voice narration and translated subtitles based on the language selected by the user.
[0376] Step 10:
[0377] The server renders the final animation and sends this completed file to the user's terminal. The user can then review the animation and request modifications or regeneration as needed.
[0378] (Example 2)
[0379] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0380] Conventional motion expression generation systems have difficulty achieving interactive expressions that reflect user emotions in real time, and also have limited support for multiple languages and customization of characters and backgrounds in terms of historical setting. As a result, they have been unable to generate motion expressions that match the individual user's situation and intentions, leading to a decline in the quality of the user experience.
[0381] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0382] In this invention, the server includes means for receiving content information and extracting key elements based on said content information; means for applying motion expression styles corresponding to image and video formats based on multiple setting options selected by the user; means for generating characters and backgrounds based on the selected settings; means for incorporating interactive elements into the motion expression; and means for recognizing the user's emotions using an emotion analysis function and adjusting the scenes, character expressions, and sound tones of the motion expression based on the recognized emotions. This makes it possible to generate interactive and personalized motion expressions that respond to the user's emotions and settings.
[0383] "Content information" refers to media data provided by users, such as text, audio, and video.
[0384] "Key elements" refer to important information and features contained in content information that are extracted in order to understand and process the content.
[0385] "Setting options" refer to the choices that users can make to determine the animation's visual style, video type, time period, theme, language setting, and character details.
[0386] "Action style" refers to the style that defines the appearance and movement of an animation.
[0387] "Characters appearing in the animation" refers to characters such as people and animals that are depicted within the generated animation.
[0388] "Background" refers to the environment or scenery in which the characters exist within an animation.
[0389] "Interactive elements" refer to elements that enable two-way interaction with the user and add interactivity to animations.
[0390] "Emotional analysis function" refers to technology used to recognize and analyze a user's emotional state.
[0391] "Action expression" refers to the entire animation generated based on the user's emotions and settings.
[0392] "Acoustic tone" refers to the pitch and quality of the sounds used in animation.
[0393] This invention is a system that generates personalized animations in real time according to the user's emotions and settings. The system operates primarily via a terminal and a server. The terminal provides an interface for the user to upload content information and specify subsequent setting options. The terminal then analyzes the user's emotions using a built-in camera and biosensors. This hardware is used to acquire the user's real-time emotional data.
[0394] The server extracts key elements based on the received content information and applies behavioral expression styles according to the setting choices. This process utilizes natural language processing (NLP) and computer vision technologies to understand and analyze the content information. The server also uses generative AI models to design characters and backgrounds, and to adjust scenes, facial expressions, and sound tones based on data obtained from emotion analysis functions. Specific software includes, for example, software commonly used in 3D modeling.
[0395] For example, if a user wants to create an animation featuring their pet, a possible prompt might be "Generate a fun, dynamic animation themed around the adventures of a pet dog." Based on this prompt, the AI model generates characters and backgrounds to depict an adventure story starring a pet dog, and the animation's actions are customized based on the user's settings and emotions.
[0396] This system allows users to easily create personalized behavioral expressions and enjoy highly interactive experiences, especially those that respond to emotions.
[0397] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0398] Step 1:
[0399] Users upload content information (e.g., videos, audio, text) to the system using their device. They select the necessary content via the device's file selection function. The input is the media data selected by the user, and the output is that data sent to the server.
[0400] Step 2:
[0401] The server analyzes the received content information and extracts key elements. Natural language processing and computer vision technologies are used for the analysis. The input is media data uploaded by the user, and the output is a list of important elements from the data. Specifically, this includes object recognition within videos and keyword extraction from text.
[0402] Step 3:
[0403] The user inputs animation settings from their device. These include visual style, video format, time period, story theme, language settings, and character details. The input represents the user's desired settings, and this information is sent to the server as output. Settings are configured using dropdown menus and sliders on the device's interface.
[0404] Step 4:
[0405] The server selects a behavioral expression style based on the user's settings. It integrates the received configuration information and applies the appropriate style using a generative AI model. The input is the extracted content and the user's configuration information, and the output is the applied animation style.
[0406] Step 5:
[0407] The server generates characters and backgrounds according to the time period and theme. A generation AI model is used here to create prompt statements. This allows the AI to design the appearance of the characters and their backgrounds. The input consists of user settings and content information elements, and the output is the generated characters and backgrounds. Specifically, a prompt statement might be something like, "Generate a fun, animated video about a pet dog's adventure."
[0408] Step 6:
[0409] The device uses its built-in camera and biosensors to analyze the user's emotions in real time. This data is sent to a server and used for emotion analysis. The input is emotion data obtained from the user, and the output is the analyzed emotion information. The server adjusts the animation according to the emotional state.
[0410] Step 7:
[0411] The server adjusts the animation scenes, character expressions, and sound tones based on the results of the emotion analysis. This enables personalized action expressions that respond to the user's emotions. The input is the analysis results, and the output is the adjusted animation.
[0412] Step 8:
[0413] Ultimately, the server renders the animation based on all the settings and adjustments. The input is the adjusted animation data, and the output is a rendered, high-quality animation file.
[0414] Step 9:
[0415] The server sends the rendered animation to the terminal, and the user previews it on the terminal. At this stage, the user verifies that the animation is represented as intended. The input is the rendered animation file, and the output is what is displayed to the user.
[0416] (Application Example 2)
[0417] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0418] There is a challenge in providing personalized visual representations that respond to user emotions in real time. In particular, when content is fixed, it is difficult to respond immediately to changes in the viewer's emotions and provide an individualized experience. Furthermore, smoothly implementing changes in voice tone and facial expressions that match the viewer's emotions while supporting multiple languages is also a technical challenge.
[0419] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0420] In this invention, the server includes means for receiving content data and extracting important elements based on the content data; means for applying visual styles corresponding to image and video types based on a plurality of setting options selected by the user; and means for detecting the user's emotional state through emotion analysis and adjusting the facial expressions and voice tone of the characters accordingly. This makes it possible to generate and deliver personalized visual representations that match the user's emotional state in real time.
[0421] "Content data" refers to information used as material for generating visual representations, and includes formats such as video, audio, and text.
[0422] A "means for extracting important elements" refers to a system that has the function of identifying and analyzing information necessary for visual expression from content data.
[0423] "Settings options" include multiple specification and style choices that users can select when customizing the visual presentation.
[0424] "Visual style" refers to the overall characteristics that form the distinctive features of design and expression in visual representation.
[0425] "Means for generating characters and backgrounds" refers to a system that has a process for creating characters and backgrounds within a visual representation based on set specifications.
[0426] "Interactive elements" refer to components that allow users to interact with visual representations.
[0427] "Means for generating translations and musical narration" refers to a system that supports multiple languages and has a process for creating audio descriptions necessary for visual representations, either mechanically or manually.
[0428] "Emotional analysis" is a technology that analyzes biometric data and facial expression data to understand the user's psychological state.
[0429] A "means of delivering visual representations" refers to a system that has a process for delivering the final created visual representation to the user.
[0430] This invention is a system that generates personalized visual representations in response to the user's emotional state. The server receives content data and extracts important elements from it. Specific software used for data analysis includes "OpenCV" and "EmotionAPI." The user selects setting options to determine the style of the visual representation via a terminal. Based on this, the server uses tools such as "Unity" to apply the visual style and generate characters and backgrounds.
[0431] The emotion analysis process utilizes the camera and biosensors installed in the device to analyze data such as the user's facial expressions and heart rate in real time. Based on the data obtained from this analysis, the facial expressions and voice tones of the characters in the visual representation are adjusted. The server then delivers the final rendered visual representation to the user's device.
[0432] For example, if a user smiles while watching an emotionally moving scene in a visual presentation, the server can detect this and adjust the characters' expressions to reflect positive emotions, and change the music to a more upbeat tone. This system also has translation capabilities for different languages, generates musical narration, and provides a smooth viewing experience.
[0433] An example of a prompt to input into the generation AI model would be, "Generate an animation for when the user is laughing. Make the character brighter and the voice tone more energetic." In this way, it is possible to provide interactive visual representations that are tailored to the individual user's emotions in real time.
[0434] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0435] Step 1:
[0436] Users upload content data they wish to animate using their device. This input data includes video, audio, and text data. The device sends this data to the server, which receives it. Based on the received data, the server extracts key elements using OpenCV or the Emotion API. This process involves data analysis and sampling.
[0437] Step 2:
[0438] The user selects setting options to determine the style of visual representation through the terminal interface. Inputs include visual style, video type, time period, story theme, language settings, and character details. This data is sent from the terminal to the server, which uses Unity to apply the visual style based on these options and generate characters and backgrounds. Specifically, the program calls the core graphics module and begins rendering according to the settings.
[0439] Step 3:
[0440] The system uses the camera and biosensors built into the device to analyze the user's emotional state in real time. Inputs include the user's facial expression data and heart rate. The server converts this into emotional data based on the EmotionAPI and obtains the analysis results. Based on this data, the system adjusts the facial expressions and voice tones of the characters in the visual representation. Specifically, the character's animation module is dynamically adjusted based on the input emotional data.
[0441] Step 4:
[0442] The server integrates all settings and analysis results and finally renders the visual representation. This includes generated characters, backgrounds, interactive elements, and translated musical narration. This output data is delivered to the terminal as a visual representation. The visual representation is updated and provided to the user in real time. The user confirms that the customization has been done correctly by viewing it.
[0443] Step 5:
[0444] In this application example, the visual representation changes according to the user's emotions. For example, if the emotion analysis detects that the user is smiling, the characters will have cheerful expressions, and the music will change to a more energetic tone. This allows the user to experience personalized visual representations. The prompts used for the generative AI model would be input statements such as, "Generate an animation for when the user is smiling. Make the characters brighter and the voice tone more energetic."
[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0448] [Third Embodiment]
[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0461] The animation generation system of the present invention begins with the uploading of content data from a user's device, followed by processing on a server, and finally providing the user with the animation. The embodiments thereof are described below in natural language.
[0462] First, users upload the content data they wish to animate using their device. This includes video files, audio data, and text documents. Once the upload is complete, users select the animation's visual style, video type, time period, story theme, language settings, and character details through the device's interface.
[0463] Once the content data and user selections are sent to the server, the server analyzes the received data and automatically extracts key elements from the video and text. The server then determines an appropriate animation style based on the user's selected settings and applies it to the animation. For example, if a Ghibli-style visual style is selected, the style will apply a specific color scheme and line style.
[0464] Based on the user's selection, the server generates character clothing and background scenery appropriate to the chosen time period. If the selected period is the Showa era, the background will depict clothing and architecture characteristic of that era. Furthermore, if the story theme is suspense, a scenario designed to create tension will be generated.
[0465] Furthermore, the server can add interactive elements specifically tailored for educational or business use to the animation. Scenes displaying questions in a quiz format, or pop-ups explaining specific situations, can be inserted. Audio narration and subtitles are then generated in the selected language. If the user selects English, the Japanese content will be translated and presented with English subtitles and narration.
[0466] Finally, the server renders the animation file and sends it to the user's terminal. The user can preview this animation and modify or regenerate it as needed. This allows users to create highly customizable animations tailored to their specific needs, even without specialized knowledge.
[0467] The following describes the processing flow.
[0468] Step 1:
[0469] The user uploads the content data they want to convert to animation using their device. During this process, the system checks whether the file format (video, text, etc.) is permitted and prepares the data accordingly.
[0470] Step 2:
[0471] Users configure animation styles through the device's interface. This includes image type (e.g., Ghibli style), video type (e.g., illustrated animation), time period (e.g., Showa era), story theme (e.g., suspense), language setting (e.g., English), and character details (e.g., hairstyle, clothing).
[0472] Step 3:
[0473] The device sends the user-uploaded content data and selected configuration options to the server. This data package includes content files and configuration information.
[0474] Step 4:
[0475] The server analyzes the received content data and extracts important elements from videos and text. For videos, a scene segmentation algorithm is used to identify important scenes. For text, keyword extraction is performed.
[0476] Step 5:
[0477] The server applies animation styles based on the user's selected settings. It applies specific filters and color palettes depending on the selected image type, determining the animation's style.
[0478] Step 6:
[0479] The server generates characters and backgrounds according to the configured options. If the setting is the Showa era, the character's clothing and background design will be automatically generated in a Showa style.
[0480] Step 7:
[0481] The server adds interactive elements to animations tailored to educational or business purposes. For example, it can display quizzes within a scene or insert pop-ups that explain specific content.
[0482] Step 8:
[0483] The server generates voice narration and subtitles based on the language settings. It uses speech synthesis technology to create multilingual narration files and translates the subtitle text.
[0484] Step 9:
[0485] The server renders the final animation based on all settings. During this process, all visual elements, as well as audio and subtitles, are synchronized.
[0486] Step 10:
[0487] The server sends the completed animation file to the user's terminal. The user can preview this animation and request readjustments or regeneration from the system as needed.
[0488] (Example 1)
[0489] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0490] In today's world, there is a demand for the efficient generation of highly customizable visual content that meets the diverse needs of users. However, conventional visual content generation methods require specialized knowledge and skills, making them difficult for the average user to use. Furthermore, generating content that supports diverse languages and cultures also presents significant challenges. Solving these problems is highly desirable.
[0491] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0492] In this invention, the server includes means for receiving information and extracting important components based on said information; means for applying expression techniques corresponding to visual and viewing types based on a plurality of setting requests selected by the user; and means for creating characters and environments based on the selected settings. This makes it possible for users to easily generate highly customizable visual content that meets their needs without requiring specialized knowledge.
[0493] "Information" is a collection of data or signals that is provided in a form that is understandable or processable to the recipient.
[0494] A "component" is an important element or part extracted from information or data, and is a fundamental element necessary for a system to perform a specific function.
[0495] A "user" is an individual or organization that operates the system or uses its services.
[0496] A "configuration request" is an instruction or selection that a user inputs into the system, used to control the system's operation and output results.
[0497] "Visual type" refers to the format or style in which visual content is presented, and includes, for example, 2D and 3D animation.
[0498] "Viewing type" refers to the format or method of playing visual content, and typically includes movies, animations, and video clips.
[0499] "Expression technology" refers to technical methods for generating or modifying visual or auditory information, and includes image processing, 3DCG generation, and speech synthesis.
[0500] "Characters" refer to characters or people represented within visual content, and constitute part of the story or scene.
[0501] "Environment" refers to the background and elements of a scene in visual content, indicating the place and situation in which the characters exist.
[0502] This invention relates to a system for generating visual content. The system consists primarily of a user-owned terminal and a cloud server. Specific embodiments thereof are described below.
[0503] Users first upload information about the visual content they wish to generate using their device. This information includes video files, audio data, text documents, etc. The uploaded data is input through the device's user-friendly interface.
[0504] Next, the device provides an interface for the user to enter configuration requests. Here, the user can configure detailed settings regarding the visual type (e.g., 2D or 3D animation), viewing type (e.g., movie, video clip), and characters and settings. This includes selecting a Ghibli-style visual style and time period.
[0505] The data and selections set by the user are sent to the server. The server uses a generative AI model to analyze the received information. This analysis automatically extracts important components. Based on the analysis results, the server determines the appropriate representation technique. This includes applying color tones and line styles specific to the set style.
[0506] Furthermore, the server can add interactive elements for learning and business use to the visual content. For example, it can insert scenes displaying quiz-style questions or pop-ups that supplement information. In addition, it can generate translations and voice narration in various languages according to the specified language settings.
[0507] Finally, the server renders the completed visual content and sends the result to the user's device. The user can preview the content on their device and make adjustments or regenerate it as needed.
[0508] (Specific example)
[0509] If a user wants to create visual content depicting the adventures of their pet dog, they upload a video from their device and select a Ghibli-style and comedy theme. This information is processed on the server, a high-quality animation is generated, and delivered to the user's device.
[0510] (Example of a prompt message)
[0511] "Please create this video in a Ghibli-esque style, as an adventure animation set in modern-day Tokyo. The theme is comedy."
[0512] This system allows users to easily generate highly customizable visual content tailored to their needs, even without specialized knowledge.
[0513] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0514] Step 1:
[0515] The user uploads information about the visual content they wish to generate using their device. Inputs include video files, audio data, and text documents. The device temporarily stores the received information and prepares it for transmission to the server. The output is a set of content data ready for transmission.
[0516] Step 2:
[0517] The terminal presents a menu through its interface for the user to select settings. Here, the user determines the visual type (e.g., 2D or 3D), viewing style, style (e.g., Ghibli style), time period, and details about the characters and environment. The user's selections are received as input, and the terminal compiles them into a single configuration file. The configuration file is generated as output.
[0518] Step 3:
[0519] The terminal sends the prepared content data and configuration files to the server. The server receives these and performs data analysis as input. Using a generative AI model, it extracts important components from the data and performs initial setup of the visual style and narrative structure. The analysis results are obtained as output.
[0520] Step 4:
[0521] The server determines the visual style and generates characters and environments based on the analysis results and user settings. The generating AI model constructs a specific screen using a particular color scheme and line style. The analysis results and configuration files are used as input, and an initial version of the visual content is generated as output.
[0522] Step 5:
[0523] The server adds interactive elements to the visual content as needed, for educational or business purposes. This includes interactive quiz formats and explanatory pop-ups. User-specific requests are referenced as input, and interactive content is generated as output.
[0524] Step 6:
[0525] The server generates translations and audio narration in various languages based on the selected language settings. Natural language processing technology is used, with the content's language data as input. The output consists of translated subtitle files and audio files.
[0526] Step 7:
[0527] Ultimately, the server renders the completed visual content in high quality and sends it to the user's device. Animation, audio, and subtitles are integrated as input, and the complete version of the visual content is generated as output. The user can preview this on their device and make adjustments or regenerate it as needed.
[0528] (Application Example 1)
[0529] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0530] Traditional animation generation systems struggled to provide interactive experiences that fully reflected the individual needs of users. Especially in educational and business settings, there was a demand for advanced interactions utilizing eye-tracking and voice input. Furthermore, there was a lack of technology to deliver real-time customized content based on diverse user-defined options.
[0531] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0532] In this invention, the server includes means for receiving content data and extracting important elements based on the content data; means for applying animation styles corresponding to image and video types based on multiple setting options selected by the user; and means for enabling interaction with the user using eye tracking and voice input. This allows the user to experience customized interactive animations in real time and enables operation using eye tracking and voice in educational and business applications.
[0533] "Content data" refers to informational materials provided by users, including data in formats such as video, audio, and text, which form the basis for animation generation.
[0534] "Methods for extracting important elements" refers to technologies that analyze content data to automatically select key points in story development and visual expression.
[0535] "Means of applying animation styles" refers to techniques for reflecting specific visual representations and dynamic configurations in generated animations based on user settings.
[0536] "Means for generating characters and backgrounds" refers to technology for designing digital characters and backgrounds according to the era and theme selected by the user, based on setting options.
[0537] "Methods for incorporating interactive elements" refer to techniques that introduce features into animations that allow users to make selections and perform actions within the animation.
[0538] "Means for generating translations and audio narration" refers to technologies that enable the creation of subtitles and audio descriptions for animations in multiple languages.
[0539] "Means that enable interaction using eye-tracking and voice input" refers to technology that detects the user's gaze and voice, responds to them as input, and enables two-way communication with animations.
[0540] "Means of delivering rendered animations" refers to technologies that send completed animations to users, making them viewable on their devices.
[0541] To implement this invention, the user must first upload content data that they wish to animate using an information and communication terminal such as a smartphone. This content data includes video, audio, and text. The user selects several options, such as visual style, language settings, and interactive elements, through the terminal's interface.
[0542] The server receives data and configuration options from this user. The server uses a generative AI model as its software. The generative AI model extracts key elements from the content data and applies a specified animation style. This also includes technologies that allow the user to interactively manipulate the animation using eye tracking and voice input. Specific hardware uses such as smart glasses, and an application that works with them is installed on a smartphone.
[0543] The generated animation is rendered, including translation and voice narration. The server delivers the completed animation to the user's device or smart glasses. The user can watch the animation and use their gaze and voice to answer quizzes and make selections.
[0544] One concrete example is an interactive animation created for history lessons in educational settings. This animation depicts the Sengoku period (Warring States period) of Japan, allowing viewers to select character actions using their gaze and answer quizzes via voice input. An example of a prompt for a generative AI model would be: "Generate an interactive animation set in Sengoku-era Japan. The visual style should be Ukiyo-e (Japanese woodblock print), and English narration should be included. Questions should be displayed in a quiz format based on user selections, and answers should be selectable via eye-tracking."
[0545] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0546] Step 1:
[0547] The user uploads content data to be animated using an information and communication terminal. Input data accepted includes video files, audio data, and text documents. Output data is sent from the terminal to the server. Specifically, file selection and upload operations are performed through the user interface.
[0548] Step 2:
[0549] On the device, the user selects setting options such as visual style, video type, and language settings. The data received as input is the user's selection information. The setting options are sent to the server as output. Specifically, the options for each choice are displayed as pull-down menus or radio buttons for selection.
[0550] Step 3:
[0551] The server receives content data and user configuration options. The input data received consists of content data and configuration options sent from the terminal. The output provides the basic data necessary for the next processing step. Specifically, the server stores the received data and prepares to begin analysis.
[0552] Step 4:
[0553] The server uses a generative AI model to extract key elements from content data. The input data is content data. The output is key points extracted from text and visuals. Specifically, the generative AI model analyzes the data and uses a particular algorithm to extract the main themes and storylines.
[0554] Step 5:
[0555] The server applies animation styles based on the user's selected settings to generate characters and backgrounds. The input data received consists of the settings options and key elements. The output is styled animated content. Specifically, the AI incorporates visual effects and designs into the animation according to the selected style.
[0556] Step 6:
[0557] The server configures the interaction using eye tracking and voice input. The data received as input includes hardware information and user settings. The output is the animation in which the interactive elements are reflected. Specifically, the server analyzes data from the eye tracking device and microphone and configures the program to reflect the results in the animation's operation.
[0558] Step 7:
[0559] The server generates translations and voice narration in multiple languages. The input data received includes the content's text information and the user's language settings. The output consists of subtitles and voice narration in the specified language. Specifically, translation software and speech synthesis technology are used to generate the text as narration and subtitles.
[0560] Step 8:
[0561] The server renders the final generated animation and delivers it to the user's device or smart glasses. The data received as input is the animation data to be rendered. The animation is delivered as output in a format that the user can view. Specifically, the file is sent via a data transfer protocol and played back on the device.
[0562] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0563] The animation generation system of the present invention incorporates an emotion engine to recognize the user's emotional state in real time and adjust the animation based on that state. An embodiment thereof is described below in natural language.
[0564] First, users use their devices to upload the content data they wish to animate to the system. This content data includes videos, audio, text, and other formats. Next, through the device's interface, users select the animation's visual style, video type, time period, story theme, language settings, and even detailed character settings.
[0565] This data and settings are sent from the terminal to the server. The server analyzes the content data, extracts key elements, and applies animation styles based on the user's selections. The server also generates characters and backgrounds according to the time period and theme, and adds interactive elements to the animation.
[0566] The emotion engine analyzes the user's emotional state through the camera and biosensors built into the device. Based on the data from the emotion engine, the server adjusts the animation scenes, character expressions, and voice tone. This functionality allows the animation to react to the user's emotions, providing a more personalized experience.
[0567] For example, if the user expresses happiness, the characters in the animation will smile, and their voice tone will become brighter. Conversely, if the user is sad, the characters' expressions will become slightly calmer, and their voice tone will become gentler.
[0568] Finally, the server sends the rendered animation, based on all settings and adjustments, to the user's device. The user can preview this animation and verify that the emotion-responsive customizations have been properly implemented. This allows users to easily create more interactive and personalized animations.
[0569] The following describes the processing flow.
[0570] Step 1:
[0571] The user uploads the content data (video, text, etc.) they want to animate to the system using their device. The user then selects the animation image type, video type, time period, story theme, language setting, and character settings through the interface.
[0572] Step 2:
[0573] The device packages the uploaded content data and the user's selected settings information and sends it to the server.
[0574] Step 3:
[0575] The server analyzes the received content data and extracts important elements. In the case of video, a scene segmentation algorithm is used to identify meaningful scenes.
[0576] Step 4:
[0577] The server applies animation styles based on user selections. This includes filtering selected image types and applying color palettes.
[0578] Step 5:
[0579] The server automatically generates character clothing and backgrounds according to the time period and story theme. For example, if the selected era is the Showa period, the design of the clothing and background will change to suit that style.
[0580] Step 6:
[0581] The emotion engine uses the device's built-in camera or biosensors to recognize the user's current emotions. This data is sent to a server.
[0582] Step 7:
[0583] The server adjusts the animation scenes, character expressions, and voice tone based on the emotion engine data. If the user perceives the content as enjoyable, the character will smile and a cheerful voice tone will be used.
[0584] Step 8:
[0585] The server incorporates interactive elements into the animation, inserting quizzes, pop-ups, and other features tailored to educational and business applications.
[0586] Step 9:
[0587] The server generates voice narration and translated subtitles based on the language selected by the user.
[0588] Step 10:
[0589] The server renders the final animation and sends this completed file to the user's terminal. The user can then review the animation and request modifications or regeneration as needed.
[0590] (Example 2)
[0591] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0592] Conventional motion expression generation systems have difficulty achieving interactive expressions that reflect user emotions in real time, and also have limited support for multiple languages and customization of characters and backgrounds in terms of historical setting. As a result, they have been unable to generate motion expressions that match the individual user's situation and intentions, leading to a decline in the quality of the user experience.
[0593] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0594] In this invention, the server includes means for receiving content information and extracting key elements based on said content information; means for applying motion expression styles corresponding to image and video formats based on multiple setting options selected by the user; means for generating characters and backgrounds based on the selected settings; means for incorporating interactive elements into the motion expression; and means for recognizing the user's emotions using an emotion analysis function and adjusting the scenes, character expressions, and sound tones of the motion expression based on the recognized emotions. This makes it possible to generate interactive and personalized motion expressions that respond to the user's emotions and settings.
[0595] "Content information" refers to media data provided by users, such as text, audio, and video.
[0596] "Key elements" refer to important information and features contained in content information that are extracted in order to understand and process the content.
[0597] "Setting options" refer to the choices that users can make to determine the animation's visual style, video type, time period, theme, language setting, and character details.
[0598] "Action style" refers to the style that defines the appearance and movement of an animation.
[0599] "Characters appearing in the animation" refers to characters such as people and animals that are depicted within the generated animation.
[0600] "Background" refers to the environment or scenery in which the characters exist within an animation.
[0601] "Interactive elements" refer to elements that enable two-way interaction with the user and add interactivity to animations.
[0602] "Emotional analysis function" refers to technology used to recognize and analyze a user's emotional state.
[0603] "Action expression" refers to the entire animation generated based on the user's emotions and settings.
[0604] "Acoustic tone" refers to the pitch and quality of the sounds used in animation.
[0605] This invention is a system that generates personalized animations in real time according to the user's emotions and settings. The system operates primarily via a terminal and a server. The terminal provides an interface for the user to upload content information and specify subsequent setting options. The terminal then analyzes the user's emotions using a built-in camera and biosensors. This hardware is used to acquire the user's real-time emotional data.
[0606] The server extracts key elements based on the received content information and applies behavioral expression styles according to the setting choices. This process utilizes natural language processing (NLP) and computer vision technologies to understand and analyze the content information. The server also uses generative AI models to design characters and backgrounds, and to adjust scenes, facial expressions, and sound tones based on data obtained from emotion analysis functions. Specific software includes, for example, software commonly used in 3D modeling.
[0607] For example, if a user wants to create an animation featuring their pet, a possible prompt might be "Generate a fun, dynamic animation themed around the adventures of a pet dog." Based on this prompt, the AI model generates characters and backgrounds to depict an adventure story starring a pet dog, and the animation's actions are customized based on the user's settings and emotions.
[0608] This system allows users to easily create personalized behavioral expressions and enjoy highly interactive experiences, especially those that respond to emotions.
[0609] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0610] Step 1:
[0611] Users upload content information (e.g., videos, audio, text) to the system using their device. They select the necessary content via the device's file selection function. The input is the media data selected by the user, and the output is that data sent to the server.
[0612] Step 2:
[0613] The server analyzes the received content information and extracts key elements. Natural language processing and computer vision technologies are used for the analysis. The input is media data uploaded by the user, and the output is a list of important elements from the data. Specifically, this includes object recognition within videos and keyword extraction from text.
[0614] Step 3:
[0615] The user inputs animation settings from their device. These include visual style, video format, time period, story theme, language settings, and character details. The input represents the user's desired settings, and this information is sent to the server as output. Settings are configured using dropdown menus and sliders on the device's interface.
[0616] Step 4:
[0617] The server selects a behavioral expression style based on the user's settings. It integrates the received configuration information and applies the appropriate style using a generative AI model. The input is the extracted content and the user's configuration information, and the output is the applied animation style.
[0618] Step 5:
[0619] The server generates characters and backgrounds according to the time period and theme. A generation AI model is used here to create prompt statements. This allows the AI to design the appearance of the characters and their backgrounds. The input consists of user settings and content information elements, and the output is the generated characters and backgrounds. Specifically, a prompt statement might be something like, "Generate a fun, animated video about a pet dog's adventure."
[0620] Step 6:
[0621] The device uses its built-in camera and biosensors to analyze the user's emotions in real time. This data is sent to a server and used for emotion analysis. The input is emotion data obtained from the user, and the output is the analyzed emotion information. The server adjusts the animation according to the emotional state.
[0622] Step 7:
[0623] The server adjusts the animation scenes, character expressions, and sound tones based on the results of the emotion analysis. This enables personalized action expressions that respond to the user's emotions. The input is the analysis results, and the output is the adjusted animation.
[0624] Step 8:
[0625] Ultimately, the server renders the animation based on all the settings and adjustments. The input is the adjusted animation data, and the output is a rendered, high-quality animation file.
[0626] Step 9:
[0627] The server sends the rendered animation to the terminal, and the user previews it on the terminal. At this stage, the user verifies that the animation is represented as intended. The input is the rendered animation file, and the output is what is displayed to the user.
[0628] (Application Example 2)
[0629] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0630] There is a challenge in providing personalized visual representations that respond to user emotions in real time. In particular, when content is fixed, it is difficult to respond immediately to changes in the viewer's emotions and provide an individualized experience. Furthermore, smoothly implementing changes in voice tone and facial expressions that match the viewer's emotions while supporting multiple languages is also a technical challenge.
[0631] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0632] In this invention, the server includes means for receiving content data and extracting important elements based on the content data; means for applying visual styles corresponding to image and video types based on a plurality of setting options selected by the user; and means for detecting the user's emotional state through emotion analysis and adjusting the facial expressions and voice tone of the characters accordingly. This makes it possible to generate and deliver personalized visual representations that match the user's emotional state in real time.
[0633] "Content data" refers to information used as material for generating visual representations, and includes formats such as video, audio, and text.
[0634] A "means for extracting important elements" refers to a system that has the function of identifying and analyzing information necessary for visual expression from content data.
[0635] "Settings options" include multiple specification and style choices that users can select when customizing the visual presentation.
[0636] "Visual style" refers to the overall characteristics that form the distinctive features of design and expression in visual representation.
[0637] "Means for generating characters and backgrounds" refers to a system that has a process for creating characters and backgrounds within a visual representation based on set specifications.
[0638] "Interactive elements" refer to components that allow users to interact with visual representations.
[0639] "Means for generating translations and musical narration" refers to a system that supports multiple languages and has a process for creating audio descriptions necessary for visual representations, either mechanically or manually.
[0640] "Emotional analysis" is a technology that analyzes biometric data and facial expression data to understand the user's psychological state.
[0641] A "means of delivering visual representations" refers to a system that has a process for delivering the final created visual representation to the user.
[0642] This invention is a system that generates personalized visual representations in response to the user's emotional state. The server receives content data and extracts important elements from it. Specific software used for data analysis includes "OpenCV" and "EmotionAPI." The user selects setting options to determine the style of the visual representation via a terminal. Based on this, the server uses tools such as "Unity" to apply the visual style and generate characters and backgrounds.
[0643] The emotion analysis process utilizes the camera and biosensors installed in the device to analyze data such as the user's facial expressions and heart rate in real time. Based on the data obtained from this analysis, the facial expressions and voice tones of the characters in the visual representation are adjusted. The server then delivers the final rendered visual representation to the user's device.
[0644] For example, if a user smiles while watching an emotionally moving scene in a visual presentation, the server can detect this and adjust the characters' expressions to reflect positive emotions, and change the music to a more upbeat tone. This system also has translation capabilities for different languages, generates musical narration, and provides a smooth viewing experience.
[0645] An example of a prompt to input into the generation AI model would be, "Generate an animation for when the user is laughing. Make the character brighter and the voice tone more energetic." In this way, it is possible to provide interactive visual representations that are tailored to the individual user's emotions in real time.
[0646] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0647] Step 1:
[0648] Users upload content data they wish to animate using their device. This input data includes video, audio, and text data. The device sends this data to the server, which receives it. Based on the received data, the server extracts key elements using OpenCV or the Emotion API. This process involves data analysis and sampling.
[0649] Step 2:
[0650] The user selects setting options to determine the style of visual representation through the terminal interface. Inputs include visual style, video type, time period, story theme, language settings, and character details. This data is sent from the terminal to the server, which uses Unity to apply the visual style based on these options and generate characters and backgrounds. Specifically, the program calls the core graphics module and begins rendering according to the settings.
[0651] Step 3:
[0652] The system uses the camera and biosensors built into the device to analyze the user's emotional state in real time. Inputs include the user's facial expression data and heart rate. The server converts this into emotional data based on the EmotionAPI and obtains the analysis results. Based on this data, the system adjusts the facial expressions and voice tones of the characters in the visual representation. Specifically, the character's animation module is dynamically adjusted based on the input emotional data.
[0653] Step 4:
[0654] The server integrates all settings and analysis results and finally renders the visual representation. This includes generated characters, backgrounds, interactive elements, and translated musical narration. This output data is delivered to the terminal as a visual representation. The visual representation is updated and provided to the user in real time. The user confirms that the customization has been done correctly by viewing it.
[0655] Step 5:
[0656] In this application example, the visual representation changes according to the user's emotions. For example, if the emotion analysis detects that the user is smiling, the characters will have cheerful expressions, and the music will change to a more energetic tone. This allows the user to experience personalized visual representations. The prompts used for the generative AI model would be input statements such as, "Generate an animation for when the user is smiling. Make the characters brighter and the voice tone more energetic."
[0657] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0658] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0659] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0660] [Fourth Embodiment]
[0661] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0662] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0663] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0664] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0665] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0666] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0667] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0668] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0669] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0670] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0671] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0672] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0673] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0674] The animation generation system of the present invention begins with the uploading of content data from a user's device, followed by processing on a server, and finally providing the user with the animation. The embodiments thereof are described below in natural language.
[0675] First, users upload the content data they wish to animate using their device. This includes video files, audio data, and text documents. Once the upload is complete, users select the animation's visual style, video type, time period, story theme, language settings, and character details through the device's interface.
[0676] Once the content data and user selections are sent to the server, the server analyzes the received data and automatically extracts key elements from the video and text. The server then determines an appropriate animation style based on the user's selected settings and applies it to the animation. For example, if a Ghibli-style visual style is selected, the style will apply a specific color scheme and line style.
[0677] Based on the user's selection, the server generates character clothing and background scenery appropriate to the chosen time period. If the selected period is the Showa era, the background will depict clothing and architecture characteristic of that era. Furthermore, if the story theme is suspense, a scenario designed to create tension will be generated.
[0678] Furthermore, the server can add interactive elements specifically tailored for educational or business use to the animation. Scenes displaying questions in a quiz format, or pop-ups explaining specific situations, can be inserted. Audio narration and subtitles are then generated in the selected language. If the user selects English, the Japanese content will be translated and presented with English subtitles and narration.
[0679] Finally, the server renders the animation file and sends it to the user's terminal. The user can preview this animation and modify or regenerate it as needed. This allows users to create highly customizable animations tailored to their specific needs, even without specialized knowledge.
[0680] The following describes the processing flow.
[0681] Step 1:
[0682] The user uploads the content data they want to convert to animation using their device. During this process, the system checks whether the file format (video, text, etc.) is permitted and prepares the data accordingly.
[0683] Step 2:
[0684] Users configure animation styles through the device's interface. This includes image type (e.g., Ghibli style), video type (e.g., illustrated animation), time period (e.g., Showa era), story theme (e.g., suspense), language setting (e.g., English), and character details (e.g., hairstyle, clothing).
[0685] Step 3:
[0686] The device sends the user-uploaded content data and selected configuration options to the server. This data package includes content files and configuration information.
[0687] Step 4:
[0688] The server analyzes the received content data and extracts important elements from videos and text. For videos, a scene segmentation algorithm is used to identify important scenes. For text, keyword extraction is performed.
[0689] Step 5:
[0690] The server applies animation styles based on the user's selected settings. It applies specific filters and color palettes depending on the selected image type, determining the animation's style.
[0691] Step 6:
[0692] The server generates characters and backgrounds according to the configured options. If the setting is the Showa era, the character's clothing and background design will be automatically generated in a Showa style.
[0693] Step 7:
[0694] The server adds interactive elements to animations tailored to educational or business purposes. For example, it can display quizzes within a scene or insert pop-ups that explain specific content.
[0695] Step 8:
[0696] The server generates voice narration and subtitles based on the language settings. It uses speech synthesis technology to create multilingual narration files and translates the subtitle text.
[0697] Step 9:
[0698] The server renders the final animation based on all settings. During this process, all visual elements, as well as audio and subtitles, are synchronized.
[0699] Step 10:
[0700] The server sends the completed animation file to the user's terminal. The user can preview this animation and request readjustments or regeneration from the system as needed.
[0701] (Example 1)
[0702] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0703] In today's world, there is a demand for the efficient generation of highly customizable visual content that meets the diverse needs of users. However, conventional visual content generation methods require specialized knowledge and skills, making them difficult for the average user to use. Furthermore, generating content that supports diverse languages and cultures also presents significant challenges. Solving these problems is highly desirable.
[0704] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0705] In this invention, the server includes means for receiving information and extracting important components based on said information; means for applying expression techniques corresponding to visual and viewing types based on a plurality of setting requests selected by the user; and means for creating characters and environments based on the selected settings. This makes it possible for users to easily generate highly customizable visual content that meets their needs without requiring specialized knowledge.
[0706] "Information" is a collection of data or signals that is provided in a form that is understandable or processable to the recipient.
[0707] A "component" is an important element or part extracted from information or data, and is a fundamental element necessary for a system to perform a specific function.
[0708] A "user" is an individual or organization that operates the system or uses its services.
[0709] A "configuration request" is an instruction or selection that a user inputs into the system, used to control the system's operation and output results.
[0710] "Visual type" refers to the format or style in which visual content is presented, and includes, for example, 2D and 3D animation.
[0711] "Viewing type" refers to the format or method of playing visual content, and typically includes movies, animations, and video clips.
[0712] "Expression technology" refers to technical methods for generating or modifying visual or auditory information, and includes image processing, 3DCG generation, and speech synthesis.
[0713] "Characters" refer to characters or people represented within visual content, and constitute part of the story or scene.
[0714] "Environment" refers to the background and elements of a scene in visual content, indicating the place and situation in which the characters exist.
[0715] This invention relates to a system for generating visual content. The system consists primarily of a user-owned terminal and a cloud server. Specific embodiments thereof are described below.
[0716] Users first upload information about the visual content they wish to generate using their device. This information includes video files, audio data, text documents, etc. The uploaded data is input through the device's user-friendly interface.
[0717] Next, the device provides an interface for the user to enter configuration requests. Here, the user can configure detailed settings regarding the visual type (e.g., 2D or 3D animation), viewing type (e.g., movie, video clip), and characters and settings. This includes selecting a Ghibli-style visual style and time period.
[0718] The data and selections set by the user are sent to the server. The server uses a generative AI model to analyze the received information. This analysis automatically extracts important components. Based on the analysis results, the server determines the appropriate representation technique. This includes applying color tones and line styles specific to the set style.
[0719] Furthermore, the server can add interactive elements for learning and business use to the visual content. For example, it can insert scenes displaying quiz-style questions or pop-ups that supplement information. In addition, it can generate translations and voice narration in various languages according to the specified language settings.
[0720] Finally, the server renders the completed visual content and sends the result to the user's device. The user can preview the content on their device and make adjustments or regenerate it as needed.
[0721] (Specific example)
[0722] If a user wants to create visual content depicting the adventures of their pet dog, they upload a video from their device and select a Ghibli-style and comedy theme. This information is processed on the server, a high-quality animation is generated, and delivered to the user's device.
[0723] (Example of a prompt message)
[0724] "Please create this video in a Ghibli-esque style, as an adventure animation set in modern-day Tokyo. The theme is comedy."
[0725] This system allows users to easily generate highly customizable visual content tailored to their needs, even without specialized knowledge.
[0726] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0727] Step 1:
[0728] The user uploads information about the visual content they wish to generate using their device. Inputs include video files, audio data, and text documents. The device temporarily stores the received information and prepares it for transmission to the server. The output is a set of content data ready for transmission.
[0729] Step 2:
[0730] The terminal presents a menu through its interface for the user to select settings. Here, the user determines the visual type (e.g., 2D or 3D), viewing style, style (e.g., Ghibli style), time period, and details about the characters and environment. The user's selections are received as input, and the terminal compiles them into a single configuration file. The configuration file is generated as output.
[0731] Step 3:
[0732] The terminal sends the prepared content data and configuration files to the server. The server receives these and performs data analysis as input. Using a generative AI model, it extracts important components from the data and performs initial setup of the visual style and narrative structure. The analysis results are obtained as output.
[0733] Step 4:
[0734] The server determines the visual style and generates characters and environments based on the analysis results and user settings. The generating AI model constructs a specific screen using a particular color scheme and line style. The analysis results and configuration files are used as input, and an initial version of the visual content is generated as output.
[0735] Step 5:
[0736] The server adds interactive elements to the visual content as needed, for educational or business purposes. This includes interactive quiz formats and explanatory pop-ups. User-specific requests are referenced as input, and interactive content is generated as output.
[0737] Step 6:
[0738] The server generates translations and audio narration in various languages based on the selected language settings. Natural language processing technology is used, with the content's language data as input. The output consists of translated subtitle files and audio files.
[0739] Step 7:
[0740] Ultimately, the server renders the completed visual content in high quality and sends it to the user's device. Animation, audio, and subtitles are integrated as input, and the complete version of the visual content is generated as output. The user can preview this on their device and make adjustments or regenerate it as needed.
[0741] (Application Example 1)
[0742] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0743] Traditional animation generation systems struggled to provide interactive experiences that fully reflected the individual needs of users. Especially in educational and business settings, there was a demand for advanced interactions utilizing eye-tracking and voice input. Furthermore, there was a lack of technology to deliver real-time customized content based on diverse user-defined options.
[0744] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0745] In this invention, the server includes means for receiving content data and extracting important elements based on the content data; means for applying animation styles corresponding to image and video types based on multiple setting options selected by the user; and means for enabling interaction with the user using eye tracking and voice input. This allows the user to experience customized interactive animations in real time and enables operation using eye tracking and voice in educational and business applications.
[0746] "Content data" refers to informational materials provided by users, including data in formats such as video, audio, and text, which form the basis for animation generation.
[0747] "Methods for extracting important elements" refers to technologies that analyze content data to automatically select key points in story development and visual expression.
[0748] "Means of applying animation styles" refers to techniques for reflecting specific visual representations and dynamic configurations in generated animations based on user settings.
[0749] "Means for generating characters and backgrounds" refers to technology for designing digital characters and backgrounds according to the era and theme selected by the user, based on setting options.
[0750] "Methods for incorporating interactive elements" refer to techniques that introduce features into animations that allow users to make selections and perform actions within the animation.
[0751] "Means for generating translations and audio narration" refers to technologies that enable the creation of subtitles and audio descriptions for animations in multiple languages.
[0752] "Means that enable interaction using eye-tracking and voice input" refers to technology that detects the user's gaze and voice, responds to them as input, and enables two-way communication with animations.
[0753] "Means of delivering rendered animations" refers to technologies that send completed animations to users, making them viewable on their devices.
[0754] To implement this invention, the user must first upload content data that they wish to animate using an information and communication terminal such as a smartphone. This content data includes video, audio, and text. The user selects several options, such as visual style, language settings, and interactive elements, through the terminal's interface.
[0755] The server receives data and configuration options from this user. The server uses a generative AI model as its software. The generative AI model extracts key elements from the content data and applies a specified animation style. This also includes technologies that allow the user to interactively manipulate the animation using eye tracking and voice input. Specific hardware uses such as smart glasses, and an application that works with them is installed on a smartphone.
[0756] The generated animation is rendered, including translation and voice narration. The server delivers the completed animation to the user's device or smart glasses. The user can watch the animation and use their gaze and voice to answer quizzes and make selections.
[0757] One concrete example is an interactive animation created for history lessons in educational settings. This animation depicts the Sengoku period (Warring States period) of Japan, allowing viewers to select character actions using their gaze and answer quizzes via voice input. An example of a prompt for a generative AI model would be: "Generate an interactive animation set in Sengoku-era Japan. The visual style should be Ukiyo-e (Japanese woodblock print), and English narration should be included. Questions should be displayed in a quiz format based on user selections, and answers should be selectable via eye-tracking."
[0758] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0759] Step 1:
[0760] The user uploads content data to be animated using an information and communication terminal. Input data accepted includes video files, audio data, and text documents. Output data is sent from the terminal to the server. Specifically, file selection and upload operations are performed through the user interface.
[0761] Step 2:
[0762] On the device, the user selects setting options such as visual style, video type, and language settings. The data received as input is the user's selection information. The setting options are sent to the server as output. Specifically, the options for each choice are displayed as pull-down menus or radio buttons for selection.
[0763] Step 3:
[0764] The server receives content data and user configuration options. The input data received consists of content data and configuration options sent from the terminal. The output provides the basic data necessary for the next processing step. Specifically, the server stores the received data and prepares to begin analysis.
[0765] Step 4:
[0766] The server uses a generative AI model to extract key elements from content data. The input data is content data. The output is key points extracted from text and visuals. Specifically, the generative AI model analyzes the data and uses a particular algorithm to extract the main themes and storylines.
[0767] Step 5:
[0768] The server applies animation styles based on the user's selected settings to generate characters and backgrounds. The input data received consists of the settings options and key elements. The output is styled animated content. Specifically, the AI incorporates visual effects and designs into the animation according to the selected style.
[0769] Step 6:
[0770] The server configures the interaction using eye tracking and voice input. The data received as input includes hardware information and user settings. The output is the animation in which the interactive elements are reflected. Specifically, the server analyzes data from the eye tracking device and microphone and configures the program to reflect the results in the animation's operation.
[0771] Step 7:
[0772] The server generates translations and voice narration in multiple languages. The input data received includes the content's text information and the user's language settings. The output consists of subtitles and voice narration in the specified language. Specifically, translation software and speech synthesis technology are used to generate the text as narration and subtitles.
[0773] Step 8:
[0774] The server renders the final generated animation and delivers it to the user's device or smart glasses. The data received as input is the animation data to be rendered. The animation is delivered as output in a format that the user can view. Specifically, the file is sent via a data transfer protocol and played back on the device.
[0775] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0776] The animation generation system of the present invention incorporates an emotion engine to recognize the user's emotional state in real time and adjust the animation based on that state. An embodiment thereof is described below in natural language.
[0777] First, users use their devices to upload the content data they wish to animate to the system. This content data includes videos, audio, text, and other formats. Next, through the device's interface, users select the animation's visual style, video type, time period, story theme, language settings, and even detailed character settings.
[0778] This data and settings are sent from the terminal to the server. The server analyzes the content data, extracts key elements, and applies animation styles based on the user's selections. The server also generates characters and backgrounds according to the time period and theme, and adds interactive elements to the animation.
[0779] The emotion engine analyzes the user's emotional state through the camera and biosensors built into the device. Based on the data from the emotion engine, the server adjusts the animation scenes, character expressions, and voice tone. This functionality allows the animation to react to the user's emotions, providing a more personalized experience.
[0780] For example, if the user expresses happiness, the characters in the animation will smile, and their voice tone will become brighter. Conversely, if the user is sad, the characters' expressions will become slightly calmer, and their voice tone will become gentler.
[0781] Finally, the server sends the rendered animation, based on all settings and adjustments, to the user's device. The user can preview this animation and verify that the emotion-responsive customizations have been properly implemented. This allows users to easily create more interactive and personalized animations.
[0782] The following describes the processing flow.
[0783] Step 1:
[0784] The user uploads the content data (video, text, etc.) they want to animate to the system using their device. The user then selects the animation image type, video type, time period, story theme, language setting, and character settings through the interface.
[0785] Step 2:
[0786] The device packages the uploaded content data and the user's selected settings information and sends it to the server.
[0787] Step 3:
[0788] The server analyzes the received content data and extracts important elements. In the case of video, a scene segmentation algorithm is used to identify meaningful scenes.
[0789] Step 4:
[0790] The server applies animation styles based on user selections. This includes filtering selected image types and applying color palettes.
[0791] Step 5:
[0792] The server automatically generates character clothing and backgrounds according to the time period and story theme. For example, if the selected era is the Showa period, the design of the clothing and background will change to suit that style.
[0793] Step 6:
[0794] The emotion engine uses the device's built-in camera or biosensors to recognize the user's current emotions. This data is sent to a server.
[0795] Step 7:
[0796] The server adjusts the animation scenes, character expressions, and voice tone based on the emotion engine data. If the user perceives the content as enjoyable, the character will smile and a cheerful voice tone will be used.
[0797] Step 8:
[0798] The server incorporates interactive elements into the animation, inserting quizzes, pop-ups, and other features tailored to educational and business applications.
[0799] Step 9:
[0800] The server generates voice narration and translated subtitles based on the language selected by the user.
[0801] Step 10:
[0802] The server renders the final animation and sends this completed file to the user's terminal. The user can then review the animation and request modifications or regeneration as needed.
[0803] (Example 2)
[0804] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0805] Conventional motion expression generation systems have difficulty achieving interactive expressions that reflect user emotions in real time, and also have limited support for multiple languages and customization of characters and backgrounds in terms of historical setting. As a result, they have been unable to generate motion expressions that match the individual user's situation and intentions, leading to a decline in the quality of the user experience.
[0806] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0807] In this invention, the server includes means for receiving content information and extracting key elements based on said content information; means for applying motion expression styles corresponding to image and video formats based on multiple setting options selected by the user; means for generating characters and backgrounds based on the selected settings; means for incorporating interactive elements into the motion expression; and means for recognizing the user's emotions using an emotion analysis function and adjusting the scenes, character expressions, and sound tones of the motion expression based on the recognized emotions. This makes it possible to generate interactive and personalized motion expressions that respond to the user's emotions and settings.
[0808] "Content information" refers to media data provided by users, such as text, audio, and video.
[0809] "Key elements" refer to important information and features contained in content information that are extracted in order to understand and process the content.
[0810] "Setting options" refer to the choices that users can make to determine the animation's visual style, video type, time period, theme, language setting, and character details.
[0811] "Action style" refers to the style that defines the appearance and movement of an animation.
[0812] "Characters appearing in the animation" refers to characters such as people and animals that are depicted within the generated animation.
[0813] "Background" refers to the environment or scenery in which the characters exist within an animation.
[0814] "Interactive elements" refer to elements that enable two-way interaction with the user and add interactivity to animations.
[0815] "Emotional analysis function" refers to technology used to recognize and analyze a user's emotional state.
[0816] "Action expression" refers to the entire animation generated based on the user's emotions and settings.
[0817] "Acoustic tone" refers to the pitch and quality of the sounds used in animation.
[0818] This invention is a system that generates personalized animations in real time according to the user's emotions and settings. The system operates primarily via a terminal and a server. The terminal provides an interface for the user to upload content information and specify subsequent setting options. The terminal then analyzes the user's emotions using a built-in camera and biosensors. This hardware is used to acquire the user's real-time emotional data.
[0819] The server extracts key elements based on the received content information and applies behavioral expression styles according to the setting choices. This process utilizes natural language processing (NLP) and computer vision technologies to understand and analyze the content information. The server also uses generative AI models to design characters and backgrounds, and to adjust scenes, facial expressions, and sound tones based on data obtained from emotion analysis functions. Specific software includes, for example, software commonly used in 3D modeling.
[0820] For example, if a user wants to create an animation featuring their pet, a possible prompt might be "Generate a fun, dynamic animation themed around the adventures of a pet dog." Based on this prompt, the AI model generates characters and backgrounds to depict an adventure story starring a pet dog, and the animation's actions are customized based on the user's settings and emotions.
[0821] This system allows users to easily create personalized behavioral expressions and enjoy highly interactive experiences, especially those that respond to emotions.
[0822] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0823] Step 1:
[0824] Users upload content information (e.g., videos, audio, text) to the system using their device. They select the necessary content via the device's file selection function. The input is the media data selected by the user, and the output is that data sent to the server.
[0825] Step 2:
[0826] The server analyzes the received content information and extracts key elements. Natural language processing and computer vision technologies are used for the analysis. The input is media data uploaded by the user, and the output is a list of important elements from the data. Specifically, this includes object recognition within videos and keyword extraction from text.
[0827] Step 3:
[0828] The user inputs animation settings from their device. These include visual style, video format, time period, story theme, language settings, and character details. The input represents the user's desired settings, and this information is sent to the server as output. Settings are configured using dropdown menus and sliders on the device's interface.
[0829] Step 4:
[0830] The server selects a behavioral expression style based on the user's settings. It integrates the received configuration information and applies the appropriate style using a generative AI model. The input is the extracted content and the user's configuration information, and the output is the applied animation style.
[0831] Step 5:
[0832] The server generates characters and backgrounds according to the time period and theme. A generation AI model is used here to create prompt statements. This allows the AI to design the appearance of the characters and their backgrounds. The input consists of user settings and content information elements, and the output is the generated characters and backgrounds. Specifically, a prompt statement might be something like, "Generate a fun, animated video about a pet dog's adventure."
[0833] Step 6:
[0834] The device uses its built-in camera and biosensors to analyze the user's emotions in real time. This data is sent to a server and used for emotion analysis. The input is emotion data obtained from the user, and the output is the analyzed emotion information. The server adjusts the animation according to the emotional state.
[0835] Step 7:
[0836] The server adjusts the animation scenes, character expressions, and sound tones based on the results of the emotion analysis. This enables personalized action expressions that respond to the user's emotions. The input is the analysis results, and the output is the adjusted animation.
[0837] Step 8:
[0838] Ultimately, the server renders the animation based on all the settings and adjustments. The input is the adjusted animation data, and the output is a rendered, high-quality animation file.
[0839] Step 9:
[0840] The server sends the rendered animation to the terminal, and the user previews it on the terminal. At this stage, the user verifies that the animation is represented as intended. The input is the rendered animation file, and the output is what is displayed to the user.
[0841] (Application Example 2)
[0842] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0843] There is a challenge in providing personalized visual representations that respond to user emotions in real time. In particular, when content is fixed, it is difficult to respond immediately to changes in the viewer's emotions and provide an individualized experience. Furthermore, smoothly implementing changes in voice tone and facial expressions that match the viewer's emotions while supporting multiple languages is also a technical challenge.
[0844] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0845] In this invention, the server includes means for receiving content data and extracting important elements based on the content data; means for applying visual styles corresponding to image and video types based on a plurality of setting options selected by the user; and means for detecting the user's emotional state through emotion analysis and adjusting the facial expressions and voice tone of the characters accordingly. This makes it possible to generate and deliver personalized visual representations that match the user's emotional state in real time.
[0846] "Content data" refers to information used as material for generating visual representations, and includes formats such as video, audio, and text.
[0847] A "means for extracting important elements" refers to a system that has the function of identifying and analyzing information necessary for visual expression from content data.
[0848] "Settings options" include multiple specification and style choices that users can select when customizing the visual presentation.
[0849] "Visual style" refers to the overall characteristics that form the distinctive features of design and expression in visual representation.
[0850] "Means for generating characters and backgrounds" refers to a system that has a process for creating characters and backgrounds within a visual representation based on set specifications.
[0851] "Interactive elements" refer to components that allow users to interact with visual representations.
[0852] "Means for generating translations and musical narration" refers to a system that supports multiple languages and has a process for creating audio descriptions necessary for visual representations, either mechanically or manually.
[0853] "Emotional analysis" is a technology that analyzes biometric data and facial expression data to understand the user's psychological state.
[0854] A "means of delivering visual representations" refers to a system that has a process for delivering the final created visual representation to the user.
[0855] This invention is a system that generates personalized visual representations in response to the user's emotional state. The server receives content data and extracts important elements from it. Specific software used for data analysis includes "OpenCV" and "EmotionAPI." The user selects setting options to determine the style of the visual representation via a terminal. Based on this, the server uses tools such as "Unity" to apply the visual style and generate characters and backgrounds.
[0856] The emotion analysis process utilizes the camera and biosensors installed in the device to analyze data such as the user's facial expressions and heart rate in real time. Based on the data obtained from this analysis, the facial expressions and voice tones of the characters in the visual representation are adjusted. The server then delivers the final rendered visual representation to the user's device.
[0857] For example, if a user smiles while watching an emotionally moving scene in a visual presentation, the server can detect this and adjust the characters' expressions to reflect positive emotions, and change the music to a more upbeat tone. This system also has translation capabilities for different languages, generates musical narration, and provides a smooth viewing experience.
[0858] An example of a prompt to input into the generation AI model would be, "Generate an animation for when the user is laughing. Make the character brighter and the voice tone more energetic." In this way, it is possible to provide interactive visual representations that are tailored to the individual user's emotions in real time.
[0859] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0860] Step 1:
[0861] Users upload content data they wish to animate using their device. This input data includes video, audio, and text data. The device sends this data to the server, which receives it. Based on the received data, the server extracts key elements using OpenCV or the Emotion API. This process involves data analysis and sampling.
[0862] Step 2:
[0863] The user selects setting options to determine the style of visual representation through the terminal interface. Inputs include visual style, video type, time period, story theme, language settings, and character details. This data is sent from the terminal to the server, which uses Unity to apply the visual style based on these options and generate characters and backgrounds. Specifically, the program calls the core graphics module and begins rendering according to the settings.
[0864] Step 3:
[0865] The system uses the camera and biosensors built into the device to analyze the user's emotional state in real time. Inputs include the user's facial expression data and heart rate. The server converts this into emotional data based on the EmotionAPI and obtains the analysis results. Based on this data, the system adjusts the facial expressions and voice tones of the characters in the visual representation. Specifically, the character's animation module is dynamically adjusted based on the input emotional data.
[0866] Step 4:
[0867] The server integrates all settings and analysis results and finally renders the visual representation. This includes generated characters, backgrounds, interactive elements, and translated musical narration. This output data is delivered to the terminal as a visual representation. The visual representation is updated and provided to the user in real time. The user confirms that the customization has been done correctly by viewing it.
[0868] Step 5:
[0869] In this application example, the visual representation changes according to the user's emotions. For example, if the emotion analysis detects that the user is smiling, the characters will have cheerful expressions, and the music will change to a more energetic tone. This allows the user to experience personalized visual representations. The prompts used for the generative AI model would be input statements such as, "Generate an animation for when the user is smiling. Make the characters brighter and the voice tone more energetic."
[0870] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0871] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0872] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0873] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0874] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0875] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0876] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0877] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0878] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0879] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0880] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0881] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0882] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0883] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0884] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0885] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0886] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0887] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0888] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0889] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0890] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0891] The following is further disclosed regarding the embodiments described above.
[0892] (Claim 1)
[0893] A means for receiving content data and extracting important elements based on said content data,
[0894] A means of applying animation styles corresponding to image and video types based on multiple setting options selected by the user,
[0895] A means of generating characters and backgrounds based on selected settings,
[0896] Means of incorporating interactive elements into animations,
[0897] A means for generating translations and voice narrations that support multiple languages,
[0898] The means of delivering the final rendered animation to the user,
[0899] A system that includes this.
[0900] (Claim 2)
[0901] The system according to claim 1, further comprising means for customizing the character's clothing and background to suit the era, based on setting options selected by the user.
[0902] (Claim 3)
[0903] The system according to claim 1, further comprising means for providing an interactive function for users to generate animations specifically for educational or business purposes.
[0904] "Example 1"
[0905] (Claim 1)
[0906] A means for receiving information and extracting important components based on said information,
[0907] A means for applying expression techniques corresponding to visual type and viewing type based on multiple setting requests selected by the user,
[0908] A means of creating characters and environments based on the selected settings,
[0909] Means of incorporating bidirectional components into the representation,
[0910] A means of creating translations and audio narrations that support diverse languages,
[0911] A means of providing the final rendered representation to the user,
[0912] A system that includes this.
[0913] (Claim 2)
[0914] The system according to claim 1, further comprising means for adjusting the costumes and environment of characters to suit the era based on setting requests selected by the user.
[0915] (Claim 3)
[0916] The system according to claim 1, further comprising means for providing interactive functions for users to generate expressions specialized for learning or business use.
[0917] "Application Example 1"
[0918] (Claim 1)
[0919] A means for receiving content data and extracting important elements based on said content data,
[0920] A means of applying animation styles corresponding to image and video types based on multiple setting options selected by the user,
[0921] A means of generating characters and backgrounds based on selected settings,
[0922] Means of incorporating interactive elements into animations,
[0923] A means for generating translations and voice narrations that support multiple languages,
[0924] A means to enable user interaction using eye tracking and voice input,
[0925] The means of delivering the final rendered animation to the user,
[0926] A system that includes this.
[0927] (Claim 2)
[0928] The system according to claim 1, further comprising means for customizing the character's clothing and background to suit the era, based on setting options selected by the user.
[0929] (Claim 3)
[0930] The system according to claim 1, further comprising means for providing an interactive function for a user to generate animations specifically for educational or business purposes, and for receiving user selections via eye tracking.
[0931] "Example 2 of combining an emotion engine"
[0932] (Claim 1)
[0933] A means for receiving content information and extracting key elements based on said content information,
[0934] A means for applying motion expression styles corresponding to image and video formats based on multiple setting options selected by the user,
[0935] A means of generating characters and backgrounds based on selected settings,
[0936] Means for incorporating interactive elements into action expressions,
[0937] A means of recognizing the user's emotions using an emotion analysis function, and adjusting the action expression scene, character's facial expressions, and sound tone based on the recognized emotions,
[0938] A means of generating multilingual translations and audio narrations,
[0939] A means of delivering the final rendered motion representation to the user,
[0940] A system that includes this.
[0941] (Claim 2)
[0942] The system according to claim 1, further comprising means for customizing the costumes and backgrounds of characters that appear in accordance with the era, based on setting options selected by the user.
[0943] (Claim 3)
[0944] The system according to claim 1, further comprising means for providing an interactive function for users to generate action expressions specialized for educational or business purposes.
[0945] "Application example 2 when combining with an emotional engine"
[0946] (Claim 1)
[0947] A means for receiving content data and extracting important elements based on said content data,
[0948] A means of applying visual styles corresponding to image and video types based on multiple setting options selected by the user,
[0949] A means of generating characters and backgrounds based on selected settings,
[0950] Means of incorporating interactive elements into visual representations,
[0951] A means of generating translations and musical narrations that support multiple languages,
[0952] A means of detecting the user's emotional state through emotion analysis and adjusting the facial expressions and voice tone of the characters accordingly,
[0953] The means of delivering the final created visual representation to the user,
[0954] A system that includes this.
[0955] (Claim 2)
[0956] The system according to claim 1, further comprising means for customizing the clothing and background of characters according to the era based on setting options selected by the user.
[0957] (Claim 3)
[0958] The system according to claim 1, further comprising means for providing an interactive function for users to generate visual representations specifically for educational or professional use. [Explanation of Symbols]
[0959] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving content data and extracting important elements based on said content data, A means of applying animation styles corresponding to image and video types based on multiple setting options selected by the user, A means of generating characters and backgrounds based on selected settings, Means of incorporating interactive elements into animations, A means for generating translations and voice narrations that support multiple languages, The means of delivering the final rendered animation to the user, A system that includes this.
2. The system according to claim 1, further comprising means for customizing the character's clothing and background according to the era based on setting options selected by the user.
3. The system according to claim 1, further comprising means for providing an interactive function for users to generate animations specifically for educational or business purposes.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A