system
The system addresses the inefficiencies in CG generation and customization by using AI to automate and personalize the process, enabling efficient and interactive video production.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional methods for generating and customizing computer graphics (CG) are time-consuming and labor-intensive, making it difficult for viewers to personalize their preferences.
A system comprising a reception unit, generation unit, modification unit, and customization unit, utilizing AI to automatically generate, modify, and customize CG based on creator and viewer inputs, with features like deep learning models and intuitive interfaces for correction and customization tools.
Streamlines CG generation and modification, allowing creators to efficiently produce high-quality, personalized content, and enables viewers to interactively customize videos to their preferences.
Smart Images

Figure 2026061843000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that it takes a great deal of time and labor to generate and modify CG, and it is difficult for viewers to customize it according to their own preferences.
[0005] The system according to the embodiment aims to improve the efficiency of generating and modifying CG and enable viewers to customize it according to their own preferences.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a reception unit, a generation unit, a modification unit, and a customization unit. The reception unit receives the creator's requirements as input. The generation unit automatically generates CG based on the requirements input by the reception unit. The modification unit checks and modifies the CG generated by the generation unit. The customization unit allows viewers to customize the CG modified by the modification unit. [Effects of the Invention]
[0007] The system according to this embodiment can streamline the generation and modification of CG and allow viewers to customize it to their own preferences. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The video generation system according to an embodiment of the present invention is a system that automatically generates CG for movies and videos using a generation AI. In this video generation system, a creator inputs basic settings and requirements for the video to the generation AI, and the generation AI automatically generates CG based on that information. The generated CG is reviewed by the creator and modified as needed. Furthermore, viewers can enjoy interactive videos tailored to their preferences. This system allows creators to overcome the constraints of time and technology and produce innovative and creative videos. In addition, viewers can participate more deeply in video production through an interactive video experience. This promotes the formation of new forms of entertainment by enabling both viewers and creators to participate in creative video production. For example, a creator inputs basic settings and requirements for the video to the generation AI. The generation AI automatically generates CG based on that information. The generated CG is reviewed by the creator and modified as needed. Viewers can enjoy interactive videos tailored to their preferences. This allows creators to overcome the constraints of time and technology and produce innovative and creative videos. In addition, viewers can participate more deeply in video production through an interactive video experience. This allows the video generation system to automatically generate CG based on the creator's requirements, and to be modified and customized.
[0029] The video generation system according to this embodiment comprises a reception unit, a generation unit, a modification unit, and a customization unit. The reception unit receives the creator's requirements as input. The creator's requirements include, but are not limited to, text format, image format, audio format, etc. The generation unit uses a generation AI to automatically generate CG based on the requirements input by the reception unit. The generation AI generates CG using, for example, a specific AI model and training method. The modification unit allows the creator to review the CG generated by the generation unit and modify it as necessary. Modifications are made, for example, based on the scope of modifications and the tools used. The customization unit allows viewers to customize the CG modified by the modification unit. Customization is made, for example, based on customizable elements and the interface used. This enables the video generation system to automatically generate, modify, and customize CG based on the creator's requirements.
[0030] The reception desk receives the creator's requirements. These requirements may include, but are not limited to, text, image, and audio formats. Specifically, creators can enter detailed descriptions of scenes and character characteristics in text format. For example, they can enter specific requirements such as "a scene of waves gently lapping on a beach at sunset" or "a girl with red hair wearing a blue dress." In image format, creators can upload reference images or sketches, and CG can be generated based on them. In audio format, creators can describe the atmosphere of a scene or the tone of a character's voice using audio. This allows the reception desk to flexibly accept diverse requirements from creators and process them appropriately as input data for the generation department. Furthermore, the reception desk also has the function to organize the creator's requirements and request additional information as needed. For example, if the entered requirements are unclear or incomplete, it can ask the creator specific questions to gather more detailed information. This allows the reception desk to provide sufficient information for the generation department to produce high-quality CG.
[0031] The generation unit uses a generation AI to automatically generate CG based on the requirements entered by the reception unit. The generation AI generates CG using, for example, specific AI models and training methods. Specifically, the generation AI utilizes deep learning technology and uses models trained on large image datasets. For example, the generation AI can generate realistic CG using models such as GAN (Generative Opposite Network) and VAE (Variational Autoencoder). Based on the creator's requirements, the generation AI automatically generates the scene composition, character design, background details, etc. For example, for the requirement of "a scene of waves gently lapping on a beach at sunset," the generation AI realistically reproduces the sunset sky, wave movement, and sandy beach. Also, for the requirement of "a girl with red hair wearing a blue dress," the generation AI faithfully reproduces the character's hair color and dress design. The generation AI analyzes the creator's requirements and automatically adjusts the parameters to generate the optimal CG. This allows the generation unit to quickly generate high-quality CG based on the creator's requirements. Furthermore, the generation unit also has a function to temporarily save the generated CG, making it accessible to the modification and customization units. This allows the generation unit to support the efficient operation of the entire system.
[0032] The editing section allows creators to review the CG generated by the generation section and make corrections as needed. Corrections are made based on factors such as the scope of the correction and the tools used. Specifically, creators can preview the generated CG and make detailed adjustments and corrections. For example, they can fine-tune character expressions and poses, background colors, and lighting. The editing section provides an intuitive interface for creators, making correction work easy. It offers a variety of correction tools, such as drag-and-drop operations, slider adjustments, and brush tools for drawing. Furthermore, the editing section incorporates AI-assisted features to support creators in efficiently performing corrections. For example, the AI can learn the creator's correction history and predict and suggest future corrections. The editing section also includes a function that allows creators to save their corrections for later re-editing. This enables creators to thoroughly review the generated CG and make corrections until they are satisfied.
[0033] The customization section allows viewers to customize the CG images modified by the modification section. Customization is based on, for example, the customizable elements and the interface used. Specifically, viewers can freely change character costumes and accessories, background designs, effects, and more. The customization section provides an intuitive interface that viewers can easily use to perform customization tasks. For example, viewers can change character costumes or adjust background colors using drag-and-drop operations or sliders. Furthermore, the customization section also has a function to save and share the customized content with other viewers. This allows viewers to create and enjoy their own original CG images. The customization section can collect viewer feedback and use it to improve customization features and add new elements. For example, new costumes, accessories, and effects can be added based on viewer requests. In this way, the customization section can provide an environment where viewers can customize and enjoy CG images to their liking.
[0034] The video generation system includes a learning unit in which a generating AI generates computer graphics (CG) based on training data. The learning unit generates CG based on training data such as image data, text data, and audio data. For example, the learning unit can have the generating AI learn from a large amount of image data and generate high-precision CG based on that data. The learning unit can also have the generating AI learn from text data and generate CG based on that data. Furthermore, the learning unit can have the generating AI learn from audio data and generate CG based on that data. This improves the generation accuracy by having the generating AI generate CG based on training data. Some or all of the above-described processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input image data into the generating AI and have the generating AI generate CG from the image data.
[0035] The generation unit can automatically generate CG based on the creator's requirements using a generation AI. For example, the generation unit automatically generates CG based on the creator's requirements using a generation AI. The generation AI generates CG using, for example, a specific AI model or training method. The generation unit generates CG based on text data entered by the creator. The generation unit can also generate CG based on image data entered by the creator. Furthermore, the generation unit can generate CG based on audio data entered by the creator. This makes it possible to automatically generate CG based on the creator's requirements by using a generation AI. Some or all of the above processes in the generation unit may be performed using, for example, AI, or not using AI. For example, the generation unit can input the creator's requirements into the generation AI and have the generation AI execute the generation of CG based on those requirements.
[0036] The editing unit allows creators to review the CG generated by the generation unit and make corrections as needed. For example, the editing unit allows creators to review the CG generated by the generation unit and make corrections as needed. Corrections are made based on, for example, the scope of the correction and the tools used. The editing unit can, for example, correct the color tone of the CG. The editing unit can also correct the shape of the CG. Furthermore, the editing unit can correct the movement of the CG. This allows creators to review the generated CG and make corrections as needed, thereby providing high-quality CG. Some or all of the above processing in the editing unit may be performed using, for example, AI, or not using AI. For example, the editing unit can input the generated CG into a generation AI and have the generation AI perform corrections on the CG.
[0037] The customization section allows viewers to customize the CG that has been modified by the modification section. For example, the customization section allows viewers to customize the CG that has been modified by the modification section. Customization is performed, for example, based on customizable elements or the interface used. For example, the customization section can customize the color of the CG. The customization section can also customize the shape of the CG. Furthermore, the customization section can customize the movement of the CG. This enables an interactive video experience by allowing viewers to customize the modified CG. Some or all of the above processing in the customization section may be performed using AI, for example, or without AI. For example, the customization section can input the modified CG into a generating AI and have the generating AI perform the CG customization.
[0038] The reception desk can analyze a creator's past requirement input history and suggest the optimal input method. For example, the reception desk can automatically display requirements that the creator has frequently entered in the past as suggestions. The reception desk can also prioritize suggesting input methods (voice, text, etc.) that the creator has used in the past. Furthermore, the reception desk can predict and suggest requirements to be used during specific time periods based on the creator's past input history. In this way, by analyzing past requirement input history, the reception desk can suggest the optimal input method for the creator. Some or all of the above processing in the reception desk may be performed using AI, for example, or not using AI. For example, the reception desk can input the creator's past input history data into a generating AI and have the generating AI suggest the optimal input method.
[0039] The reception system can filter requirements based on the creator's current projects and areas of interest when they are entered. For example, the reception system will prioritize displaying requirements related to the creator's current projects. The reception system can also automatically suggest relevant requirements based on the creator's areas of interest. Furthermore, the reception system can filter requirements based on areas the creator has shown interest in in the past. This allows for the priority input of highly relevant requirements by filtering requirements based on the creator's current projects and areas of interest. Some or all of the above processing in the reception system may be performed using AI, for example, or not. For example, the reception system can input the creator's project data and area of interest data into a generating AI and have the generating AI perform the requirement filtering.
[0040] The reception desk can prioritize highly relevant requirements when a creator enters their requirements, taking into account their geographical location. For example, if a creator is in a specific region, the reception desk will prioritize requirements related to that region. Furthermore, if a creator is on the move, the reception desk can suggest the most relevant requirements based on their current location. Additionally, if a creator is in a specific location, the reception desk can prioritize requirements related to that location. This allows for the prioritization of highly relevant requirements by considering the creator's geographical location. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk can input the creator's geographical location data into a generating AI and have the generating AI determine the priority of the requirements.
[0041] The reception desk can analyze the creator's social media activity and input relevant requirements when requirements are entered. For example, the reception desk can suggest relevant requirements based on information the creator has shared on social media. The reception desk can also identify areas of interest from the creator's social media activity and input requirements based on those areas. Furthermore, the reception desk can suggest relevant requirements based on the accounts the creator follows on social media. This allows for the input of relevant requirements by analyzing the creator's social media activity. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the creator's social media data into a generating AI and have the generating AI generate requirement suggestions.
[0042] The generation unit can select the optimal generation algorithm by referring to the creator's past works during generation. For example, the generation unit can select the optimal generation algorithm based on works created by the creator in the past. The generation unit can also extract a specific style from the creator's past works and select a generation algorithm based on that. Furthermore, the generation unit can analyze the creator's past works and select the most effective generation algorithm. In this way, the optimal generation algorithm can be selected by referring to the creator's past works. Some or all of the above processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the creator's past work data into a generation AI and have the generation AI perform the selection of a generation algorithm.
[0043] The generation unit can adjust the level of detail of the generated image based on the creator's current project. For example, the generation unit can adjust the level of detail based on the project the creator is currently working on. It can also adjust the level of detail based on the progress of the creator's project. Furthermore, the generation unit can adjust the level of detail based on the requirements of the creator's project. This allows for the generation of CG optimized for the project by adjusting the level of detail based on the creator's current project. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the creator's project data into a generation AI and have the generation AI perform the adjustment of the level of detail of the generated image.
[0044] The generation unit can select the optimal generation method during generation, taking into account the creator's geographical location information. For example, if the creator is in a specific region, the generation unit will prioritize generating CG related to that region. Furthermore, if the creator is on the move, the generation unit can select the optimal generation method based on their current location. Additionally, if the creator is in a specific location, the generation unit can prioritize generating CG related to that location. This allows the generation unit to select the optimal generation method by considering the creator's geographical location information. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the creator's geographical location data into a generation AI and have the generation AI select the generation method.
[0045] The generation unit can analyze the creator's social media activity during generation and propose generation methods. For example, the generation unit can propose the optimal generation method based on information shared by the creator on social media. The generation unit can also identify areas of interest from the creator's social media activity and propose generation methods based on those areas. Furthermore, the generation unit can propose the optimal generation method based on the accounts the creator follows on social media. In this way, the optimal generation method can be proposed by analyzing the creator's social media activity. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the creator's social media data into a generation AI and have the generation AI execute the generation method proposal.
[0046] The editing unit can select the optimal editing method by referring to the creator's past editing history during editing. For example, the editing unit can select the optimal editing method based on the editing the creator has done in the past. The editing unit can also extract a specific style from the creator's past editing history and select an editing method based on that. Furthermore, the editing unit can analyze the creator's past editing history and select the most effective editing method. In this way, the optimal editing method can be selected by referring to the creator's past editing history. Some or all of the above processes in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can input the creator's past editing history data into a generating AI and have the generating AI perform the selection of editing methods.
[0047] The editing unit can adjust the level of detail of the edits based on the creator's current project. For example, the editing unit can adjust the level of detail based on the project the creator is currently working on. It can also adjust the level of detail based on the progress of the creator's project. Furthermore, the editing unit can adjust the level of detail based on the requirements of the creator's project. This allows for optimal edits to be made based on the creator's current project. Some or all of the above processes in the editing unit may be performed using AI, for example, or not using AI. For example, the editing unit can input the creator's project data into a generating AI and have the generating AI perform the adjustment of the level of detail of the edits.
[0048] The editing unit can select the optimal editing method by considering the creator's geographical location information during editing. For example, if the creator is in a specific region, the editing unit will prioritize editing CG related to that region. Furthermore, if the creator is on the move, the editing unit can select the optimal editing method based on their current location. Additionally, if the creator is in a specific location, the editing unit can prioritize editing CG related to that location. This allows the optimal editing method to be selected by considering the creator's geographical location information. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can input the creator's geographical location data into a generating AI and have the generating AI select the editing method.
[0049] The editing unit can analyze the creator's social media activity and propose editing methods during the editing process. For example, the editing unit can propose the most suitable editing method based on information shared by the creator on social media. The editing unit can also identify areas of interest from the creator's social media activity and propose editing methods based on those areas. Furthermore, the editing unit can propose the most suitable editing method based on the accounts the creator follows on social media. In this way, the editing unit can propose the most suitable editing method by analyzing the creator's social media activity. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can input the creator's social media data into a generating AI and have the generating AI execute the proposal of editing methods.
[0050] The customization unit can select the optimal customization method by referring to the viewer's past customization history during the customization process. For example, the customization unit can select the optimal customization method based on the viewer's past customizations. Furthermore, the customization unit can extract specific styles from the viewer's past customization history and select a customization method based on those styles. In addition, the customization unit can analyze the viewer's past customization history and select the most effective customization method. This allows the optimal customization method to be selected by referring to the viewer's past customization history. Some or all of the above-described processes in the customization unit may be performed using AI, for example, or without AI. For example, the customization unit can input the viewer's past customization history data into a generating AI and have the generating AI perform the selection of a customization method.
[0051] The customization unit can adjust the level of detail of customization based on the viewer's current areas of interest during the customization process. For example, the customization unit can adjust the level of detail of customization based on the areas the viewer is currently interested in. The customization unit can also adjust the level of detail of customization in response to changes in the viewer's areas of interest. Furthermore, the customization unit can adjust the level of detail of customization based on the viewer's current areas of interest. This allows for more appropriate customization by adjusting the level of detail of customization based on the viewer's current areas of interest. Some or all of the above-described processes in the customization unit may be performed using AI, for example, or without AI. For example, the customization unit can input viewer area of interest data into a generating AI and have the generating AI perform the adjustment of the level of detail of customization.
[0052] The customization unit can select the optimal customization method by considering the viewer's geographical location information during the customization process. For example, if the viewer is in a specific region, the customization unit will prioritize customizing CG related to that region. Furthermore, if the viewer is on the move, the customization unit can also select the optimal customization method based on their current location. Additionally, if the viewer is in a specific location, the customization unit can prioritize customizing CG related to that location. This allows for the selection of the optimal customization method by considering the viewer's geographical location information. Some or all of the above-described processes in the customization unit may be performed using AI, for example, or without AI. For example, the customization unit can input the viewer's geographical location data into a generating AI and have the generating AI perform the selection of the customization method.
[0053] The customization unit can analyze the viewer's social media activity during the customization process and propose customization methods. For example, the customization unit can propose the optimal customization method based on information shared by the viewer on social media. Furthermore, the customization unit can identify areas of interest from the viewer's social media activity and propose customization methods based on those areas. In addition, the customization unit can propose the optimal customization method based on the accounts the viewer follows on social media. This allows for the proposal of the optimal customization method by analyzing the viewer's social media activity. Some or all of the above processing in the customization unit may be performed using AI, for example, or without AI. For example, the customization unit can input the viewer's social media data into a generating AI and have the generating AI execute the proposal of customization methods.
[0054] The learning unit can optimize the learning algorithm by referring to past learning data during the learning process. For example, the learning unit can select the optimal learning algorithm based on past learning data. The learning unit can also extract specific patterns from past learning data and optimize the learning algorithm based on those patterns. Furthermore, the learning unit can analyze past learning data and select the most effective learning algorithm. This allows for the optimization of the learning algorithm by referring to past learning data. Some or all of the above processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input past learning data into a generating AI and have the generating AI perform the optimization of the learning algorithm.
[0055] The learning unit can weight the training data during training based on the timing of creator requirements input. For example, the learning unit weights the training data based on the timing of requirements previously entered by the creator. The learning unit can also adjust the weighting of the training data according to the timing of creator requirements input. Furthermore, the learning unit can weight the training data based on the timing of creator requirements input. This allows for more appropriate training by weighting the training data based on the timing of creator requirements input. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input creator requirements input timing data into a generating AI and have the generating AI perform the weighting of the training data.
[0056] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0057] The video generation system can also be equipped with a data integration unit. The data integration unit centrally manages the requirements and revisions entered by creators and provides them to the generation AI. For example, the data integration unit can integrate text data, image data, audio data, etc., entered by creators and input them into the generation AI. The data integration unit can also record the revisions made by creators and use them as training data for the generation AI. Furthermore, the data integration unit can analyze the creators' past requirements and revision history and provide feedback to improve the performance of the generation AI. In this way, the data integration unit can effectively manage the creators' requirements and revisions and maximize the performance of the generation AI.
[0058] The video generation system may also include a version control unit. The version control unit manages different versions of the generated CG, allowing creators to revert to previous versions as needed. For example, the version control unit automatically saves each version of the generated CG, making it easily accessible to creators. It can also record the revisions made by creators and track the revision history. Furthermore, the version control unit can suggest optimal revision methods based on the creator's past versions. This allows the version control unit to effectively manage different versions of the generated CG and revert to previous versions as needed.
[0059] The video generation system can also include a template section. This section provides predefined templates for creators to input requirements. For example, the template section can save frequently used requirements and settings as templates, which can then be easily recalled for future input. The template section can also analyze the creator's past requirement input history and suggest the most suitable template. Furthermore, the template section can allow creators to customize templates, supporting requirement input tailored to individual projects. This streamlines the creator's requirement input process and maximizes the performance of the generation AI.
[0060] The video generation system can also be equipped with an analysis unit. The analysis unit comprehensively analyzes generated CG, creator requirements input history, viewer feedback, and other data to provide insights for improving the performance of the generation AI. For example, the analysis unit can analyze the quality of generated CG and viewer reactions to provide feedback for improving the generation AI's algorithm. It can also analyze creator requirements input history and suggest ways to streamline future requirements input. Furthermore, the analysis unit can collect viewer feedback and provide insights based on that feedback to improve the generation AI's performance. This allows the analysis unit to continuously improve the generation AI's performance and provide creators and viewers with a higher quality video experience.
[0061] The following briefly describes the processing flow for example form 1.
[0062] Step 1: The reception desk enters the creator's requirements. Creator requirements may include, but are not limited to, text format, image format, audio format, etc. Step 2: The generation unit uses a generation AI to automatically generate computer graphics (CG) based on the requirements entered by the reception unit. The generation AI generates the CG using a specific AI model and training method. Step 3: The revision section involves the creator reviewing the CG generated by the generation section and making revisions as needed. Revisions are made based on the scope of the revisions and the tools used. Step 4: The customization section allows viewers to customize the CG that has been modified by the modification section. Customization is based on the customizable elements and the interface to be used.
[0063] (Example of form 2) The video generation system according to an embodiment of the present invention is a system that automatically generates CG for movies and videos using a generation AI. In this video generation system, a creator inputs basic settings and requirements for the video to the generation AI, and the generation AI automatically generates CG based on that information. The generated CG is reviewed by the creator and modified as needed. Furthermore, viewers can enjoy interactive videos tailored to their preferences. This system allows creators to overcome the constraints of time and technology and produce innovative and creative videos. In addition, viewers can participate more deeply in video production through an interactive video experience. This promotes the formation of new forms of entertainment by enabling both viewers and creators to participate in creative video production. For example, a creator inputs basic settings and requirements for the video to the generation AI. The generation AI automatically generates CG based on that information. The generated CG is reviewed by the creator and modified as needed. Viewers can enjoy interactive videos tailored to their preferences. This allows creators to overcome the constraints of time and technology and produce innovative and creative videos. In addition, viewers can participate more deeply in video production through an interactive video experience. This allows the video generation system to automatically generate CG based on the creator's requirements, and to be modified and customized.
[0064] The video generation system according to this embodiment comprises a reception unit, a generation unit, a modification unit, and a customization unit. The reception unit receives the creator's requirements as input. The creator's requirements include, but are not limited to, text format, image format, audio format, etc. The generation unit uses a generation AI to automatically generate CG based on the requirements input by the reception unit. The generation AI generates CG using, for example, a specific AI model and training method. The modification unit allows the creator to review the CG generated by the generation unit and modify it as necessary. Modifications are made, for example, based on the scope of modifications and the tools used. The customization unit allows viewers to customize the CG modified by the modification unit. Customization is made, for example, based on customizable elements and the interface used. This enables the video generation system to automatically generate, modify, and customize CG based on the creator's requirements.
[0065] The reception desk receives the creator's requirements. These requirements may include, but are not limited to, text, image, and audio formats. Specifically, creators can enter detailed descriptions of scenes and character characteristics in text format. For example, they can enter specific requirements such as "a scene of waves gently lapping on a beach at sunset" or "a girl with red hair wearing a blue dress." In image format, creators can upload reference images or sketches, and CG can be generated based on them. In audio format, creators can describe the atmosphere of a scene or the tone of a character's voice using audio. This allows the reception desk to flexibly accept diverse requirements from creators and process them appropriately as input data for the generation department. Furthermore, the reception desk also has the function to organize the creator's requirements and request additional information as needed. For example, if the entered requirements are unclear or incomplete, it can ask the creator specific questions to gather more detailed information. This allows the reception desk to provide sufficient information for the generation department to produce high-quality CG.
[0066] The generation unit uses a generation AI to automatically generate CG based on the requirements entered by the reception unit. The generation AI generates CG using, for example, specific AI models and training methods. Specifically, the generation AI utilizes deep learning technology and uses models trained on large image datasets. For example, the generation AI can generate realistic CG using models such as GAN (Generative Opposite Network) and VAE (Variational Autoencoder). Based on the creator's requirements, the generation AI automatically generates the scene composition, character design, background details, etc. For example, for the requirement of "a scene of waves gently lapping on a beach at sunset," the generation AI realistically reproduces the sunset sky, wave movement, and sandy beach. Also, for the requirement of "a girl with red hair wearing a blue dress," the generation AI faithfully reproduces the character's hair color and dress design. The generation AI analyzes the creator's requirements and automatically adjusts the parameters to generate the optimal CG. This allows the generation unit to quickly generate high-quality CG based on the creator's requirements. Furthermore, the generation unit also has a function to temporarily save the generated CG, making it accessible to the modification and customization units. This allows the generation unit to support the efficient operation of the entire system.
[0067] The editing section allows creators to review the CG generated by the generation section and make corrections as needed. Corrections are made based on factors such as the scope of the correction and the tools used. Specifically, creators can preview the generated CG and make detailed adjustments and corrections. For example, they can fine-tune character expressions and poses, background colors, and lighting. The editing section provides an intuitive interface for creators, making correction work easy. It offers a variety of correction tools, such as drag-and-drop operations, slider adjustments, and brush tools for drawing. Furthermore, the editing section incorporates AI-assisted features to support creators in efficiently performing corrections. For example, the AI can learn the creator's correction history and predict and suggest future corrections. The editing section also includes a function that allows creators to save their corrections for later re-editing. This enables creators to thoroughly review the generated CG and make corrections until they are satisfied.
[0068] The customization section allows viewers to customize the CG images modified by the modification section. Customization is based on, for example, the customizable elements and the interface used. Specifically, viewers can freely change character costumes and accessories, background designs, effects, and more. The customization section provides an intuitive interface that viewers can easily use to perform customization tasks. For example, viewers can change character costumes or adjust background colors using drag-and-drop operations or sliders. Furthermore, the customization section also has a function to save and share the customized content with other viewers. This allows viewers to create and enjoy their own original CG images. The customization section can collect viewer feedback and use it to improve customization features and add new elements. For example, new costumes, accessories, and effects can be added based on viewer requests. In this way, the customization section can provide an environment where viewers can customize and enjoy CG images to their liking.
[0069] The video generation system includes a learning unit in which a generating AI generates computer graphics (CG) based on training data. The learning unit generates CG based on training data such as image data, text data, and audio data. For example, the learning unit can have the generating AI learn from a large amount of image data and generate high-precision CG based on that data. The learning unit can also have the generating AI learn from text data and generate CG based on that data. Furthermore, the learning unit can have the generating AI learn from audio data and generate CG based on that data. This improves the generation accuracy by having the generating AI generate CG based on training data. Some or all of the above-described processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input image data into the generating AI and have the generating AI generate CG from the image data.
[0070] The generation unit can automatically generate CG based on the creator's requirements using a generation AI. For example, the generation unit automatically generates CG based on the creator's requirements using a generation AI. The generation AI generates CG using, for example, a specific AI model or training method. The generation unit generates CG based on text data entered by the creator. The generation unit can also generate CG based on image data entered by the creator. Furthermore, the generation unit can generate CG based on audio data entered by the creator. This makes it possible to automatically generate CG based on the creator's requirements by using a generation AI. Some or all of the above processes in the generation unit may be performed using, for example, AI, or not using AI. For example, the generation unit can input the creator's requirements into the generation AI and have the generation AI execute the generation of CG based on those requirements.
[0071] The editing unit allows creators to review the CG generated by the generation unit and make corrections as needed. For example, the editing unit allows creators to review the CG generated by the generation unit and make corrections as needed. Corrections are made based on, for example, the scope of the correction and the tools used. The editing unit can, for example, correct the color tone of the CG. The editing unit can also correct the shape of the CG. Furthermore, the editing unit can correct the movement of the CG. This allows creators to review the generated CG and make corrections as needed, thereby providing high-quality CG. Some or all of the above processing in the editing unit may be performed using, for example, AI, or not using AI. For example, the editing unit can input the generated CG into a generation AI and have the generation AI perform corrections on the CG.
[0072] The customization section allows viewers to customize the CG that has been modified by the modification section. For example, the customization section allows viewers to customize the CG that has been modified by the modification section. Customization is performed, for example, based on customizable elements or the interface used. For example, the customization section can customize the color of the CG. The customization section can also customize the shape of the CG. Furthermore, the customization section can customize the movement of the CG. This enables an interactive video experience by allowing viewers to customize the modified CG. Some or all of the above processing in the customization section may be performed using AI, for example, or without AI. For example, the customization section can input the modified CG into a generating AI and have the generating AI perform the CG customization.
[0073] The reception desk can estimate the creator's emotions and adjust the way requirements are entered based on the estimated emotions. For example, if the creator is stressed, the reception desk can provide a simple interface and minimize the input steps. If the creator is relaxed, the reception desk can also provide detailed input options and suggest customizable input methods. Furthermore, if the creator is in a hurry, the reception desk can prioritize voice input to allow for quick requirement entry. This allows for more appropriate input by adjusting the requirement input method according to the creator's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception desk may be performed using AI or not. For example, the reception desk can input the creator's emotion data into a generative AI and have the generative AI perform emotion estimation.
[0074] The reception desk can analyze a creator's past requirement input history and suggest the optimal input method. For example, the reception desk can automatically display requirements that the creator has frequently entered in the past as suggestions. The reception desk can also prioritize suggesting input methods (voice, text, etc.) that the creator has used in the past. Furthermore, the reception desk can predict and suggest requirements to be used during specific time periods based on the creator's past input history. In this way, by analyzing past requirement input history, the reception desk can suggest the optimal input method for the creator. Some or all of the above processing in the reception desk may be performed using AI, for example, or not using AI. For example, the reception desk can input the creator's past input history data into a generating AI and have the generating AI suggest the optimal input method.
[0075] The reception system can filter requirements based on the creator's current projects and areas of interest when they are entered. For example, the reception system will prioritize displaying requirements related to the creator's current projects. The reception system can also automatically suggest relevant requirements based on the creator's areas of interest. Furthermore, the reception system can filter requirements based on areas the creator has shown interest in in the past. This allows for the priority input of highly relevant requirements by filtering requirements based on the creator's current projects and areas of interest. Some or all of the above processing in the reception system may be performed using AI, for example, or not. For example, the reception system can input the creator's project data and area of interest data into a generating AI and have the generating AI perform the requirement filtering.
[0076] The reception desk can estimate the creator's emotions and determine the priority of the input requirements based on the estimated emotions. For example, if the creator is stressed, the reception desk will prioritize inputting important requirements. If the creator is relaxed, the reception desk may also prioritize inputting detailed requirements. Furthermore, if the creator is in a hurry, the reception desk may prioritize inputting requirements that can be processed quickly. This allows for the prioritization of important requirements by determining the priority of requirements according to the creator's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk can input the creator's emotion data into a generative AI and have the generative AI determine the priority of requirements.
[0077] The reception desk can prioritize highly relevant requirements when a creator enters their requirements, taking into account their geographical location. For example, if a creator is in a specific region, the reception desk will prioritize requirements related to that region. Furthermore, if a creator is on the move, the reception desk can suggest the most relevant requirements based on their current location. Additionally, if a creator is in a specific location, the reception desk can prioritize requirements related to that location. This allows for the prioritization of highly relevant requirements by considering the creator's geographical location. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk can input the creator's geographical location data into a generating AI and have the generating AI determine the priority of the requirements.
[0078] The reception desk can analyze the creator's social media activity and input relevant requirements when requirements are entered. For example, the reception desk can suggest relevant requirements based on information the creator has shared on social media. The reception desk can also identify areas of interest from the creator's social media activity and input requirements based on those areas. Furthermore, the reception desk can suggest relevant requirements based on the accounts the creator follows on social media. This allows for the input of relevant requirements by analyzing the creator's social media activity. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the creator's social media data into a generating AI and have the generating AI generate requirement suggestions.
[0079] The generation unit can estimate the creator's emotions and adjust the CG generation method based on the estimated emotions. For example, if the creator is relaxed, the generation unit can generate CG that progresses at a leisurely pace. If the creator is in a hurry, the generation unit can also generate CG that emphasizes the shortest route. Furthermore, if the creator is excited, the generation unit can generate CG with visually stimulating effects. In this way, by adjusting the CG generation method according to the creator's emotions, more appropriate CG can be generated. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or not using AI. For example, the generation unit can input the creator's emotion data into the generation AI and have the generation AI adjust the CG generation method.
[0080] The generation unit can select the optimal generation algorithm by referring to the creator's past works during generation. For example, the generation unit can select the optimal generation algorithm based on works created by the creator in the past. The generation unit can also extract a specific style from the creator's past works and select a generation algorithm based on that. Furthermore, the generation unit can analyze the creator's past works and select the most effective generation algorithm. In this way, the optimal generation algorithm can be selected by referring to the creator's past works. Some or all of the above processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the creator's past work data into a generation AI and have the generation AI perform the selection of a generation algorithm.
[0081] The generation unit can adjust the level of detail of the generated image based on the creator's current project. For example, the generation unit can adjust the level of detail based on the project the creator is currently working on. It can also adjust the level of detail based on the progress of the creator's project. Furthermore, the generation unit can adjust the level of detail based on the requirements of the creator's project. This allows for the generation of CG optimized for the project by adjusting the level of detail based on the creator's current project. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the creator's project data into a generation AI and have the generation AI perform the adjustment of the level of detail of the generated image.
[0082] The generation unit can estimate the creator's emotions and determine the priority of the CG to be generated based on the estimated emotions. For example, if the creator is stressed, the generation unit will prioritize generating important CG. It can also prioritize generating detailed CG if the creator is relaxed. Furthermore, if the creator is in a hurry, the generation unit can prioritize generating CG that can be produced quickly. This allows for the priority of important CG by determining the priority of CG according to the creator's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the generation unit may be performed using AI, or not. For example, the generation unit can input the creator's emotion data into the generation AI and have the generation AI determine the priority of the CG.
[0083] The generation unit can select the optimal generation method during generation, taking into account the creator's geographical location information. For example, if the creator is in a specific region, the generation unit will prioritize generating CG related to that region. Furthermore, if the creator is on the move, the generation unit can select the optimal generation method based on their current location. Additionally, if the creator is in a specific location, the generation unit can prioritize generating CG related to that location. This allows the generation unit to select the optimal generation method by considering the creator's geographical location information. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the creator's geographical location data into a generation AI and have the generation AI select the generation method.
[0084] The generation unit can analyze the creator's social media activity during generation and propose generation methods. For example, the generation unit can propose the optimal generation method based on information shared by the creator on social media. The generation unit can also identify areas of interest from the creator's social media activity and propose generation methods based on those areas. Furthermore, the generation unit can propose the optimal generation method based on the accounts the creator follows on social media. In this way, the optimal generation method can be proposed by analyzing the creator's social media activity. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the creator's social media data into a generation AI and have the generation AI execute the generation method proposal.
[0085] The editing unit can estimate the creator's emotions and adjust the editing method based on the estimated emotions. For example, if the creator is relaxed, the editing unit can provide detailed editing options. If the creator is in a hurry, the editing unit can also provide options for quick editing. Furthermore, if the creator is stressed, the editing unit can provide a simple editing interface. This allows for more appropriate editing by adjusting the editing method according to the creator's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the editing unit may be performed using AI or not using AI. For example, the editing unit can input the creator's emotion data into the generative AI and have the generative AI adjust the editing method.
[0086] The editing unit can select the optimal editing method by referring to the creator's past editing history during editing. For example, the editing unit can select the optimal editing method based on the editing the creator has done in the past. The editing unit can also extract a specific style from the creator's past editing history and select an editing method based on that. Furthermore, the editing unit can analyze the creator's past editing history and select the most effective editing method. In this way, the optimal editing method can be selected by referring to the creator's past editing history. Some or all of the above processes in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can input the creator's past editing history data into a generating AI and have the generating AI perform the selection of editing methods.
[0087] The editing unit can adjust the level of detail of the edits based on the creator's current project. For example, the editing unit can adjust the level of detail based on the project the creator is currently working on. It can also adjust the level of detail based on the progress of the creator's project. Furthermore, the editing unit can adjust the level of detail based on the requirements of the creator's project. This allows for optimal edits to be made based on the creator's current project. Some or all of the above processes in the editing unit may be performed using AI, for example, or not using AI. For example, the editing unit can input the creator's project data into a generating AI and have the generating AI perform the adjustment of the level of detail of the edits.
[0088] The editing unit can estimate the creator's emotions and determine the priority of CG to be edited based on the estimated emotions. For example, if the creator is stressed, the editing unit will prioritize editing important CG. If the creator is relaxed, the editing unit can also prioritize editing detailed CG. Furthermore, if the creator is in a hurry, the editing unit can prioritize editing CG that can be quickly edited. In this way, by determining the priority of CG according to the creator's emotions, important CG can be edited preferentially. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the editing unit may be performed using AI, or not using AI. For example, the editing unit can input the creator's emotion data into a generative AI and have the generative AI perform the determination of CG priorities.
[0089] The editing unit can select the optimal editing method by considering the creator's geographical location information during editing. For example, if the creator is in a specific region, the editing unit will prioritize editing CG related to that region. Furthermore, if the creator is on the move, the editing unit can select the optimal editing method based on their current location. Additionally, if the creator is in a specific location, the editing unit can prioritize editing CG related to that location. This allows the optimal editing method to be selected by considering the creator's geographical location information. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can input the creator's geographical location data into a generating AI and have the generating AI select the editing method.
[0090] The editing unit can analyze the creator's social media activity and propose editing methods during the editing process. For example, the editing unit can propose the most suitable editing method based on information shared by the creator on social media. The editing unit can also identify areas of interest from the creator's social media activity and propose editing methods based on those areas. Furthermore, the editing unit can propose the most suitable editing method based on the accounts the creator follows on social media. In this way, the editing unit can propose the most suitable editing method by analyzing the creator's social media activity. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can input the creator's social media data into a generating AI and have the generating AI execute the proposal of editing methods.
[0091] The customization unit can estimate the viewer's emotions and adjust the customization method based on the estimated viewer emotions. For example, if the viewer is relaxed, the customization unit can provide detailed customization options. If the viewer is in a hurry, the customization unit can also provide options for quick customization. Furthermore, if the viewer is excited, the customization unit can provide visually stimulating customization options. This allows for more appropriate customization by adjusting the customization method according to the viewer's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the customization unit may be performed using AI or not using AI. For example, the customization unit can input viewer emotion data into a generative AI and have the generative AI perform the adjustment of the customization method.
[0092] The customization unit can select the optimal customization method by referring to the viewer's past customization history during the customization process. For example, the customization unit can select the optimal customization method based on the viewer's past customizations. Furthermore, the customization unit can extract specific styles from the viewer's past customization history and select a customization method based on those styles. In addition, the customization unit can analyze the viewer's past customization history and select the most effective customization method. This allows the optimal customization method to be selected by referring to the viewer's past customization history. Some or all of the above-described processes in the customization unit may be performed using AI, for example, or without AI. For example, the customization unit can input the viewer's past customization history data into a generating AI and have the generating AI perform the selection of a customization method.
[0093] The customization unit can adjust the level of detail of customization based on the viewer's current areas of interest during the customization process. For example, the customization unit can adjust the level of detail of customization based on the areas the viewer is currently interested in. The customization unit can also adjust the level of detail of customization in response to changes in the viewer's areas of interest. Furthermore, the customization unit can adjust the level of detail of customization based on the viewer's current areas of interest. This allows for more appropriate customization by adjusting the level of detail of customization based on the viewer's current areas of interest. Some or all of the above-described processes in the customization unit may be performed using AI, for example, or without AI. For example, the customization unit can input viewer area of interest data into a generating AI and have the generating AI perform the adjustment of the level of detail of customization.
[0094] The customization unit can estimate the viewer's emotions and determine the priority of CG to customize based on the estimated viewer emotions. For example, if the viewer is relaxed, the customization unit will prioritize detailed CG. If the viewer is in a hurry, the customization unit can also prioritize CG that can be customized quickly. Furthermore, if the viewer is excited, the customization unit can prioritize visually stimulating CG. This allows for the prioritization of important CG by determining the priority of CG according to the viewer's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the customization unit may be performed using AI or not. For example, the customization unit can input viewer emotion data into a generative AI and have the generative AI determine the priority of CG.
[0095] The customization unit can select the optimal customization method by considering the viewer's geographical location information during the customization process. For example, if the viewer is in a specific region, the customization unit will prioritize customizing CG related to that region. Furthermore, if the viewer is on the move, the customization unit can also select the optimal customization method based on their current location. Additionally, if the viewer is in a specific location, the customization unit can prioritize customizing CG related to that location. This allows for the selection of the optimal customization method by considering the viewer's geographical location information. Some or all of the above-described processes in the customization unit may be performed using AI, for example, or without AI. For example, the customization unit can input the viewer's geographical location data into a generating AI and have the generating AI perform the selection of the customization method.
[0096] The customization unit can analyze the viewer's social media activity during the customization process and propose customization methods. For example, the customization unit can propose the optimal customization method based on information shared by the viewer on social media. Furthermore, the customization unit can identify areas of interest from the viewer's social media activity and propose customization methods based on those areas. In addition, the customization unit can propose the optimal customization method based on the accounts the viewer follows on social media. This allows for the proposal of the optimal customization method by analyzing the viewer's social media activity. Some or all of the above processing in the customization unit may be performed using AI, for example, or without AI. For example, the customization unit can input the viewer's social media data into a generating AI and have the generating AI execute the proposal of customization methods.
[0097] The learning unit can estimate the creator's emotions and select training data based on the estimated emotions. For example, if the creator is relaxed, the learning unit can select detailed training data. If the creator is in a hurry, the learning unit can also select data that allows for quick learning. Furthermore, if the creator is stressed, the learning unit can select simple training data. This allows for more appropriate learning by selecting training data according to the creator's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the learning unit may be performed using AI, or not using AI. For example, the learning unit can input the creator's emotion data into a generative AI and have the generative AI perform the selection of training data.
[0098] The learning unit can optimize the learning algorithm by referring to past learning data during the learning process. For example, the learning unit can select the optimal learning algorithm based on past learning data. The learning unit can also extract specific patterns from past learning data and optimize the learning algorithm based on those patterns. Furthermore, the learning unit can analyze past learning data and select the most effective learning algorithm. This allows for the optimization of the learning algorithm by referring to past learning data. Some or all of the above processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input past learning data into a generating AI and have the generating AI perform the optimization of the learning algorithm.
[0099] The learning unit can estimate the creator's emotions and adjust the learning frequency based on the estimated emotions. For example, if the creator is relaxed, the learning unit will learn more frequently. Conversely, if the creator is in a hurry, the learning unit can reduce the learning frequency. Furthermore, if the creator is stressed, the learning unit can adjust the learning frequency. This allows for more appropriate learning by adjusting the learning frequency according to the creator's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the learning unit may be performed using AI, for example, or not using AI. For example, the learning unit can input the creator's emotion data into the generative AI and have the generative AI adjust the learning frequency.
[0100] The learning unit can weight the training data during training based on the timing of creator requirements input. For example, the learning unit weights the training data based on the timing of requirements previously entered by the creator. The learning unit can also adjust the weighting of the training data according to the timing of creator requirements input. Furthermore, the learning unit can weight the training data based on the timing of creator requirements input. This allows for more appropriate training by weighting the training data based on the timing of creator requirements input. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input creator requirements input timing data into a generating AI and have the generating AI perform the weighting of the training data.
[0101] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0102] The video generation system can also be equipped with a feedback unit. The feedback unit collects feedback from creators and viewers and uses that feedback to improve the performance of the generating AI. For example, the feedback unit can analyze the modifications that creators have made to the generated CG and use that as training data for the generating AI. It can also collect feedback provided by viewers through interactive video experiences and improve the generating algorithm of the generating AI. Furthermore, the feedback unit can estimate the emotions of creators and viewers and evaluate the importance of the feedback based on those emotions. In this way, the feedback unit can effectively utilize feedback from creators and viewers and continuously improve the performance of the generating AI.
[0103] The video generation system can also be equipped with a prediction unit. This unit analyzes the creator's past requirements and revision history to predict the next required requirements and revisions. For example, the prediction unit can automatically suggest the next required requirements based on the requirements the creator has frequently used in the past. It can also analyze the creator's revision history to predict areas that will need revision. Furthermore, the prediction unit can estimate the creator's emotions and adjust the accuracy of its predictions based on those emotions. This allows the prediction unit to improve the creator's work efficiency.
[0104] The video generation system can also be equipped with an assistance unit. This unit provides real-time support as the creator inputs requirements. For example, it can suggest appropriate samples and reference materials in response to the requirements entered by the creator. Furthermore, it can automatically detect errors in the input and suggest corrections. In addition, the assistance unit can estimate the creator's emotions and adjust the content and method of support based on those emotions. This allows the assistance unit to efficiently support the creator's work.
[0105] The video generation system may also include an evaluation unit. This unit automatically evaluates the quality of the generated CG and provides the evaluation results to the creator. For example, the evaluation unit can analyze elements such as the color tone, shape, and movement of the generated CG and calculate a quality score. It can also estimate the creator's emotions and adjust the evaluation criteria based on those emotions. Furthermore, the evaluation unit can collect viewer feedback and update the evaluation results based on that feedback. This allows the evaluation unit to objectively evaluate the quality of the generated CG and provide feedback to the creator.
[0106] The video generation system can also be equipped with a guide unit. This guide unit provides appropriate guidelines and best practices when creators input requirements. For example, it can present past success stories and reference materials in response to the requirements entered by the creator. It can also estimate the creator's emotions and adjust the content and methods of the guidelines based on those emotions. Furthermore, the guide unit can provide real-time feedback as creators input requirements, supporting improvements to their input. This allows the guide unit to enable creators to input requirements effectively and maximize the performance of the generation AI.
[0107] The video generation system can also be equipped with a data integration unit. The data integration unit centrally manages the requirements and revisions entered by creators and provides them to the generation AI. For example, the data integration unit can integrate text data, image data, audio data, etc., entered by creators and input them into the generation AI. The data integration unit can also record the revisions made by creators and use them as training data for the generation AI. Furthermore, the data integration unit can analyze the creators' past requirements and revision history and provide feedback to improve the performance of the generation AI. In this way, the data integration unit can effectively manage the creators' requirements and revisions and maximize the performance of the generation AI.
[0108] The video generation system may also include a version control unit. The version control unit manages different versions of the generated CG, allowing creators to revert to previous versions as needed. For example, the version control unit automatically saves each version of the generated CG, making it easily accessible to creators. It can also record the revisions made by creators and track the revision history. Furthermore, the version control unit can suggest optimal revision methods based on the creator's past versions. This allows the version control unit to effectively manage different versions of the generated CG and revert to previous versions as needed.
[0109] The video generation system can also include a template section. This section provides predefined templates for creators to input requirements. For example, the template section can save frequently used requirements and settings as templates, which can then be easily recalled for future input. The template section can also analyze the creator's past requirement input history and suggest the most suitable template. Furthermore, the template section can allow creators to customize templates, supporting requirement input tailored to individual projects. This streamlines the creator's requirement input process and maximizes the performance of the generation AI.
[0110] The video generation system can also be equipped with a notification unit. This notification unit provides creators and viewers with real-time updates on the generation AI's progress and important events. For example, the notification unit can send a notification to the creator when the generation AI has completed CG generation. It can also notify viewers of important events or new options while they are enjoying an interactive video experience. Furthermore, the notification unit can estimate the emotions of creators and viewers and adjust the content and timing of notifications based on those emotions. This allows the notification unit to provide creators and viewers with important information at the right time, improving the user experience of the video generation system.
[0111] The video generation system can also be equipped with an analysis unit. The analysis unit comprehensively analyzes generated CG, creator requirements input history, viewer feedback, and other data to provide insights for improving the performance of the generation AI. For example, the analysis unit can analyze the quality of generated CG and viewer reactions to provide feedback for improving the generation AI's algorithm. It can also analyze creator requirements input history and suggest ways to streamline future requirements input. Furthermore, the analysis unit can collect viewer feedback and provide insights based on that feedback to improve the generation AI's performance. This allows the analysis unit to continuously improve the generation AI's performance and provide creators and viewers with a higher quality video experience.
[0112] The following briefly describes the processing flow for example form 2.
[0113] Step 1: The reception desk enters the creator's requirements. Creator requirements may include, but are not limited to, text format, image format, audio format, etc. Step 2: The generation unit uses a generation AI to automatically generate computer graphics (CG) based on the requirements entered by the reception unit. The generation AI generates the CG using a specific AI model and training method. Step 3: The revision section involves the creator reviewing the CG generated by the generation section and making revisions as needed. Revisions are made based on the scope of the revisions and the tools used. Step 4: The customization section allows viewers to customize the CG that has been modified by the modification section. Customization is based on the customizable elements and the interface to be used.
[0114] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0115] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0116] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0117] For example, the reception unit can input creator requirements via the reception device 38 of the smart device 14 or the communication I / F 26 of the data processing unit 12. The generation unit automatically generates CG using generation AI via the specific processing unit 290 of the data processing unit 12 or the control unit 46A of the smart device 14. The modification unit checks the generated CG on the display 40A of the smart device 14 or the specific processing unit 290 of the data processing unit 12 and makes modifications as needed. The customization unit allows viewers to customize the modified CG using the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing unit 12. The learning unit generates CG based on learning data such as image data, text data, and audio data via the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the devices and control units is not limited to the example described above, and various changes are possible.
[0118] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0119] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0120] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0121] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0122] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0123] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0124] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0125] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0126] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0127] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0128] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0129] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0130] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0131] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0132] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0133] For example, the reception unit can input creator requirements via the microphone 238 of the smart glasses 214 or the communication I / F 26 of the data processing unit 12. The generation unit automatically generates CG using generation AI via the specific processing unit 290 of the data processing unit 12 or the control unit 46A of the smart glasses 214. The modification unit checks the generated CG on the display of the smart glasses 214 or the specific processing unit 290 of the data processing unit 12 and makes modifications as needed. The customization unit allows viewers to customize the modified CG using the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing unit 12. The learning unit generates CG based on learning data such as image data, text data, and audio data via the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the devices and control units is not limited to the example described above and can be modified in various ways.
[0134] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0135] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0136] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0137] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0138] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0139] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0140] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0141] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0142] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0143] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0144] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0145] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0146] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0147] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0148] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0149] For example, the reception unit can input creator requirements via the microphone 238 of the headset terminal 314 or the communication I / F 26 of the data processing unit 12. The generation unit automatically generates CG using generation AI via the specific processing unit 290 of the data processing unit 12 or the control unit 46A of the headset terminal 314. The modification unit checks the generated CG on the display 343 of the headset terminal 314 or the specific processing unit 290 of the data processing unit 12 and makes modifications as needed. The customization unit allows viewers to customize the modified CG using the control unit 46A of the headset terminal 314 or the specific processing unit 290 of the data processing unit 12. The learning unit generates CG based on learning data such as image data, text data, and audio data via the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the devices and control units is not limited to the example described above, and various changes are possible.
[0150] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0151] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0152] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0153] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0154] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0155] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0156] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0157] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0158] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0159] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0160] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0161] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0162] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0163] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0164] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0165] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0166] For example, the reception unit can input creator requirements via the microphone 238 of the robot 414 or the communication I / F 26 of the data processing unit 12. The generation unit automatically generates CG using generation AI via the specific processing unit 290 of the data processing unit 12 or the control unit 46A of the robot 414. The modification unit checks the generated CG on the display of the robot 414 or the specific processing unit 290 of the data processing unit 12 and makes modifications as needed. The customization unit allows viewers to customize the modified CG using the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing unit 12. The learning unit generates CG based on learning data such as image data, text data, and audio data via the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the devices and control units is not limited to the example described above and can be changed in various ways.
[0167] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0168] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0169] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0170] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0171] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0172] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0173] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0174] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0175] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0176] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0177] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0178] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0179] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0180] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0181] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0182] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0183] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0184] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0185] (Note 1) A reception area where creator requirements are entered, A generation unit that automatically generates CG based on the requirements entered by the reception unit, A correction unit that checks and corrects the CG generated by the generation unit, The system includes a customization unit that allows viewers to customize the CG modified by the aforementioned modification unit. A system characterized by the following features. (Note 2) The generation AI has a learning unit that generates computer graphics based on training data. The system described in Appendix 1, characterized by the features described herein. (Note 3) The generating unit is Using generation AI, CG is automatically generated based on the creator's requirements. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned modification section is, The creator reviews the CG generated by the generation unit and makes corrections as needed. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned customization unit is Viewers can customize the CG that has been modified by the editing team. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned reception unit is It estimates the creator's emotions and adjusts the way requirements are entered based on the estimated emotions of the creator. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is We analyze the creator's past requirements input history and propose the optimal input method. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is When entering requirements, filtering is performed based on the creator's current projects and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is Estimate the creator's emotions and determine the priority of input requirements based on the estimated creator's emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is When entering requirements, the system prioritizes highly relevant requirements by considering the creator's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reception unit is When entering requirements, analyze the creator's social media activity and enter relevant requirements. The system described in Appendix 1, characterized by the features described herein. (Note 12) The generating unit is It estimates the creator's emotions and adjusts the CG generation method based on the estimated emotions of the creator. The system described in Appendix 1, characterized by the features described herein. (Note 13) The generating unit is During generation, the optimal generation algorithm is selected by referencing the creator's past works. The system described in Appendix 1, characterized by the features described herein. (Note 14) The generating unit is During generation, adjust the level of detail based on the creator's current project. The system described in Appendix 1, characterized by the features described herein. (Note 15) The generating unit is It estimates the creator's emotions and determines the priority of the CG to be generated based on the estimated emotions of the creator. The system described in Appendix 1, characterized by the features described herein. (Note 16) The generating unit is During generation, the optimal generation method is selected considering the creator's geographical location information. The system described in Appendix 1, characterized by the features described herein. (Note 17) The generating unit is During generation, we analyze the creator's social media activity and propose generation methods. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned modification section is, It estimates the creator's emotions and adjusts the method of correction based on the estimated emotions of the creator. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned modification section is, When making revisions, the optimal revision method is selected by referring to the creator's past revision history. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned modification section is, When making corrections, adjust the level of detail of the corrections based on the creator's current project. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned modification section is, It estimates the creator's emotions and determines the priority of CG modifications based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned modification section is, When making revisions, the creator's geographical location information will be taken into consideration to select the most suitable revision method. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned modification section is, During the revision process, we analyze the creator's social media activity and propose revision methods. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned customization unit is It estimates the viewer's emotions and adjusts the customization method based on the estimated viewer emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned customization unit is During customization, the system will refer to the viewer's past customization history to select the most suitable customization method. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned customization unit is When customizing, adjust the level of detail based on the viewer's current areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned customization unit is It estimates the viewer's emotions and determines the priority of CG to customize based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned customization unit is During customization, the optimal customization method is selected by considering the viewer's geographical location information. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned customization unit is During customization, we analyze the audience's social media activity and suggest customization options. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned learning unit, The system estimates the emotions of creators and selects training data based on the estimated emotions of the creators. The system described in Appendix 2, characterized by the features described herein. (Note 31) The aforementioned learning unit, During training, the learning algorithm is optimized by referring to past training data. The system described in Appendix 2, characterized by the features described herein. (Note 32) The aforementioned learning unit, It estimates the creator's emotions and adjusts the learning frequency based on the estimated emotions of the creator. The system described in Appendix 2, characterized by the features described herein. (Note 33) The aforementioned learning unit, During training, the training data is weighted based on when the creators entered their requirements. The system described in Appendix 2, characterized by the features described herein. [Explanation of Symbols]
[0186] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A reception area where creator requirements are entered, A generation unit that automatically generates CG based on the requirements entered by the reception unit, A correction unit that checks and corrects the CG generated by the generation unit, The system includes a customization unit that allows viewers to customize the CG modified by the aforementioned modification unit. A system characterized by the following features.
2. The system includes a learning unit that generates computer graphics (CG) based on training data using a generative AI. The system according to feature 1.
3. The generating unit is Using generation AI, CG is automatically generated based on the creator's requirements. The system according to feature 1.
4. The aforementioned reception unit is It estimates the creator's emotions and adjusts the way requirements are entered based on the estimated emotions of the creator. The system according to feature 1.
5. The aforementioned reception unit is We analyze the creator's past requirements input history and propose the optimal input method. The system according to feature 1.
6. The aforementioned reception unit is When entering requirements, filtering is performed based on the creator's current projects and areas of interest. The system according to feature 1.
7. The aforementioned reception unit is Estimate the creator's emotions and determine the priority of input requirements based on the estimated creator's emotions. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A