system

The system efficiently generates presentation materials through AI-driven text, layout, and video creation, addressing the inefficiencies of manual creation by automating the process and enhancing work efficiency.

JP2026045860APending Publication Date: 2026-03-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Creating presentation materials is laborious and time-consuming, making it difficult to perform efficiently.

Method used

A system comprising a reception unit, text generation unit, layout generation unit, script generation unit, and video generation unit, which automatically generates presentation materials using AI, including text, layout, and video creation based on user input, and suggests optimal templates and layouts based on user history and preferences.

Benefits of technology

Significantly reduces the time required to create presentation materials and enhances work efficiency by automating the process, providing high-quality materials and improving presentation skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026045860000001_ABST
    Figure 2026045860000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to efficiently create presentation materials. [Solution] The system according to the embodiment comprises a reception unit, a text generation unit, a layout generation unit, a script generation unit, and a video generation unit. The reception unit selects the type of material and inputs the necessary information or images to be used. The text generation unit analyzes the information input by the reception unit and generates the text portion of the presentation material. The layout generation unit creates the layout of the material based on the text generated by the text generation unit. The script generation unit generates a talk script related to the material generated by the layout generation unit. The video generation unit generates a sample video based on the talk script generated by the script generation unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there was a problem that creating presentation materials was laborious and time-consuming and difficult to perform efficiently.

[0005] The system according to the embodiment aims to efficiently create presentation materials.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a reception unit, a text generation unit, a layout generation unit, a script generation unit, and a video generation unit. The reception unit selects the type of material and inputs the necessary information or images to be used. The text generation unit analyzes the information input by the reception unit and generates the text portion of the presentation material. The layout generation unit creates the layout of the material based on the text generated by the text generation unit. The script generation unit generates a talk script related to the material generated by the layout generation unit. The video generation unit generates a sample video based on the talk script generated by the script generation unit. [Effects of the Invention]

[0007] The system according to this embodiment can efficiently create presentation materials. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The presentation material automatic generation system according to an embodiment of the present invention is a system that automatically generates presentation materials using generation AI and also provides talk scripts and example videos. In this system, the user selects the type of material (presentation material, proposal material, educational material, etc.) and inputs the necessary information and images to be used, and the text generation AI and image generation AI automatically generate the presentation material based on the input information. Furthermore, the text generation AI automatically generates a talk script related to the generated material, and the video generation AI creates an example video appropriate to the situation. This mechanism can significantly reduce the time required to create materials and improve work efficiency. For example, the user selects the type of material and inputs the necessary information and images to be used. For example, the user selects presentation material and inputs information such as product features and market data. This information is input to the text generation AI and image generation AI. Next, the text generation AI and image generation AI automatically generate the presentation material based on the input information. The text generation AI refines the expression of the input text, and the image generation AI creates the layout of the material. For example, slides explaining product features and slides showing market data in graphs are automatically generated. The text generation AI automatically generates a talk script related to the generated material. For example, a talk script is automatically generated for slides explaining product features. Furthermore, a video generation AI creates example videos appropriate to the situation. For instance, an example video is automatically generated showing how to speak based on presentation materials. This allows users to learn appropriate speaking styles for different situations. This system integrates document creation and talk script creation, significantly reducing the time spent on document creation. In addition, the provision of situation-appropriate videos allows users to quickly acquire presentation skills. For example, when a junior sales representative creates presentation materials, they can use this system to create high-quality materials in a short time and prepare for their presentation by referring to the talk script and example videos. This is expected to improve work efficiency and enhance presentation skills.This allows the automated presentation material generation system to significantly reduce the time required for creating materials and improve work efficiency.

[0029] The automated presentation material generation system according to this embodiment comprises a reception unit, a text generation unit, a layout generation unit, a script generation unit, and a video generation unit. The reception unit selects the type of material and inputs necessary information and images to be used. For example, a user selects a presentation material and inputs information such as product features and market data. This information is input to the text generation AI and the image generation AI. The text generation unit analyzes the input information and generates the text portion of the presentation material. For example, the text generation unit generates text that explains the product features. The text generation unit can also generate text that explains market data. Furthermore, the text generation unit can use the generation AI to refine the expression of the text. The layout generation unit creates the layout of the material based on the text generated by the text generation unit. For example, the layout generation unit creates the layout of a slide that explains the product features. The layout generation unit can also create the layout of a slide that shows market data in graph form. Furthermore, the layout generation unit can use the image generation AI to automatically generate the layout of the material. The script generation unit generates a talk script related to the material generated by the layout generation unit. For example, the script generation unit generates a talk script for slides describing product features, outlining how to explain them. It can also generate a talk script for slides describing market data, outlining how to explain them. Furthermore, the script generation unit can automatically generate talk scripts using a generation AI. The video generation unit generates example videos based on the talk scripts generated by the script generation unit. For example, the video generation unit generates example videos for presentation materials, outlining how to speak. It can also generate example videos appropriate to different situations. Furthermore, the video generation unit can automatically generate videos using a generation AI. As a result, the automated presentation material generation system according to this embodiment can significantly reduce the time required for material creation and improve operational efficiency.

[0030] The reception desk can analyze the user's past document creation history and suggest appropriate information input methods. For example, the reception desk can automatically suggest information input methods that the user has frequently used in the past. The reception desk can also suggest the optimal template based on the user's past document creation history. Furthermore, the reception desk can analyze the user's past document creation history and suggest the most efficient information input method. This improves the user's work efficiency by suggesting the optimal information input method based on past history. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input the user's past document creation history data into a generating AI and have the generating AI suggest the optimal information input method.

[0031] The reception desk can filter the selection of document types based on the user's current projects and areas of interest. For example, the reception desk can prioritize presenting document types relevant to the user's current projects. It can also suggest highly relevant document types based on the user's areas of interest. Furthermore, the reception desk can suggest the most suitable document types according to the progress of the user's projects. This supports efficient document creation by suggesting appropriate document types according to the user's projects and areas of interest. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's project data into a generating AI and have the generating AI suggest the most suitable document types.

[0032] The reception desk can prioritize and present highly relevant options when the user selects a type of document, taking into account the user's geographical location. For example, the reception desk can suggest relevant document types based on the user's current location. It can also present the most suitable document type considering the user's geographical location. Furthermore, the reception desk can suggest region-specific document types based on the user's location. This supports efficient document creation by presenting appropriate options based on the user's geographical location. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input the user's location data into a generating AI and have the generating AI suggest the most suitable document type.

[0033] The reception desk can analyze the user's social media activity and present relevant options when selecting the type of material. For example, the reception desk can suggest types of materials of interest based on the user's social media activity. It can also suggest the most suitable type of material based on the user's social media activity. Furthermore, the reception desk can suggest relevant types of materials by referring to the activities of the user's social media followers and friends. This supports efficient material creation by presenting appropriate options based on the user's social media activity. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's social media data into a generating AI and have the generating AI suggest the most suitable type of material.

[0034] The text generation unit can adjust the level of detail in the text based on the importance of the material during text generation. For example, in the case of important material, the text generation unit generates text that includes detailed explanations. In addition, in the case of general material, the text generation unit can generate concise text. Furthermore, in the case of urgent material, the text generation unit can generate text that gets straight to the point. By adjusting the level of detail in the text according to the importance of the material, it is possible to generate presentation materials with an appropriate amount of information. Some or all of the above processing in the text generation unit may be performed using AI, for example, or without AI. For example, the text generation unit can input material importance data into a generation AI and have the generation AI perform the adjustment of the level of detail in the text.

[0035] The text generation unit can apply different generation algorithms depending on the category of the document during text generation. For example, in the case of presentation materials, the text generation unit can apply an algorithm that generates persuasive text. Furthermore, in the case of proposal materials, the text generation unit can apply an algorithm that generates text containing specific proposal content. In addition, in the case of educational materials, the text generation unit can apply an algorithm that generates text containing easy-to-understand explanations. By applying a generation algorithm appropriate to the category of the document, more appropriate presentation materials can be generated. Some or all of the above-described processes in the text generation unit may be performed using AI, for example, or without AI. For example, the text generation unit can input the document category data into a generation AI and have the generation AI execute the application of the generation algorithm.

[0036] The text generation unit can determine the priority of text based on the submission deadline of the document during text generation. For example, in the case of an urgent document, the text generation unit will prioritize generating the most important information. Furthermore, for documents with an approaching submission deadline, the text generation unit can generate concise text. In addition, for documents with a distant submission deadline, the text generation unit can generate text that includes detailed explanations. This allows for the efficient generation of presentation materials by determining the text priority according to the submission deadline. Some or all of the above processing in the text generation unit may be performed using AI, for example, or without AI. For example, the text generation unit can input document submission deadline data into a generation AI and have the generation AI determine the text priority.

[0037] The text generation unit can adjust the order of text based on the relevance of the materials during text generation. For example, the text generation unit may place the main points of the materials first. It can also prioritize the placement of information that is highly relevant to the materials. Furthermore, the text generation unit can adjust the order of text based on the importance of the materials. By adjusting the order of text according to the relevance of the materials, more effective presentation materials can be generated. Some or all of the above processing in the text generation unit may be performed using AI, for example, or without AI. For example, the text generation unit can input the relevance data of the materials into a generation AI and have the generation AI perform the adjustment of the text order.

[0038] The layout generation unit can adjust the level of detail of the layout based on the content of the document during layout generation. For example, the layout generation unit can provide a detailed layout for important documents. It can also provide a concise layout for general documents. Furthermore, it can provide a layout that captures the essentials for urgent documents. By adjusting the level of detail of the layout according to the content of the document, more effective presentation materials can be generated. Some or all of the above processing in the layout generation unit may be performed using AI, for example, or without AI. For example, the layout generation unit can input the content data of the document into a generation AI and have the generation AI perform the adjustment of the level of detail of the layout.

[0039] The layout generation unit can apply different layout algorithms depending on the category of the document during layout generation. For example, in the case of presentation materials, the layout generation unit can apply an algorithm that provides a visually appealing layout. It can also apply an algorithm that provides a layout that emphasizes the specific content of the proposal in the case of proposal materials. Furthermore, in the case of educational materials, the layout generation unit can apply an algorithm that provides a layout that includes easy-to-understand explanations. By applying a layout algorithm appropriate to the category of the document, more effective presentation materials can be generated. Some or all of the above-described processes in the layout generation unit may be performed using AI, for example, or without AI. For example, the layout generation unit can input the category data of the document into a generation AI and have the generation AI execute the application of the layout algorithm.

[0040] The layout generation unit can determine layout priorities based on the submission deadline of the document during layout generation. For example, in the case of an urgent document, the layout generation unit provides a layout that prioritizes the placement of the most important information. The layout generation unit can also provide a layout that focuses on the essentials for documents with an approaching submission deadline. Furthermore, for documents with a distant submission deadline, the layout generation unit can provide a layout that includes detailed explanations. By determining the layout priorities according to the submission deadline of the document, efficient presentation materials can be generated. Some or all of the above processing in the layout generation unit may be performed using AI, for example, or without AI. For example, the layout generation unit can input document submission deadline data into a generation AI and have the generation AI perform the determination of layout priorities.

[0041] The layout generation unit can adjust the order of the layout based on the relevance of the materials during layout generation. For example, the layout generation unit can provide a layout that places the main points of the materials first. It can also provide a layout that prioritizes the placement of highly relevant information in the materials. Furthermore, the layout generation unit can adjust the order of the layout based on the importance of the materials. By adjusting the order of the layout according to the relevance of the materials, more effective presentation materials can be generated. Some or all of the above processing in the layout generation unit may be performed using AI, for example, or without AI. For example, the layout generation unit can input the relevance data of the materials into a generation AI and have the generation AI perform the adjustment of the layout order.

[0042] The script generation unit can adjust the level of detail in the script based on the importance of the document during script generation. For example, for important documents, the script generation unit generates a script that includes detailed explanations. For general documents, the script generation unit can also generate a concise script. Furthermore, for urgent documents, the script generation unit can generate a script that gets straight to the point. By adjusting the level of detail in the script according to the importance of the document, more effective presentation materials can be generated. Some or all of the above processing in the script generation unit may be performed using AI, for example, or without AI. For example, the script generation unit can input document importance data into a generation AI and have the generation AI perform the adjustment of the level of detail in the script.

[0043] The script generation unit can apply different generation algorithms depending on the category of the document when generating a script. For example, in the case of presentation materials, the script generation unit can apply an algorithm that generates a persuasive script. In the case of proposal materials, the script generation unit can also apply an algorithm that generates a script that includes specific proposal content. Furthermore, in the case of educational materials, the script generation unit can apply an algorithm that generates a script that includes easy-to-understand explanations. By applying a generation algorithm according to the category of the document, more effective presentation materials can be generated. Some or all of the above-described processes in the script generation unit may be performed using AI, for example, or without AI. For example, the script generation unit can input the category data of the document into a generation AI and have the generation AI execute the application of the generation algorithm.

[0044] The script generation unit can determine the priority of scripts based on the submission deadline of the materials when generating scripts. For example, in the case of urgent materials, the script generation unit provides a script that prioritizes the generation of the most important information. The script generation unit can also provide a script that gets to the point for materials with an approaching submission deadline. Furthermore, the script generation unit can provide a script that includes detailed explanations for materials with a distant submission deadline. By determining the priority of scripts according to the submission deadline of the materials, efficient presentation materials can be generated. Some or all of the above processing in the script generation unit may be performed using AI, for example, or without AI. For example, the script generation unit can input the submission deadline data of the materials into a generation AI and have the generation AI perform the determination of script priority.

[0045] The script generation unit can adjust the order of scripts based on the relevance of the materials during script generation. For example, the script generation unit can provide a script that places the main points of the materials first. It can also provide a script that prioritizes the placement of highly relevant information in the materials. Furthermore, the script generation unit can adjust the order of scripts based on the importance of the materials. By adjusting the order of scripts according to the relevance of the materials, more effective presentation materials can be generated. Some or all of the above processing in the script generation unit may be performed using AI, for example, or without AI. For example, the script generation unit can input the relevance data of the materials into a generating AI and have the generating AI perform the adjustment of the script order.

[0046] The video generation unit can adjust the level of detail in a video based on the importance of the material during video generation. For example, if the material is important, the video generation unit will generate a video with detailed explanations. It can also generate a concise video for general materials. Furthermore, for urgent materials, the video generation unit can generate a video that summarizes the key points. By adjusting the level of detail in the video according to the importance of the material, more effective presentation materials can be generated. Some or all of the above processing in the video generation unit may be performed using AI, for example, or without AI. For example, the video generation unit can input material importance data into a generation AI and have the generation AI adjust the level of detail in the video.

[0047] The video generation unit can apply different generation algorithms depending on the category of the material when generating videos. For example, in the case of presentation materials, the video generation unit can apply an algorithm that generates a visually appealing video. In the case of proposal materials, the video generation unit can also apply an algorithm that generates a video that emphasizes the specific content of the proposal. Furthermore, in the case of educational materials, the video generation unit can apply an algorithm that generates a video that includes easy-to-understand explanations. By applying a generation algorithm according to the category of the material, more effective presentation materials can be generated. Some or all of the above processing in the video generation unit may be performed using AI, for example, or without AI. For example, the video generation unit can input the category data of the material into a generation AI and have the generation AI execute the application of the generation algorithm.

[0048] The video generation unit can prioritize videos based on the submission deadline of the materials during video generation. For example, in the case of urgent materials, the video generation unit will provide a video that prioritizes the most important information. It can also provide a concise video for materials with an approaching submission deadline. Furthermore, for materials with a distant submission deadline, the video generation unit can provide a video with detailed explanations. This allows for the efficient generation of presentation materials by prioritizing videos according to the submission deadline of the materials. Some or all of the above-described processes in the video generation unit may be performed using AI, for example, or without AI. For example, the video generation unit can input the submission deadline data of the materials into a generation AI and have the generation AI determine the video prioritization.

[0049] The video generation unit can adjust the order of videos based on the relevance of the materials during video generation. For example, the video generation unit can provide a video that places the main points of the materials first. It can also provide a video that prioritizes the placement of highly relevant information from the materials. Furthermore, the video generation unit can adjust the order of videos based on the importance of the materials. By adjusting the order of videos according to the relevance of the materials, more effective presentation materials can be generated. Some or all of the above processing in the video generation unit may be performed using AI, for example, or without AI. For example, the video generation unit can input the relevance data of the materials into a generation AI and have the generation AI perform the adjustment of the video order.

[0050] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0051] The reception desk collects evaluation data on users' past presentation materials, analyzes the characteristics of high-rated materials, and incorporates those characteristics into future presentations. For example, it can refer to the layout and text style of previously highly-rated materials. It can also suggest areas for improvement to users based on the evaluation data. Furthermore, it can provide customized advice to help users improve their presentation skills based on the evaluation data. This allows users to leverage past successes and create more effective presentation materials.

[0052] The reception desk analyzes the user's past document creation history and not only suggests the optimal information input method, but also learns the user's work patterns and can automatically select the most suitable template and layout for the next document creation. For example, it can prioritize presenting templates that the user frequently uses. It can also suggest an optimal work schedule based on the user's working hours and frequency. Furthermore, it can automatically set efficient work procedures based on the user's past work patterns. As a result, users can leverage their past work history and create documents more efficiently.

[0053] The reception desk can not only filter the user's current projects and areas of interest when selecting document types, but can also suggest highly relevant document types by considering the user's industry trends and the latest research findings. For example, it can present the most suitable document types based on the latest trends in the user's industry. It can also suggest document types that reflect the latest research findings related to the user's areas of interest. Furthermore, it can suggest the most suitable document types according to the progress of the user's project. This allows users to utilize the latest information and create more effective documents.

[0054] The reception desk not only prioritizes and presents highly relevant options based on the user's geographical location when selecting document types, but can also suggest region-specific designs and content based on the user's location. For example, if a user is in a specific region, it can suggest designs that match the local culture and customs. It can also provide content that includes region-specific data and case studies. Furthermore, it can suggest document types related to local events and trends based on the user's location. This allows users to utilize region-specific information and create more effective documents.

[0055] The reception desk can analyze the user's social media activity when selecting a type of document, not only presenting relevant options but also suggesting the most suitable type of document based on the user's social media influence and the interests of their followers. For example, if a user has many followers, it can suggest a type of document that matches the interests of those followers. It can also suggest highly relevant types of documents based on the user's social media activity. Furthermore, it can suggest the most suitable type of document considering the user's social media influence. This allows users to leverage their social media influence and create more effective documents.

[0056] The text generation unit can adjust the level of detail of the text based on the importance of the material, as well as the difficulty level of the text according to the expertise level of the recipient. For example, it can provide detailed and specialized text to recipients with high expertise, and concise and easy-to-understand text to general audiences. Furthermore, it can provide text containing basic concepts for beginners. This allows for the generation of more effective presentation materials by providing appropriate text tailored to the audience.

[0057] The text generation unit not only applies different generation algorithms depending on the category of the document during text generation, but can also adjust the tone and style of the text according to the purpose of the document. For example, for persuasive presentation materials, an emphasized tone and persuasive style are used. For educational materials, a clear and approachable tone and style can be used. Furthermore, for proposal materials, a specific and practical tone and style can be used. This allows for the generation of more effective presentation materials by providing text appropriate to the purpose of the document.

[0058] The following briefly describes the processing flow for example form 1.

[0059] Step 1: The reception desk selects the type of document and inputs the necessary information and images to be used. For example, a user might select a presentation document and input information such as product features and market data. This information is then input into the text generation AI and image generation AI. Step 2: The text generation unit analyzes the input information and generates the text portion of the presentation materials. For example, the text generation unit generates text describing product features or text describing market data. It can also use generation AI to refine the expression of the text. Step 3: The layout generation unit creates the layout of the document based on the text generated by the text generation unit. For example, it creates layouts for slides explaining product features or slides showing market data in graphs. It can also automatically generate document layouts using image generation AI. Step 4: The script generation unit generates talk scripts related to the materials generated by the layout generation unit. For example, it generates talk scripts on how to explain slides describing product features or slides describing market data. It can also automatically generate talk scripts using generation AI. Step 5: The video generation unit generates example videos based on the talk scripts generated by the script generation unit. For example, it can generate example videos showing how to speak based on presentation materials or example videos appropriate for different situations. It can also automatically generate videos using generation AI.

[0060] (Example of form 2) The presentation material automatic generation system according to an embodiment of the present invention is a system that automatically generates presentation materials using generation AI and also provides talk scripts and example videos. In this system, the user selects the type of material (presentation material, proposal material, educational material, etc.) and inputs the necessary information and images to be used, and the text generation AI and image generation AI automatically generate the presentation material based on the input information. Furthermore, the text generation AI automatically generates a talk script related to the generated material, and the video generation AI creates an example video appropriate to the situation. This mechanism can significantly reduce the time required to create materials and improve work efficiency. For example, the user selects the type of material and inputs the necessary information and images to be used. For example, the user selects presentation material and inputs information such as product features and market data. This information is input to the text generation AI and image generation AI. Next, the text generation AI and image generation AI automatically generate the presentation material based on the input information. The text generation AI refines the expression of the input text, and the image generation AI creates the layout of the material. For example, slides explaining product features and slides showing market data in graphs are automatically generated. The text generation AI automatically generates a talk script related to the generated material. For example, a talk script is automatically generated for slides explaining product features. Furthermore, a video generation AI creates example videos appropriate to the situation. For instance, an example video is automatically generated showing how to speak based on presentation materials. This allows users to learn appropriate speaking styles for different situations. This system integrates document creation and talk script creation, significantly reducing the time spent on document creation. In addition, the provision of situation-appropriate videos allows users to quickly acquire presentation skills. For example, when a junior sales representative creates presentation materials, they can use this system to create high-quality materials in a short time and prepare for their presentation by referring to the talk script and example videos. This is expected to improve work efficiency and enhance presentation skills.This allows the automated presentation material generation system to significantly reduce the time required for creating materials and improve work efficiency.

[0061] The automated presentation material generation system according to this embodiment comprises a reception unit, a text generation unit, a layout generation unit, a script generation unit, and a video generation unit. The reception unit selects the type of material and inputs necessary information and images to be used. For example, a user selects a presentation material and inputs information such as product features and market data. This information is input to the text generation AI and the image generation AI. The text generation unit analyzes the input information and generates the text portion of the presentation material. For example, the text generation unit generates text that explains the product features. The text generation unit can also generate text that explains market data. Furthermore, the text generation unit can use the generation AI to refine the expression of the text. The layout generation unit creates the layout of the material based on the text generated by the text generation unit. For example, the layout generation unit creates the layout of a slide that explains the product features. The layout generation unit can also create the layout of a slide that shows market data in graph form. Furthermore, the layout generation unit can use the image generation AI to automatically generate the layout of the material. The script generation unit generates a talk script related to the material generated by the layout generation unit. For example, the script generation unit generates a talk script for slides describing product features, outlining how to explain them. It can also generate a talk script for slides describing market data, outlining how to explain them. Furthermore, the script generation unit can automatically generate talk scripts using a generation AI. The video generation unit generates example videos based on the talk scripts generated by the script generation unit. For example, the video generation unit generates example videos for presentation materials, outlining how to speak. It can also generate example videos appropriate to different situations. Furthermore, the video generation unit can automatically generate videos using a generation AI. As a result, the automated presentation material generation system according to this embodiment can significantly reduce the time required for material creation and improve operational efficiency.

[0062] The reception desk can estimate the user's emotions and present a selection of document types based on the estimated emotions. For example, if the user is stressed, the reception desk may present simple options and minimize the number of choices. If the user is relaxed, the reception desk may also present detailed options and offer customizable options. Furthermore, if the user is in a hurry, the reception desk may prioritize presenting the most frequently used options. This reduces user stress and supports efficient document creation by presenting appropriate options according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception desk may be performed using AI or not. For example, the reception desk may input user facial expression data into a generative AI and have the generative AI perform the user's emotion estimation.

[0063] The reception desk can analyze the user's past document creation history and suggest appropriate information input methods. For example, the reception desk can automatically suggest information input methods that the user has frequently used in the past. The reception desk can also suggest the optimal template based on the user's past document creation history. Furthermore, the reception desk can analyze the user's past document creation history and suggest the most efficient information input method. This improves the user's work efficiency by suggesting the optimal information input method based on past history. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input the user's past document creation history data into a generating AI and have the generating AI suggest the optimal information input method.

[0064] The reception desk can filter the selection of document types based on the user's current projects and areas of interest. For example, the reception desk can prioritize presenting document types relevant to the user's current projects. It can also suggest highly relevant document types based on the user's areas of interest. Furthermore, the reception desk can suggest the most suitable document types according to the progress of the user's projects. This supports efficient document creation by suggesting appropriate document types according to the user's projects and areas of interest. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's project data into a generating AI and have the generating AI suggest the most suitable document types.

[0065] The reception desk can estimate the user's emotions and prioritize the information to be entered based on the estimated emotions. For example, if the user is stressed, the reception desk will prioritize the input of the most important information. If the user is relaxed, the reception desk may also request detailed information. Furthermore, if the user is in a hurry, the reception desk may allow only minimal information to be entered. This supports efficient information input by prioritizing information input according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception desk may be performed using AI or not using AI. For example, the reception desk can input the user's facial expression data into the generative AI and have the generative AI perform the estimation of the user's emotions.

[0066] The reception desk can prioritize and present highly relevant options when the user selects a type of document, taking into account the user's geographical location. For example, the reception desk can suggest relevant document types based on the user's current location. It can also present the most suitable document type considering the user's geographical location. Furthermore, the reception desk can suggest region-specific document types based on the user's location. This supports efficient document creation by presenting appropriate options based on the user's geographical location. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input the user's location data into a generating AI and have the generating AI suggest the most suitable document type.

[0067] The reception desk can analyze the user's social media activity and present relevant options when selecting the type of material. For example, the reception desk can suggest types of materials of interest based on the user's social media activity. It can also suggest the most suitable type of material based on the user's social media activity. Furthermore, the reception desk can suggest relevant types of materials by referring to the activities of the user's social media followers and friends. This supports efficient material creation by presenting appropriate options based on the user's social media activity. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's social media data into a generating AI and have the generating AI suggest the most suitable type of material.

[0068] The text generation unit can estimate the user's emotions and adjust the text's expression based on the estimated emotions. For example, if the user is relaxed, the text generation unit will use softer language. If the user is in a hurry, the text generation unit can also use concise and to-the-point language. Furthermore, if the user is excited, the text generation unit can use visually stimulating language. By adjusting the text's expression according to the user's emotions, more effective presentation materials can be generated. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generative AI. Some or all of the above-described processes in the text generation unit may be performed using AI, or not using AI. For example, the text generation unit can input user facial expression data into the generative AI and have the generative AI perform the user's emotion estimation.

[0069] The text generation unit can adjust the level of detail in the text based on the importance of the material during text generation. For example, in the case of important material, the text generation unit generates text that includes detailed explanations. In addition, in the case of general material, the text generation unit can generate concise text. Furthermore, in the case of urgent material, the text generation unit can generate text that gets straight to the point. By adjusting the level of detail in the text according to the importance of the material, it is possible to generate presentation materials with an appropriate amount of information. Some or all of the above processing in the text generation unit may be performed using AI, for example, or without AI. For example, the text generation unit can input material importance data into a generation AI and have the generation AI perform the adjustment of the level of detail in the text.

[0070] The text generation unit can apply different generation algorithms depending on the category of the document during text generation. For example, in the case of presentation materials, the text generation unit can apply an algorithm that generates persuasive text. Furthermore, in the case of proposal materials, the text generation unit can apply an algorithm that generates text containing specific proposal content. In addition, in the case of educational materials, the text generation unit can apply an algorithm that generates text containing easy-to-understand explanations. By applying a generation algorithm appropriate to the category of the document, more appropriate presentation materials can be generated. Some or all of the above-described processes in the text generation unit may be performed using AI, for example, or without AI. For example, the text generation unit can input the document category data into a generation AI and have the generation AI execute the application of the generation algorithm.

[0071] The text generation unit can estimate the user's emotions and adjust the length of the text based on the estimated emotions. For example, if the user is in a hurry, the text generation unit can generate short, concise text. If the user is relaxed, the text generation unit can also generate longer text with detailed explanations. Furthermore, if the user is excited, the text generation unit can generate text with visually stimulating expressions. By adjusting the length of the text according to the user's emotions, more effective presentation materials can be generated. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generative AI, but is not limited to these examples. Some or all of the above processing in the text generation unit may be performed using AI, or not using AI. For example, the text generation unit can input user facial expression data into the generative AI and have the generative AI perform the estimation of the user's emotions.

[0072] The text generation unit can determine the priority of text based on the submission deadline of the document during text generation. For example, in the case of an urgent document, the text generation unit will prioritize generating the most important information. Furthermore, for documents with an approaching submission deadline, the text generation unit can generate concise text. In addition, for documents with a distant submission deadline, the text generation unit can generate text that includes detailed explanations. This allows for the efficient generation of presentation materials by determining the text priority according to the submission deadline. Some or all of the above processing in the text generation unit may be performed using AI, for example, or without AI. For example, the text generation unit can input document submission deadline data into a generation AI and have the generation AI determine the text priority.

[0073] The text generation unit can adjust the order of text based on the relevance of the materials during text generation. For example, the text generation unit may place the main points of the materials first. It can also prioritize the placement of information that is highly relevant to the materials. Furthermore, the text generation unit can adjust the order of text based on the importance of the materials. By adjusting the order of text according to the relevance of the materials, more effective presentation materials can be generated. Some or all of the above processing in the text generation unit may be performed using AI, for example, or without AI. For example, the text generation unit can input the relevance data of the materials into a generation AI and have the generation AI perform the adjustment of the text order.

[0074] The layout generation unit can estimate the user's emotions and adjust the layout design based on the estimated emotions. For example, if the user is nervous, the layout generation unit can provide a layout with calm colors. It can also provide a layout with bright colors if the user is enjoying themselves. Furthermore, if the user is tired, the layout generation unit can provide a simple and highly visible layout. By adjusting the layout design according to the user's emotions, more effective presentation materials can be generated. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the layout generation unit may be performed using AI, or not. For example, the layout generation unit can input user facial expression data into the generative AI and have the generative AI perform the user's emotion estimation.

[0075] The layout generation unit can adjust the level of detail of the layout based on the content of the document during layout generation. For example, the layout generation unit can provide a detailed layout for important documents. It can also provide a concise layout for general documents. Furthermore, it can provide a layout that captures the essentials for urgent documents. By adjusting the level of detail of the layout according to the content of the document, more effective presentation materials can be generated. Some or all of the above processing in the layout generation unit may be performed using AI, for example, or without AI. For example, the layout generation unit can input the content data of the document into a generation AI and have the generation AI perform the adjustment of the level of detail of the layout.

[0076] The layout generation unit can apply different layout algorithms depending on the category of the document during layout generation. For example, in the case of presentation materials, the layout generation unit can apply an algorithm that provides a visually appealing layout. It can also apply an algorithm that provides a layout that emphasizes the specific content of the proposal in the case of proposal materials. Furthermore, in the case of educational materials, the layout generation unit can apply an algorithm that provides a layout that includes easy-to-understand explanations. By applying a layout algorithm appropriate to the category of the document, more effective presentation materials can be generated. Some or all of the above-described processes in the layout generation unit may be performed using AI, for example, or without AI. For example, the layout generation unit can input the category data of the document into a generation AI and have the generation AI execute the application of the layout algorithm.

[0077] The layout generation unit can estimate the user's emotions and adjust the length of the layout based on the estimated emotions. For example, if the user is in a hurry, the layout generation unit can provide a short, concise layout. If the user is relaxed, the layout generation unit can also provide a longer layout with detailed explanations. Furthermore, if the user is excited, the layout generation unit can provide a layout with visually stimulating effects. By adjusting the length of the layout according to the user's emotions, more effective presentation materials can be generated. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the layout generation unit may be performed using AI or not. For example, the layout generation unit can input user facial expression data into the generative AI and have the generative AI perform the estimation of the user's emotions.

[0078] The layout generation unit can determine layout priorities based on the submission deadline of the document during layout generation. For example, in the case of an urgent document, the layout generation unit provides a layout that prioritizes the placement of the most important information. The layout generation unit can also provide a layout that focuses on the essentials for documents with an approaching submission deadline. Furthermore, for documents with a distant submission deadline, the layout generation unit can provide a layout that includes detailed explanations. By determining the layout priorities according to the submission deadline of the document, efficient presentation materials can be generated. Some or all of the above processing in the layout generation unit may be performed using AI, for example, or without AI. For example, the layout generation unit can input document submission deadline data into a generation AI and have the generation AI perform the determination of layout priorities.

[0079] The layout generation unit can adjust the order of the layout based on the relevance of the materials during layout generation. For example, the layout generation unit can provide a layout that places the main points of the materials first. It can also provide a layout that prioritizes the placement of highly relevant information in the materials. Furthermore, the layout generation unit can adjust the order of the layout based on the importance of the materials. By adjusting the order of the layout according to the relevance of the materials, more effective presentation materials can be generated. Some or all of the above processing in the layout generation unit may be performed using AI, for example, or without AI. For example, the layout generation unit can input the relevance data of the materials into a generation AI and have the generation AI perform the adjustment of the layout order.

[0080] The script generation unit can estimate the user's emotions and adjust the script's expression based on those emotions. For example, if the user is relaxed, the script generation unit can generate a script using softer language. It can also generate a script using concise and to-the-point language if the user is in a hurry. Furthermore, if the user is excited, the script generation unit can generate a script using visually stimulating language. This allows for the creation of more effective presentation materials by adjusting the script's expression according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI may be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the script generation unit may be performed using AI, or not. For example, the script generation unit can input user facial expression data into the generative AI and have the generative AI perform the user's emotion estimation.

[0081] The script generation unit can adjust the level of detail in the script based on the importance of the document during script generation. For example, for important documents, the script generation unit generates a script that includes detailed explanations. For general documents, the script generation unit can also generate a concise script. Furthermore, for urgent documents, the script generation unit can generate a script that gets straight to the point. By adjusting the level of detail in the script according to the importance of the document, more effective presentation materials can be generated. Some or all of the above processing in the script generation unit may be performed using AI, for example, or without AI. For example, the script generation unit can input document importance data into a generation AI and have the generation AI perform the adjustment of the level of detail in the script.

[0082] The script generation unit can apply different generation algorithms depending on the category of the document when generating a script. For example, in the case of presentation materials, the script generation unit can apply an algorithm that generates a persuasive script. In the case of proposal materials, the script generation unit can also apply an algorithm that generates a script that includes specific proposal content. Furthermore, in the case of educational materials, the script generation unit can apply an algorithm that generates a script that includes easy-to-understand explanations. By applying a generation algorithm according to the category of the document, more effective presentation materials can be generated. Some or all of the above-described processes in the script generation unit may be performed using AI, for example, or without AI. For example, the script generation unit can input the category data of the document into a generation AI and have the generation AI execute the application of the generation algorithm.

[0083] The script generation unit can estimate the user's emotions and adjust the length of the script based on the estimated emotions. For example, if the user is in a hurry, the script generation unit can generate a short, concise script. If the user is relaxed, the script generation unit can also generate a longer script with detailed explanations. Furthermore, if the user is excited, the script generation unit can generate a script with visually stimulating expressions. By adjusting the length of the script according to the user's emotions, more effective presentation materials can be generated. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the script generation unit may be performed using AI, or not using AI. For example, the script generation unit can input user facial expression data into the generative AI and have the generative AI perform the estimation of the user's emotions.

[0084] The script generation unit can determine the priority of scripts based on the submission deadline of the materials when generating scripts. For example, in the case of urgent materials, the script generation unit provides a script that prioritizes the generation of the most important information. The script generation unit can also provide a script that gets to the point for materials with an approaching submission deadline. Furthermore, the script generation unit can provide a script that includes detailed explanations for materials with a distant submission deadline. By determining the priority of scripts according to the submission deadline of the materials, efficient presentation materials can be generated. Some or all of the above processing in the script generation unit may be performed using AI, for example, or without AI. For example, the script generation unit can input the submission deadline data of the materials into a generation AI and have the generation AI perform the determination of script priority.

[0085] The script generation unit can adjust the order of scripts based on the relevance of the materials during script generation. For example, the script generation unit can provide a script that places the main points of the materials first. It can also provide a script that prioritizes the placement of highly relevant information in the materials. Furthermore, the script generation unit can adjust the order of scripts based on the importance of the materials. By adjusting the order of scripts according to the relevance of the materials, more effective presentation materials can be generated. Some or all of the above processing in the script generation unit may be performed using AI, for example, or without AI. For example, the script generation unit can input the relevance data of the materials into a generating AI and have the generating AI perform the adjustment of the script order.

[0086] The video generation unit can estimate the user's emotions and adjust the video's presentation based on those emotions. For example, if the user is relaxed, the video generation unit can generate a video using softer expressions. It can also generate a video using concise and to-the-point expressions if the user is in a hurry. Furthermore, if the user is excited, the video generation unit can generate a video using visually stimulating expressions. This allows for the creation of more effective presentation materials by adjusting the video's presentation according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI may be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the video generation unit may be performed using AI, or not. For example, the video generation unit can input user facial expression data into the generative AI and have the generative AI perform the user's emotion estimation.

[0087] The video generation unit can adjust the level of detail in a video based on the importance of the material during video generation. For example, if the material is important, the video generation unit will generate a video with detailed explanations. It can also generate a concise video for general materials. Furthermore, for urgent materials, the video generation unit can generate a video that summarizes the key points. By adjusting the level of detail in the video according to the importance of the material, more effective presentation materials can be generated. Some or all of the above processing in the video generation unit may be performed using AI, for example, or without AI. For example, the video generation unit can input material importance data into a generation AI and have the generation AI adjust the level of detail in the video.

[0088] The video generation unit can apply different generation algorithms depending on the category of the material when generating videos. For example, in the case of presentation materials, the video generation unit can apply an algorithm that generates a visually appealing video. In the case of proposal materials, the video generation unit can also apply an algorithm that generates a video that emphasizes the specific content of the proposal. Furthermore, in the case of educational materials, the video generation unit can apply an algorithm that generates a video that includes easy-to-understand explanations. By applying a generation algorithm according to the category of the material, more effective presentation materials can be generated. Some or all of the above processing in the video generation unit may be performed using AI, for example, or without AI. For example, the video generation unit can input the category data of the material into a generation AI and have the generation AI execute the application of the generation algorithm.

[0089] The video generation unit can estimate the user's emotions and adjust the video length based on the estimated emotions. For example, if the user is in a hurry, the video generation unit can generate a short, concise video. If the user is relaxed, the video generation unit can also generate a longer video with detailed explanations. Furthermore, if the user is excited, the video generation unit can generate a video with visually stimulating expressions. By adjusting the video length according to the user's emotions, more effective presentation materials can be generated. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the video generation unit may be performed using AI, or not using AI. For example, the video generation unit can input user facial expression data into the generative AI and have the generative AI perform the user's emotion estimation.

[0090] The video generation unit can prioritize videos based on the submission deadline of the materials during video generation. For example, in the case of urgent materials, the video generation unit will provide a video that prioritizes the most important information. It can also provide a concise video for materials with an approaching submission deadline. Furthermore, for materials with a distant submission deadline, the video generation unit can provide a video with detailed explanations. This allows for the efficient generation of presentation materials by prioritizing videos according to the submission deadline of the materials. Some or all of the above-described processes in the video generation unit may be performed using AI, for example, or without AI. For example, the video generation unit can input the submission deadline data of the materials into a generation AI and have the generation AI determine the video prioritization.

[0091] The video generation unit can adjust the order of videos based on the relevance of the materials during video generation. For example, the video generation unit can provide a video that places the main points of the materials first. It can also provide a video that prioritizes the placement of highly relevant information from the materials. Furthermore, the video generation unit can adjust the order of videos based on the importance of the materials. By adjusting the order of videos according to the relevance of the materials, more effective presentation materials can be generated. Some or all of the above processing in the video generation unit may be performed using AI, for example, or without AI. For example, the video generation unit can input the relevance data of the materials into a generation AI and have the generation AI perform the adjustment of the video order. === Hard Collateral 1-1 === Each of the multiple elements described above, including the reception unit, text generation unit, layout generation unit, script generation unit, and video generation unit, is implemented in at least one of the smart device 14 and the data processing device 12. For example, the reception unit is implemented by the control unit 46A of the smart device 14, where the user selects the type of material and inputs the necessary information and images to be used. The text generation unit is implemented by the specific processing unit 290 of the data processing device 12, where it analyzes the input information and generates the text portion of the presentation material. The layout generation unit is implemented by the specific processing unit 290 of the data processing device 12, where it creates the layout of the material based on the generated text. The script generation unit is implemented by the specific processing unit 290 of the data processing device 12, where it generates a talk script related to the generated material. The video generation unit is implemented by the control unit 46A of the smart device 14, where it generates a sample video based on the generated talk script. === Hard Collateral 1-2 === Each of the multiple elements described above, including the reception unit, text generation unit, layout generation unit, script generation unit, and video generation unit, is implemented in at least one of the smart glasses 214 and the data processing device 12. For example, the reception unit is implemented by the control unit 46A of the smart glasses 214, where the user selects the type of material and inputs the necessary information and images to be used. The text generation unit is implemented by the specific processing unit 290 of the data processing device 12, where the input information is analyzed and the text portion of the presentation material is generated. The layout generation unit is implemented by the specific processing unit 290 of the data processing device 12, where the layout of the material is created based on the generated text. The script generation unit is implemented by the specific processing unit 290 of the data processing device 12, where a talk script related to the generated material is generated. The video generation unit is implemented by the control unit 46A of the smart glasses 214, where a sample video is generated based on the generated talk script. === Hard Collateral 1-3 === Each of the multiple elements described above, including the reception unit, text generation unit, layout generation unit, script generation unit, and video generation unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the headset terminal 314, where the user selects the type of material and inputs necessary information and images to be used. The text generation unit is implemented by the specific processing unit 290 of the data processing unit 12, where the input information is analyzed and the text portion of the presentation material is generated. The layout generation unit is implemented by the specific processing unit 290 of the data processing unit 12, where the layout of the material is created based on the generated text. The script generation unit is implemented by the specific processing unit 290 of the data processing unit 12, where a talk script related to the generated material is generated. The video generation unit is implemented by the control unit 46A of the headset terminal 314, where a sample video is generated based on the generated talk script. === Hard Collateral 1-4 === Each of the multiple elements described above, including the reception unit, text generation unit, layout generation unit, script generation unit, and video generation unit, is implemented by, for example, at least one of the robot 414 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the robot 414, where the user selects the type of material and inputs the necessary information and images to be used. The text generation unit is implemented by the specific processing unit 290 of the data processing unit 12, where it analyzes the input information and generates the text portion of the presentation material. The layout generation unit is implemented by the specific processing unit 290 of the data processing unit 12, where it creates the layout of the material based on the generated text. The script generation unit is implemented by the specific processing unit 290 of the data processing unit 12, where it generates a talk script related to the generated material. The video generation unit is implemented by the control unit 46A of the robot 414, where it generates a sample video based on the generated talk script.

[0092] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0093] The reception desk collects evaluation data on users' past presentation materials, analyzes the characteristics of high-rated materials, and incorporates those characteristics into future presentations. For example, it can refer to the layout and text style of previously highly-rated materials. It can also suggest areas for improvement to users based on the evaluation data. Furthermore, it can provide customized advice to help users improve their presentation skills based on the evaluation data. This allows users to leverage past successes and create more effective presentation materials.

[0094] The reception desk can estimate the user's emotions and, based on that estimation, provide real-time feedback tailored to the progress of the document creation process. For example, if the user is feeling stressed, it can display encouraging messages appropriate to their progress. If the user is relaxed, it can provide detailed feedback and specifically indicate areas for improvement. Furthermore, if the user is in a hurry, it can provide concise feedback focused on the most important points. By providing appropriate feedback tailored to the user's emotions, the efficiency of document creation can be improved.

[0095] The reception desk analyzes the user's past document creation history and not only suggests the optimal information input method, but also learns the user's work patterns and can automatically select the most suitable template and layout for the next document creation. For example, it can prioritize presenting templates that the user frequently uses. It can also suggest an optimal work schedule based on the user's working hours and frequency. Furthermore, it can automatically set efficient work procedures based on the user's past work patterns. As a result, users can leverage their past work history and create documents more efficiently.

[0096] The reception desk can not only filter the user's current projects and areas of interest when selecting document types, but can also suggest highly relevant document types by considering the user's industry trends and the latest research findings. For example, it can present the most suitable document types based on the latest trends in the user's industry. It can also suggest document types that reflect the latest research findings related to the user's areas of interest. Furthermore, it can suggest the most suitable document types according to the progress of the user's project. This allows users to utilize the latest information and create more effective documents.

[0097] The reception desk can estimate the user's emotions and prioritize the information to be entered based on those emotions. Furthermore, it can dynamically change the interface design to match the user's emotions. For example, if the user is stressed, it can provide a simple and intuitive interface. If the user is relaxed, it can provide an interface that makes it easy to enter detailed information. Additionally, if the user is in a hurry, it can provide an interface that allows information to be entered with minimal effort. This improves the efficiency of information entry by providing an appropriate interface tailored to the user's emotions.

[0098] The reception desk not only prioritizes and presents highly relevant options based on the user's geographical location when selecting document types, but can also suggest region-specific designs and content based on the user's location. For example, if a user is in a specific region, it can suggest designs that match the local culture and customs. It can also provide content that includes region-specific data and case studies. Furthermore, it can suggest document types related to local events and trends based on the user's location. This allows users to utilize region-specific information and create more effective documents.

[0099] The reception desk can analyze the user's social media activity when selecting a type of document, not only presenting relevant options but also suggesting the most suitable type of document based on the user's social media influence and the interests of their followers. For example, if a user has many followers, it can suggest a type of document that matches the interests of those followers. It can also suggest highly relevant types of documents based on the user's social media activity. Furthermore, it can suggest the most suitable type of document considering the user's social media influence. This allows users to leverage their social media influence and create more effective documents.

[0100] The text generation unit not only estimates the user's emotions and adjusts the text's presentation based on those emotions, but it can also dynamically change the font and color of the text according to the user's mood. For example, if the user is relaxed, a soft font and calming colors will be used. If the user is in a hurry, a concise and highly legible font and colors can be used. Furthermore, if the user is excited, a visually stimulating font and colors can be used. By adjusting the text design according to the user's emotions, more effective presentation materials can be generated.

[0101] The text generation unit can adjust the level of detail of the text based on the importance of the material, as well as the difficulty level of the text according to the expertise level of the recipient. For example, it can provide detailed and specialized text to recipients with high expertise, and concise and easy-to-understand text to general audiences. Furthermore, it can provide text containing basic concepts for beginners. This allows for the generation of more effective presentation materials by providing appropriate text tailored to the audience.

[0102] The text generation unit not only applies different generation algorithms depending on the category of the document during text generation, but can also adjust the tone and style of the text according to the purpose of the document. For example, for persuasive presentation materials, an emphasized tone and persuasive style are used. For educational materials, a clear and approachable tone and style can be used. Furthermore, for proposal materials, a specific and practical tone and style can be used. This allows for the generation of more effective presentation materials by providing text appropriate to the purpose of the document.

[0103] The following briefly describes the processing flow for example form 2.

[0104] Step 1: The reception desk selects the type of document and inputs the necessary information and images to be used. For example, a user might select a presentation document and input information such as product features and market data. This information is then input into the text generation AI and image generation AI. Step 2: The text generation unit analyzes the input information and generates the text portion of the presentation materials. For example, the text generation unit generates text describing product features or text describing market data. It can also use generation AI to refine the expression of the text. Step 3: The layout generation unit creates the layout of the document based on the text generated by the text generation unit. For example, it creates layouts for slides explaining product features or slides showing market data in graphs. It can also automatically generate document layouts using image generation AI. Step 4: The script generation unit generates talk scripts related to the materials generated by the layout generation unit. For example, it generates talk scripts on how to explain slides describing product features or slides describing market data. It can also automatically generate talk scripts using generation AI. Step 5: The video generation unit generates example videos based on the talk scripts generated by the script generation unit. For example, it can generate example videos showing how to speak based on presentation materials or example videos appropriate for different situations. It can also automatically generate videos using generation AI.

[0105] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0106] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (for example, still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or Naive Bayes, and can perform a variety of operations, but is not limited to these examples. Furthermore, AI may also be an AI agent. Also, when the operations described above are performed by AI, the operations may be performed partially or entirely by AI, but is not limited to these examples. Additionally, operations performed by AI, including generative AI, may be replaced by rule-based operations, and rule-based operations may be replaced by operations performed by AI, including generative AI.

[0107] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0108] The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0109] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0110] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0111] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0112] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0113] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0114] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0115] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0116] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0117] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0118] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0119] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0120] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0121] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0122] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0123] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0124] The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0125] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0126] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0127] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0128] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0129] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0130] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0131] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0132] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0133] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0134] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0135] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0136] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0137] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0138] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0139] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0140] The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0141] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0142] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0143] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0144] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0145] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0146] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0147] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0148] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0149] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0150] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0151] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0152] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0153] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0154] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0155] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0156] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0157] The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0158] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0159] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0160] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0161] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0162] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0163] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0164] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0165] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0166] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0167] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0168] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0169] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0170] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0171] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0172] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0173] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0174] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0175] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0176] [Explanation of Symbols]

[0177] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. The reception desk allows you to select the type of document and enter the necessary information or images you wish to use. A text generation unit analyzes the information entered by the reception unit and generates the text portion of the presentation material, A layout generation unit creates a layout for a document based on the text generated by the text generation unit, A script generation unit that generates talk scripts related to the materials generated by the layout generation unit, The system includes a video generation unit that generates a sample video based on a talk script generated by the script generation unit. A system characterized by the following features.

2. The aforementioned reception unit is It estimates the user's emotions and presents a selection of material types based on those estimated emotions. The system according to feature 1.

3. The aforementioned reception unit is We analyze the user's past document creation history and suggest appropriate information input methods. The system according to feature 1.

4. The aforementioned reception unit is When selecting the type of document, filtering is performed based on the user's current projects and areas of interest. The system according to feature 1.

5. The aforementioned reception unit is It estimates the user's emotions and prioritizes the information to be entered based on the estimated user emotions. The system according to feature 1.

6. The aforementioned reception unit is When selecting a type of document, the system prioritizes presenting highly relevant options by considering the user's geographical location. The system according to feature 1.

7. The aforementioned reception unit is When selecting the type of material, the system analyzes the user's social media activity and presents relevant options. The system according to feature 1.

8. The text generation unit, It estimates the user's emotions and adjusts the way text is expressed based on those estimated emotions. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A