system

The system addresses the challenge of creating high-quality training videos by importing text-based materials, using generative AI to enhance layout and content, and allowing user corrections, resulting in efficient and accessible video production.

JP2026025551APending Publication Date: 2026-02-16SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024128360
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

Traditional training video production requires specialized knowledge and skills, leading to inefficiencies and high costs due to skill disparity among instructors, making it difficult for anyone to create high-quality training videos.

Method used

A system that imports text-based materials, analyzes them using generative AI to improve layout and content, automatically generates animated videos with narration, subtitles, and background music, and allows users to review and correct the content for high-quality video creation.

Benefits of technology

Enables anyone to easily create engaging and high-quality training videos with rich content, improving efficiency and reducing the need for specialized knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025551000001_ABST
    Figure 2026025551000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: This system is provided with a means for fetching the material of a text base, a means for analyzing the fetched material by using a generation AI, performing brush-up to screen constitution easy to see and hear and converting it into animation moving images and a means for automatically generating BGM matched with narration, subtitles and scenes.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Traditional training video production requires specialized knowledge and skills, which creates a problem of skill disparity among training instructors. Furthermore, if there are no specialized video creators, it is inefficient due to the significant cost and time consumption. There is a need to solve this problem and provide a method that allows anyone to easily create high-quality training videos. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for importing text-based materials into a tool, a means for analyzing the imported materials using a generation AI, brushing up the screen composition to make it easy to see and hear, and converting it into an animated video, and a means for automatically generating narration, subtitles, and background music that matches the scene.Furthermore, by including an interface means for a user to review the automatically generated video and input any necessary corrections, and a means for adjusting and correcting the video based on the user's correction instructions, the system realizes a system that can easily create high-quality training videos.

[0006] "Text-based materials" are document files that primarily contain textual information, including PDFs, Word documents, and text files.

[0007] "Means for importing" refers to the processing function that accepts files uploaded by users into the system and prepares them for analysis.

[0008] "Generative AI" is a technology that uses artificial intelligence to autonomously analyze the content of materials and extract and convert the information necessary for creating videos.

[0009] Analysis is the process of breaking down and understanding the content, structure, and visual elements of text-based material and identifying key points.

[0010] "A screen layout that is easy to see and hear" refers to a presentation format that uses appropriate layout, color, font, audio, etc. that is suited to human vision and hearing.

[0011] "Brushing up" is an operation to improve the content of existing materials and make modifications to improve visibility and comprehension.

[0012] Animated videos are videos that dynamically convey visual information through moving content.

[0013] "Narration" refers to audio data that reads the contents of a document aloud.

[0014] "Subtitles" are textual information that displays narration and important text information on the screen.

[0015] "BGM that matches the scene" refers to background music that is appropriate for the atmosphere and content of each scene in the video.

[0016] "Automatic generation means" is a function that uses artificial intelligence to automatically create and add narration, subtitles, background music, etc. from materials.

[0017] "Interface means" refers to the operation screen and input functions that allow the user to input and provide feedback to the system.

[0018] "Means for adjusting and correcting" refers to the function by which the system re-edits or corrects the video content based on feedback entered by the user.

[0019] "Multilingual support" refers to the system's ability to support multiple languages ​​and generate narration and subtitles in each language. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] This invention relates to a system that allows users to easily import text-based materials and automatically generate engaging training videos. Below, we will explain the program processing of this system in natural language and provide examples to illustrate how to implement the invention.

[0042] 1. User Import of Materials

[0043] Users upload text-based materials (PDF, Word, text files, etc.) to the system from their own devices, which then receive the materials and send them to the server.

[0044] 2. Analysis of data by the server

[0045] The server analyzes the received file and extracts text and image data. It identifies the file format and uses a text extraction module to extract character information. It also uses an image extraction module to obtain image data.

[0046] 3. Brushing up materials with generative AI

[0047] The server's generated AI analyzes the content and structure of the document. It identifies each page by section and identifies important keywords and phrases. It then improves the page layout based on basic principles of visual design to improve visibility. If necessary, the AI ​​generates appropriate images and charts and adds them to the document.

[0048] As a concrete example, if a user uploads "product training" materials, the generating AI will visually arrange each product's features, usage instructions, precautions, etc. in an easy-to-understand manner, adding diagrams and animations.

[0049] 4. Video generation and adding additional elements

[0050] The server uses generative AI to generate animated videos based on the content of the materials. It animates the information on each page along a timeline and uses a natural language processing module to generate narration from the text. If multilingual support is required, a translation module is used to generate narration. Subtitles are also automatically generated to match the narration. Furthermore, background music appropriate to the video content is automatically selected from a library and added to the video.

[0051] For example, in the case of "product training" materials, the generative AI visually represents the product's features, adds appropriate narration at each point, and provides subtitles and background music appropriate to the content.

[0052] 5. Check and correct the video

[0053] Users can check the generated videos on their devices and input any necessary corrections. The server then adjusts and corrects the video content based on the user's feedback. This allows users to easily create high-quality training videos.

[0054] The above processing steps realize a system that allows anyone to quickly generate training videos based on their own materials that are easy to watch and listen to and have rich content.

[0055] The processing flow will be explained below.

[0056] Step 1:

[0057] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system. The device receives the file specified by the user and sends it to the server.

[0058] Step 2:

[0059] The server analyzes the received file and extracts text and image data. Specifically, it identifies the file format (PDF, DOCX, TXT, etc.), extracts text information using a text extraction module, and obtains image data using an image extraction module.

[0060] Step 3:

[0061] The server analyzes the extracted text and images using generative AI, which identifies each page of the document by section and identifies important keywords and phrases.

[0062] Step 4:

[0063] The server-generated AI improves the page layout, rearranging information based on basic principles of visual design to improve readability, and adding appropriate AI-generated images and charts where necessary.

[0064] Step 5:

[0065] The server uses generative AI to generate animated videos based on the content of the documents, and displays the information on each page as an animation along a timeline.

[0066] Step 6:

[0067] The server generates narration using a natural language processing module, converts text information into speech, and generates narration using a translation module if multilingual support is required.

[0068] Step 7:

[0069] The server automatically generates subtitles to match the narration, creating text information to be displayed on the screen in sync with the narration.

[0070] Step 8:

[0071] The server automatically selects background music from its library that matches the content of the video and adds it to the video, selecting the best background music for each scene.

[0072] Step 9:

[0073] The terminal displays the generated video to the user and provides a preview function for the user to verify.

[0074] Step 10:

[0075] The user can review the generated video and input any necessary corrections. The user can use a designated input form to specify corrections to specific text, narration, subtitles, and animation.

[0076] Step 11:

[0077] The server executes the relevant modules again based on the user's feedback, adjusts and modifies the video content, and performs necessary processing according to the feedback content.

[0078] Step 12:

[0079] The device will present the final video file to the user and provide a download link, and the video will be exported in the specified format (e.g., MP4).

[0080] The above processing steps create a system that allows anyone to easily create high-quality training videos.

[0081] Example 1

[0082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0083] Traditionally, creating training videos based on text-based materials required a great deal of time and effort, and often required specialized knowledge. Furthermore, creating narration, subtitles, and selecting background music all required manual work, which added to the burden. This made it difficult for anyone to easily create high-quality training videos. Furthermore, creating multilingual narration and subtitles required advanced technology and time.

[0084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0085] In this invention, the server includes a means for importing text-based materials, a means for determining the format of the imported materials and extracting text data and image data, a means for analyzing the extracted materials using a generation AI, organizing the content by section, and brushing up the layout to make it easy to view, and a means for automatically generating narration, subtitles, and background music that matches the scenes. This enables users, even without specialized knowledge, to quickly generate training videos that are easy to view and listen to and have rich content based on their own materials.

[0086] "Text-based materials" refers to materials that contain written data such as PDFs, Word files, and text files.

[0087] "Format" refers to the data format of the material, examples of which include PDF format and Word format.

[0088] "Text data" refers to textual information extracted from materials.

[0089] "Image data" refers to image information extracted from materials.

[0090] "Generative AI" is a system that uses artificial intelligence technology to analyze and refine materials.

[0091] "Natural language processing" is a technology that allows computers to understand, analyze, and generate human language.

[0092] "Narration" is audio generated from text.

[0093] "Subtitles" are texts that appear within a video.

[0094] "BGM" refers to the background music added to a video.

[0095] "User's correction instructions" refer to feedback and requests for improvements from users regarding the content of the automatically generated video.

[0096] "Interface means" refers to the means by which a user accesses and operates the system, and specifically includes a web interface and an application.

[0097] "Multilingual support" refers to the ability to generate narration and subtitles in multiple languages.

[0098] This invention is a system that allows users to easily import text-based materials and automatically generate high-quality training videos. This system uses a server, terminal, generation AI, text extraction module, image extraction module, natural language processing module, speech synthesis module, translation module, narration generation module, background music selection module, and user interface. The specific implementation of this invention is described below.

[0099] 1. User Incorporation of Materials:

[0100] Users use their devices to upload text-based materials (PDF, Word, text files, etc.) to the system, which receives the materials and sends them to the server via an HTTP POST request.

[0101] 2. Analysis of data by the server:

[0102] The server determines the format of the received document file and extracts the text data using the appropriate text extraction module (e.g., PDF parser, Docx parser), and also uses the image extraction module to obtain the image data. Through this process, the text and image information in the document is stored in a database.

[0103] 3. Brushing up materials with generative AI:

[0104] The server-based generative AI analyzes the document's text and image data. It uses natural language processing (NLP) algorithms to identify important keywords and phrases and organize the content of each page into sections. It also improves page layout and legibility based on basic principles of visual design. If necessary, the generative AI's image generation module generates missing images and charts and adds them to the document.

[0105] 4. Video generation and additional elements:

[0106] The server generates animated videos based on the generated text and images. Based on the generated materials, it uses a natural language processing module and a speech synthesis module to generate narration from the text. If multilingual support is required, it uses a translation module to translate and generate narration and subtitles. Furthermore, it uses a background music selection module to automatically select background music that matches the content of the video and add it to the video.

[0107] 5. Check and correct the video:

[0108] Users can view the generated videos on their devices and provide feedback and corrections through the user interface. The server then adjusts and corrects the video content based on the user's corrections. This allows users to easily create high-quality training videos.

[0109] Examples:

[0110] For example, if a user uploads "product training" materials, the AI ​​will visually arrange each product's features, usage instructions, and precautions in an easy-to-understand manner, adding diagrams and animations as needed. Narration and subtitles are also automatically generated, and appropriate background music is selected.

[0111] Example prompt sentence:

[0112] Example prompt 1: "Upload the PDF version of our Product Training Materials and generate a training video highlighting key points for each section."

[0113] Example prompt 2: "Based on a Word document, extract text data and create a narrated training video with a clear layout."

[0114] As described above, by linking various modules, this system is able to automatically generate high-quality training videos and meet user needs quickly and efficiently.

[0115] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0116] Step 1:

[0117] A user uses their own device to select text-based materials (PDF, Word, text files, etc.) and upload them to the system. The device then sends the selected materials to the server. The input is the material file selected by the user, and the output is the material file sent to the server.

[0118] Specific behavior:

[0119] The user clicks the "Upload" button and selects the file they want to upload. The device then sends the selected file to the server via an HTTP POST request.

[0120] Step 2:

[0121] The server determines the format of the received document file. It then calls the appropriate text extraction module, such as a PDF parser for PDF format or a Docx parser for Word format, to extract the text data. It also uses an image extraction module to obtain image data. The input is the document file sent to the server, and the output is the extracted text data and image data.

[0122] Specific behavior:

[0123] The server checks the extension of the received file and passes the file to the appropriate parser module. The text extraction module extracts the text information and stores it in a database. The image extraction module extracts the image data and stores it in a database as well.

[0124] Step 3:

[0125] The server's generation AI analyzes the extracted text data and image data. It uses natural language processing algorithms to identify keywords and important phrases and organize the content of each page into sections. It also refines the page layout and generates and adds missing images and charts to improve visibility. The input is the extracted text data and image data, and the output is improved document data.

[0126] Specific behavior:

[0127] The generative AI analyzes the text data, extracts key keywords and phrases, rearranges the layout of each page, and incorporates newly generated images and charts into the document as needed.

[0128] Step 4:

[0129] The server generates animated videos based on the reference data improved by the generative AI. The server generates narration text using a natural language processing module and generates audio files using a speech synthesis module. If multilingual support is required, the translation module translates and generates narration and subtitles. In addition, the background music selection module automatically selects appropriate background music and adds it to the video. The input is the improved reference data, and the output is the generated animated video.

[0130] Specific behavior:

[0131] The generative AI creates animated slides based on the text and images on each page. The narration text is generated using a natural language processing module and converted into an audio file using a speech synthesis module. Appropriate music is selected from the background music library and added to the video.

[0132] Step 5:

[0133] The user checks the generated animation on the device and inputs feedback and corrections using the user interface. The server adjusts and corrects the video content based on the user's instructions. The input is the generated animation video and the user's feedback, and the output is the final video that has been corrected and adjusted.

[0134] Specific behavior:

[0135] The user opens a preview of the video on their device and inputs corrections as they are made. The server updates the video content based on the user's feedback and generates the final version.

[0136] As described above, this system processes and analyzes input data at each processing step, automatically generating high-quality training videos that meet the user's needs.

[0137] (Application example 1)

[0138] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0139] In modern business and educational settings, training and instruction using text-based materials is common. However, these materials are generally visually unappealing and lack narration or video formats, resulting in low learning effectiveness. Furthermore, for many users, video editing and adding narration, which require specialized knowledge, are difficult and time-consuming. Furthermore, creating training videos tailored to specific tasks, such as multilingual support, operation manuals, and security training, requires additional specialized skills. Therefore, there is a need for a system that allows anyone to easily generate high-quality training videos and apply them to operation guides and security training.

[0140] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0141] In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, polishing the screen layout to make them easier to see and hear, and converting them into animated videos, a means for automatically generating narration, subtitles, and background music that matches the scenes, a means for converting the narration from text to audio, and a means for automatically selecting and generating visual decoration elements to be used in generating the videos. This allows users to simply upload text-based materials and automatically generate high-quality training videos that are visually appealing and include additional elements such as multilingual narration, subtitles, and appropriate background music. The system also provides an interface that allows users to review the generated videos and make necessary corrections, enabling them to make corrections according to specific tasks, making it possible to generate videos for specific purposes such as operation guides and security training.

[0142] "Text-based materials" refers to digital documents that primarily contain text, such as PDFs, Word files, and text files.

[0143] "Generative AI" refers to artificial intelligence that generates new content based on given data.

[0144] "Analysis" is the process of analyzing data and understanding and extracting its contents.

[0145] "Screen configuration" refers to the layout that visually displays information.

[0146] "Animated video" refers to video content that has movement and change.

[0147] "Narration" refers to audio that provides explanations and commentary in videos and presentations.

[0148] "Subtitles" are elements that display textual information in videos and presentations.

[0149] A "scene" refers to a specific situation or situation within a video.

[0150] "BGM" refers to the music played in the background of a video or presentation.

[0151] "Convert to speech" refers to the process of converting text information into audio data.

[0152] "Visual decoration" refers to visual information such as images or icons used to enhance visual appeal.

[0153] "User" means any individual or entity that uses the system or service.

[0154] An "interface" is the means by which a system and a user interact with each other to perform operations and exchange information.

[0155] "Modification instructions" refer to instructions for changes that the user makes to the generated video.

[0156] "Multilingual support" refers to having the functionality to support multiple languages.

[0157] An "operation guide" is a document or video that explains how to use a particular device or system.

[0158] "Security training" refers to training designed to provide knowledge about security measures and national security.

[0159] The present invention relates to a system for automatically generating attractive training videos for use as operational guides or security training for electronic payment services. Specific embodiments of this system are described below.

[0160] System program description

[0161] This system allows users to easily import text-based materials and automatically generate engaging training videos. It mainly consists of the following processing steps:

[0162] 1. Importing materials:

[0163] The server receives text-based materials such as PDFs, Word documents, and text files uploaded from users' devices. After receiving the files, the server extracts text and image data to analyze the contents of the files.

[0164] 2. Analysis of the material:

[0165] The server uses a text extraction module (e.g., Python's PIL library) to extract text information from the materials, and an image extraction module (e.g., PIL library) to obtain image data.

[0166] 3. Brushing up with generative AI:

[0167] The server's generative AI (e.g., GPT-3 or Stable Diffusion) analyzes the content of the document and refines the screen composition to make it easier to read and listen to. It improves the page layout based on basic principles of visual design and generates appropriate images and charts to add to the document.

[0168] 4. Video generation:

[0169] It automatically generates narration, subtitles, and background music to match the scenes, and generates animated videos. The narration is converted from text to speech using a natural language processing module (e.g., pyttsx3), and images are converted to video using MoviePy.

[0170] 5. Check and correct the video:

[0171] The user checks the generated video and inputs any necessary corrections. The server readjusts and corrects the video content based on the user's feedback and provides an optimized video.

[0172] This system allows users to significantly improve their work efficiency and easily create high-quality operation guides and security training videos.

[0173] Hardware and software used

[0174] Text and image data extraction:

[0175] Software used: Python library (PIL, FPDF)

[0176] Enhanced by generative AI:

[0177] AI model used: Generative AI (e.g., GPT-3, Stable Diffusion)

[0178] Software used: NLP library (spacy), image processing library (PIL)

[0179] Video Generation:

[0180] Tools used: MoviePy, pyttsx3

[0181] User Interface:

[0182] Devices used: smartphone, PC

[0183] Specific examples

[0184] For example, if you use security training materials for an electronic payment service:

[0185] 1. The user uploads "Security Guide.pdf".

[0186] 2. The server extracts important security measures from the guide.

[0187] 3. Use the extracted information to generate a video that visually explains security measures.

[0188] 4. The video is accompanied by narration and background music, and the user watches the training video.

[0189] 5. The user gives feedback saying, "I want more emphasis on password management," and the server makes the corrections.

[0190] Example of input prompt for generative AI model

[0191] Example prompt:

[0192] "Please extract important security measures from this document and create a layout that visually explains them. The contents of the document are as follows: (Text content)"

[0193] This makes it possible for anyone to easily generate high-quality, visually appealing training videos, contributing to training and improving security awareness among users of electronic payment services.

[0194] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0195] Step 1:

[0196] The terminal uploads the user's text-based document file (PDF, Word, text file, etc.). The input is the document file that the user specifies on the terminal, and the output is that file being transferred to the server. Specifically, the file is uploaded to the server by opening a file selection dialog on the terminal and clicking the send button.

[0197] Step 2:

[0198] The server analyzes the uploaded file. After receiving the file, it determines the file format and extracts the text and image data. The input is the uploaded file, and the output is the extracted text and image data. Specifically, it uses the Python PIL library to analyze and extract the text and images.

[0199] Step 3:

[0200] The server uses a generative AI model to analyze the content and structure of the document and refine it to make it easier to read and listen to. The input is extracted text data and image data, and the output is refined text and images. Specifically, generative AI (e.g., GPT-3, Stable Diffusion) is used to identify important keywords and phrases in the document and improve the page layout based on them.

[0201] Step 4:

[0202] The server uses generative AI to generate animated videos based on the content of the materials. It animates the information on each page along a timeline and uses a natural language processing module to generate narration from the text. The input is the polished text and images, and the output is an animated video file. Specifically, MoviePy and pyttsx3 are used to create the video and audio.

[0203] Step 5:

[0204] The server automatically generates narration, subtitles, and background music to match the scene. The input is the generated animation video file, and the output is the final video file containing narration, subtitles, and background music. Background music is automatically selected from a library, and subtitles are automatically generated to match the narration.

[0205] Step 6:

[0206] The terminal displays the generated video to the user and prompts them to check it. The user checks the generated video and inputs any necessary corrections. The input is a correction instruction from the user, and the output is the correction instruction being sent to the server. Specifically, the user inputs the correction content using the correction interface and clicks the send button.

[0207] Step 7:

[0208] The server readjusts and modifies the video content based on the user's correction instructions. The input is the user's correction instructions, and the output is the final modified video file. Specifically, it uses regeneration AI to correct the specified parts, regenerates narration and subtitles, and adds necessary visual decoration elements.

[0209] This series of processing steps enables users to easily generate high-quality operation guides and security training videos.

[0210] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0211] This invention relates to a system that allows users to easily import text-based materials, which are then analyzed by a generative AI to automatically generate high-quality training videos. In particular, by combining an emotion engine, this system is characterized by recognizing user emotions and enabling more adaptive content creation that reflects that feedback. Below, the program processing of this system is explained in natural language, and specific examples are provided to illustrate the mode of carrying out the invention.

[0212] 1. User Import of Materials

[0213] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system, which then receives the materials and sends them to the server.

[0214] 2. Analysis of data by the server

[0215] The server analyzes the received file and extracts text and image data. It identifies the file format (PDF, DOCX, TXT, etc.), extracts text information using a text extraction module, and obtains image data using an image extraction module.

[0216] 3. Brushing up materials with generative AI

[0217] The server-generated AI analyzes the content and structure of the document, identifies each page into sections and identifies important keywords and phrases, improves page layout based on basic principles of visual design to improve visibility, and adds appropriate AI-generated images and charts to the document as needed.

[0218] As a concrete example, if a user uploads "product training" materials, the generating AI will visually arrange each product's features, usage instructions, precautions, etc. in an easy-to-understand manner, adding diagrams and animations.

[0219] 4. Video generation and adding additional elements

[0220] The server uses generative AI to generate animated videos based on the content of the materials. The information on each page is animated along a timeline, and a natural language processing module is used to generate narration. If multilingual support is required, a translation module is used to generate narration. Subtitles are also automatically generated to match the narration, and background music appropriate to the video content is automatically selected from a library and added to the video.

[0221] For example, in the case of "product training" materials, the generative AI will visually represent the product's features, add appropriate narration for each point, and provide subtitles and background music appropriate to the content.

[0222] 5. Emotion Analysis Using an Emotion Engine

[0223] When a user watches a video on a device, the device collects the user's facial expression and voice data. This data is sent to a server, and the server's emotion engine recognizes the user's emotions in real time. The emotion engine uses, for example, facial recognition technology and voice analysis technology to determine whether the user is interested, confused, or expressing other specific emotions.

[0224] 6. Adaptive Content Adjustment

[0225] Based on the emotion engine's judgment, the server dynamically changes some or all elements in the video (such as narration, subtitles, background music, etc.) For example, if the server determines that the user is confused, it changes the narration to make it easier to understand and adds subtitles that re-emphasize important points.

[0226] 7. Check and correct the video

[0227] The user checks the generated video and inputs any necessary corrections. The server adjusts and corrects the video content based on the user's feedback. The emotion engine also provides automatic feedback according to changes in the user's emotions, allowing the video to be optimally adjusted.

[0228] Through the above processing steps, users can easily create more intuitive and effective training videos through a system that combines an emotion recognition engine.

[0229] The processing flow will be explained below.

[0230] This invention relates to a system that takes text-based materials and automatically generates high-quality training videos using generative AI and an emotion engine. Specific processing steps are explained below.

[0231] Step 1:

[0232] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system, which then receives the materials and sends them to the server.

[0233] Step 2:

[0234] The server analyzes the received document file and extracts text and image data. Specifically, it determines the file format and extracts the text using a text extraction engine if it is a PDF, or a different text extraction engine if it is a Word document, and performs the same process on image data.

[0235] Step 3:

[0236] The server analyzes the extracted text and images using generative AI, which identifies each page and section of the document and uses natural language processing techniques to extract key keywords and phrases.

[0237] Step 4:

[0238] Server-generated AI improves page layout: the layout engine follows basic principles of visual design to optimize text position, font, color, and image placement, and adds appropriate AI-generated images and charts where necessary.

[0239] Step 5:

[0240] The server uses generative AI to generate animated videos based on the content of the document, and an engine runs to animate the information on each page along a timeline, creating visually appealing videos.

[0241] Step 6:

[0242] The server automatically generates narration using a natural language processing module, text-to-speech synthesis technology, and, if multilingual support is required, a translation module is used to generate narration in each language.

[0243] Step 7:

[0244] The server automatically generates subtitles to match the narration. The subtitle generation engine matches the narration text with the timestamps to generate subtitle data, which is then incorporated into the video.

[0245] Step 8:

[0246] The server automatically selects the best background music from the library to fit the video content, and adds it to the video using the BGM addition engine. The most suitable music for each scene is selected and played as background music.

[0247] Step 9:

[0248] When a user watches a video, the device collects facial expression and voice data from the user. It captures real-time video and audio through a camera and microphone and sends the data to a server.

[0249] Step 10:

[0250] The server's emotion engine analyzes the user's facial expressions and voice data to recognize their emotions. An emotion recognition algorithm is used to identify whether the user is interested, confused, happy, etc.

[0251] Step 11:

[0252] The server dynamically adjusts the video's narration, subtitles, and background music based on the analysis results of the emotion engine. For example, if the server determines that the user is confused, it will change the narration to make it easier to understand and highlight the subtitles and narration.

[0253] Step 12:

[0254] The device displays the generated video to the user through a preview function, and the user can check the video and provide feedback if any corrections are required.

[0255] Step 13:

[0256] The server adjusts and modifies the video based on user feedback, and then uses related modules (generative AI, emotion engine, etc.) to improve the video content.

[0257] Step 14:

[0258] The device will provide the final video file to the user in a downloadable format, exporting the video in the specified format (e.g. MP4).

[0259] Through the above processing steps, users can easily create more adaptive and effective training videos through a system that combines an emotion recognition engine.

[0260] Example 2

[0261] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0262] Conventional training video generation systems lack sufficient visual and audio quality, requiring users to spend a lot of time manually making corrections. Furthermore, they lack the ability to dynamically adjust content based on the viewer's emotions, limiting the effectiveness of the content. Furthermore, their lack of multilingual support makes them difficult to use internationally. These issues need to be addressed.

[0263] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0264] In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, polishing the screen composition to make them easier to see and hear, and converting them into animated videos, a means for automatically generating narration, subtitles, and background music to match the scenes, an emotion analysis means for collecting user facial and voice data and recognizing emotions in real time, and a means for dynamically adjusting the video content based on the emotion analysis results. This not only improves the visual and auditory quality of the materials, but also enables adaptive adjustments based on the viewer's emotions in real time, resulting in the creation of more effective training videos. Furthermore, multilingual support for narration and subtitles facilitates international use.

[0265] "Text-based materials" are electronic documents that consist primarily of textual information and are provided in formats such as PDFs, Word documents, and text files.

[0266] "Generative AI" is a software system that uses artificial intelligence techniques to analyze data and automate complex tasks.

[0267] "Narration" refers to the commentary or explanation provided as audio within a video.

[0268] "Subtitles" are narrations and audio content displayed as textual information, in a format that can be visually understood by viewers.

[0269] "BGM that matches the scene" automatically selects and adds background music that matches each scene in the video.

[0270] "User facial expression and voice data" refers to data obtained from facial expressions and voice collected while a user is watching a video.

[0271] "Real-time emotion analysis" is a process that uses facial expression recognition and voice analysis technology to instantly analyze and determine a user's emotions.

[0272] "Dynamic adjustment of video content" refers to the operation of changing and optimizing elements such as video narration, subtitles, and background music in real time based on the results of sentiment analysis.

[0273] "Multilingual" refers to the ability of a system to function in multiple languages ​​and to translate and generate narration and subtitles into those languages.

[0274] An "interface means" is a software component that provides a mechanism for a user to perform operations and inputs through interaction with the system.

[0275] This invention is a system that allows users to easily import text-based materials and automatically generates high-quality training videos using generative AI. In particular, it combines an emotion engine to recognize user emotions in real time and enable more adaptive content creation that reflects feedback.

[0276] First, the user selects text-based materials (e.g., PDF, Word, text files, etc.) from their device and uploads them to the system. The device reads the materials from local storage and sends them to the server. Specifically, the user sends "Product Training Materials.pdf" to the server using the system's upload button.

[0277] Next, the server receives the uploaded file and identifies its format. It uses the server's text extraction module to extract text information and its image extraction module to obtain image data. For example, the server analyzes "Product Training Materials.pdf" and extracts the text and images for each page.

[0278] The server's generative AI then analyzes the content and structure of the materials. Generative AI is a software system that uses artificial intelligence technology to analyze data and automate complex tasks. It divides each page into sections and identifies important keywords and phrases. It improves the page layout based on basic principles of visual design to improve visibility. If necessary, it generates additional images and diagrams and incorporates them into the materials. As a specific example, the generative AI identifies and highlights "product features," "usage instructions," and "precautions" from the main text of "product training materials."

[0279] The server then begins generating videos based on the analyzed and polished materials. Animated slides are created using text and image data, and a natural language processing module generates narration. A translation module generates narration in multiple languages ​​and automatically generates subtitles to accompany the narration. For example, a slide showing "Product Features" is animated, and a narrator generates the following: "Product A is the latest model with the best performance on the market." Appropriate background music for the video is then selected from the library and added as "BGM_Upbeat.mp3."

[0280] When a user watches a video, the device uses a camera and microphone to collect the user's facial expression and voice data. This data is then sent to a server, where the server's emotion analysis engine uses facial expression and voice analysis technology to recognize the user's emotions in real time. For example, if the device streams camera data and detects the user's smile, the server determines that the user is in an "interested" state.

[0281] Based on the results of the sentiment analysis, the server dynamically adjusts the video content. If it determines that the user is confused, it changes the narration to make it easier to understand and adds additional subtitles. For example, if the user is determined to be in a "confused" state, it changes the narration to "What's important here is the safety features of product A. Let me explain in detail," and adds a subtitle "Safety feature details" with key points highlighted.

[0282] Finally, the user reviews the generated video and inputs any necessary corrections through the interface. The server receives the user's feedback and readjusts or corrects the video content. For example, if the user inputs a correction request saying "the narration is too fast," the server presents the user with a regenerated video with the narration speed adjusted.

[0283] Examples of prompt statements

[0284] User: "I've uploaded some product training materials. I'd like to see videos that visually explain the features and usage of each product."

[0285] Generative AI: "Analyzing your document. Identifying each section and identifying keywords. Improving page layout and adding appropriate images and charts."

[0286] Emotion Engine: "We detected that your audience was confused. We'll update the narration to make it clearer and add subtitles to re-emphasize key points."

[0287] Through the above process, users can easily create intuitive and effective training videos through a system that combines an emotion analysis engine.

[0288] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0289] Step 1:

[0290] User import of materials

[0291] A user selects text-based materials (e.g., PDF, Word, text files, etc.) on their device and uploads them to the system. The device reads the materials from local storage and sends them to the server. For example, the user selects "Product Training Materials.pdf" and clicks the system's upload button. The device reads the selected file and sends an HTTP POST request to the server. The input is the material file selected by the user, and the output is the material file sent to the server.

[0292] Step 2:

[0293] Server analysis of data

[0294] The server receives the document file sent from the terminal. It determines the file format (e.g., PDF, DOCX, TXT), extracts text information using the text extraction module, and obtains image data using the image extraction module. For example, the server receives "Product Training Materials.pdf" and uses a PDF parser to extract the text and images for each page. The input is the received document file, and the output is the extracted text data and image data.

[0295] Step 3:

[0296] Brushing up materials with generative AI

[0297] The server's generation AI analyzes the content and structure of the materials. It divides each page into sections and identifies important keywords and phrases. It improves the page layout based on basic principles of visual design to improve visibility. If necessary, it generates additional images and diagrams and incorporates them into the materials. Specifically, the generation AI identifies "product features," "usage instructions," and "precautions" from the main text of the "product training materials" and highlights this information. The input is the extracted text and image data, and the output is polished material data.

[0298] Step 4:

[0299] Video generation and adding additional elements

[0300] The server begins generating videos based on the analyzed and refined materials. Animated slides are created using text and image data, and a natural language processing module is used to generate narration. A translation module generates narration in multiple languages ​​as needed. Subtitles are automatically generated to match the narration, and appropriate background music is selected from a library and added to the video. For example, a "Product Features" slide is animated, and the narrator generates, "Product A is the latest model with the best performance on the market." Additionally, background music suitable for the video is selected from a library and added as "BGM_Upbeat.mp3." The input is the refined material data, and the output is the generated animated video.

[0301] Step 5:

[0302] Emotion analysis using an emotion engine

[0303] When a user watches a video, the device uses a camera and microphone to collect the user's facial and voice data. This data is sent to a server, where the server's emotion analysis engine uses facial and voice analysis technologies to recognize the user's emotions in real time. For example, if the device streams camera data and detects a user's smile, the server determines that the user is in an "interested" state. The input is the collected facial and voice data, and the output is emotion data determined in real time.

[0304] Step 6:

[0305] Adaptive content adjustment

[0306] The server dynamically adjusts elements in the video (such as narration, subtitles, and background music) based on the results of sentiment analysis. For example, if it determines that the user is confused, it changes the narration to something more understandable and adds subtitles that re-emphasize important points. For example, it changes the narration to "What's important here are the safety features of product A. Let me explain in detail," and adds a subtitle "Safety feature details" that highlights the important points. The input is the sentiment analysis results, and the output is the dynamically adjusted video content.

[0307] Step 7:

[0308] Check and correct the video

[0309] The user reviews the generated video and inputs any necessary corrections through the interface. The server receives the user's feedback and readjusts and corrects the video content. For example, if the user inputs a correction request saying "the narration is too fast," the server presents the user with a regenerated video with the narration speed adjusted. The input is the user's correction feedback, and the output is the corrected video.

[0310] Through the above processing steps, users can easily create intuitive and effective training videos through a system that combines an emotion analysis engine.

[0311] (Application example 2)

[0312] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0313] Conventional training video creation systems have difficulty adjusting content in real time to take user emotions into account, limiting their effectiveness in improving viewer comprehension. Furthermore, the generation of multilingual narration and subtitles is inefficient because it requires manual editing due to insufficient automation.

[0314] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, brushing up the screen composition to make it easy to see and hear, and converting it into an animated video, a means for automatically generating narration, subtitles, and background music that matches the scene, and a means for analyzing the user's emotions in real time using an emotion engine and dynamically adjusting the video content. This makes it possible to automatically generate high-quality training videos that adapt to the user's emotions and also easily support multiple languages.

[0315] "Text-based materials" refers to documents that primarily contain text, such as PDFs, Word documents, and text files.

[0316] "Generative AI" refers to a system that uses artificial intelligence technology to analyze input data and generate new content.

[0317] "Brushing up" refers to improving input materials to make them visually easier to read and understand.

[0318] An "animated video" refers to a video that expresses movement by displaying multiple still images in sequence.

[0319] "Narration" refers to audio commentary that accompanies a video scene.

[0320] "Subtitles" refers to the display of video content and audio as text.

[0321] A "scene" refers to a specific situation or moment within a video.

[0322] "BGM" is an abbreviation for background music and refers to the music played in the background of a video.

[0323] "Emotion engine" refers to artificial intelligence technology that recognizes and analyzes a user's emotional state from their facial expressions and voice.

[0324] "Analyzing user emotions in real time" means determining the user's emotions from their facial expressions and voice at the moment they are watching a video.

[0325] "Dynamic adjustment of video content" refers to changing elements of a video, such as narration and subtitles, in real time based on the results of user sentiment analysis.

[0326] "Smart devices" refers to mobile information terminals with computer functions, such as smartphones and tablets.

[0327] "Collecting facial and voice data" refers to recording the user's facial movements and vocalizations using the smart device's camera and microphone.

[0328] This invention is a system that allows users to easily import text-based materials using a smart device, analyze them with a generative AI, and automatically generate high-quality training videos using an emotion engine. Specific embodiments of this system are described below.

[0329] 1. User Import of Materials

[0330] Users select text-based materials from their smart devices (such as smartphones or tablets) and upload them to the system. During this process, users press a button to open a file selection dialog, and the selected file (PDF, Word, or text file) is sent to the server. The hardware used here is a smart device, and the software is React Native and Amazon S3.

[0331] 2. Analysis of the data

[0332] The server analyzes the document files received from Amazon S3 and extracts text and image data. It identifies the file format (PDF, DOCX, TXT) and uses a specific text extraction module to extract character information. This process includes basic principles of visual design, allowing the generative AI to analyze the content and structure of the document and improve the page layout.

[0333] 3. Video generation and adding additional elements

[0334] The server uses generative AI to generate animated videos based on the content of the materials. It automatically generates narration, subtitles, and background music to match the scenes, improving visibility and comprehension. If multilingual support is required, a translation module is used to generate narration and automatically generate subtitles. The software used includes generative AI models and natural language processing technology.

[0335] 4. Emotion Analysis Using an Emotion Engine

[0336] When a user watches a video on a device, the smart device's camera and microphone are used to collect facial expressions and voice data, and emotions are analyzed in real time. The emotion engine uses facial recognition technology (OpenCV) and voice analysis technology (Google Cloud Speech-to-Text) to recognize the user's emotions.

[0337] 5. Real-time content adjustment

[0338] The server dynamically adjusts the video's narration and subtitles based on the analysis results of the emotion engine, allowing it to add additional narration explanations if the user is confused and display subtitles that emphasize important points.

[0339] 6. Check and correct the video

[0340] The user can review the generated video and input any necessary corrections. The server then adjusts the video content based on the user's feedback and provides the optimal training video.

[0341] For example, if a user uploads "product training" materials, the generative AI will visually arrange each product's features, usage instructions, and precautions in an easy-to-understand manner, automatically generating diagrams, narration, and subtitles. Furthermore, the emotion engine will analyze the user's facial expressions and voice, and if it determines that their understanding is incomplete, it will insert additional explanations into the video.

[0342] Example prompt sentence:

[0343] Please select the file you want to upload.

[0344] "Data analysis completed."

[0345] "We recognize the emotions of our users."

[0346] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0347] Step 1:

[0348] A user uses a smart device to select text-based materials and upload them to the system. Specifically, the user presses the file selection button and selects either a PDF, Word, or text file from the file selection dialog on the smart device. The selected file is sent to Amazon S3 by the device. The input is the file selected by the user, and the output is the file uploaded to the server.

[0349] Step 2:

[0350] The server receives documents uploaded to Amazon S3 and determines the file format (PDF, DOCX, TXT). It uses a text extraction module to extract text information, and if necessary, an image extraction module to obtain image data. The input is the uploaded document file, and the output is the extracted text data and image data. Specifically, it analyzes the file format and runs the extraction module corresponding to each format.

[0351] Step 3:

[0352] The server's generative AI analyzes the content and structure based on the analyzed material. It identifies important keywords and phrases and improves the page layout based on visual design. If necessary, it adds images and charts generated by the generative AI to the material. The input is the extracted text data and image data, and the output is a polished page layout and additional visual elements. Specifically, the generative AI analyzes the text and reconstructs the page layout.

[0353] Step 4:

[0354] The server uses generative AI to generate animated videos based on the content of the materials. The information on each page is animated along a timeline, and a natural language processing module is used to generate narration. Furthermore, if multilingual support is required, a translation module is used to generate narration and automatically generate subtitles. The input is the polished materials and layout information, and the output is an animated video with narration. Specific operations include natural language processing and timeline setting.

[0355] Step 5:

[0356] When a user watches a video on a device, facial expression and voice data are collected using the device's camera and microphone. This data is sent to a server in real time and analyzed by an emotion engine. The input is the user's facial expression and voice data, and the output is analyzed emotion data. Specifically, the system collects camera video and voice data in real time and performs facial recognition and voice analysis.

[0357] Step 6:

[0358] The server dynamically adjusts the video content based on the results of emotion analysis. If it determines that the user is confused, it adds supplementary narration explanations and displays subtitles that re-emphasize important points. The input is the analyzed emotion data, and the output is the adjusted video content. Specific operations include updating the narration content and subtitles.

[0359] Step 7:

[0360] The user reviews the generated video and inputs corrections as necessary. The server receives the user's feedback and readjusts the video content. The input is the user's feedback, and the output is the final corrected video. Specific operations include collecting feedback through the interface and re-editing the video.

[0361] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0362] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0363] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0364] [Second embodiment]

[0365] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0366] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0367] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0368] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0369] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0370] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0371] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0372] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0373] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0374] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0375] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0376] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0377] This invention relates to a system that allows users to easily import text-based materials and automatically generate engaging training videos. Below, we will explain the program processing of this system in natural language and provide examples to illustrate how to implement the invention.

[0378] 1. User Import of Materials

[0379] Users upload text-based materials (PDF, Word, text files, etc.) to the system from their own devices, which then receive the materials and send them to the server.

[0380] 2. Analysis of data by the server

[0381] The server analyzes the received file and extracts text and image data. It identifies the file format and uses a text extraction module to extract character information. It also uses an image extraction module to obtain image data.

[0382] 3. Brushing up materials with generative AI

[0383] The server's generated AI analyzes the content and structure of the document. It identifies each page by section and identifies important keywords and phrases. It then improves the page layout based on basic principles of visual design to improve visibility. If necessary, the AI ​​generates appropriate images and charts and adds them to the document.

[0384] As a concrete example, if a user uploads "product training" materials, the generating AI will visually arrange each product's features, usage instructions, precautions, etc. in an easy-to-understand manner, adding diagrams and animations.

[0385] 4. Video generation and adding additional elements

[0386] The server uses generative AI to generate animated videos based on the content of the materials. It animates the information on each page along a timeline and uses a natural language processing module to generate narration from the text. If multilingual support is required, a translation module is used to generate narration. Subtitles are also automatically generated to match the narration. Furthermore, background music appropriate to the video content is automatically selected from a library and added to the video.

[0387] For example, in the case of "product training" materials, the generative AI visually represents the product's features, adds appropriate narration at each point, and provides subtitles and background music appropriate to the content.

[0388] 5. Check and correct the video

[0389] Users can check the generated videos on their devices and input any necessary corrections. The server then adjusts and corrects the video content based on the user's feedback. This allows users to easily create high-quality training videos.

[0390] The above processing steps realize a system that allows anyone to quickly generate training videos based on their own materials that are easy to watch and listen to and have rich content.

[0391] The processing flow will be explained below.

[0392] Step 1:

[0393] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system. The device receives the file specified by the user and sends it to the server.

[0394] Step 2:

[0395] The server analyzes the received file and extracts text and image data. Specifically, it identifies the file format (PDF, DOCX, TXT, etc.), extracts text information using a text extraction module, and obtains image data using an image extraction module.

[0396] Step 3:

[0397] The server analyzes the extracted text and images using generative AI, which identifies each page of the document by section and identifies important keywords and phrases.

[0398] Step 4:

[0399] The server-generated AI improves the page layout, rearranging information based on basic principles of visual design to improve readability, and adding appropriate AI-generated images and charts where necessary.

[0400] Step 5:

[0401] The server uses generative AI to generate animated videos based on the content of the documents, and displays the information on each page as an animation along a timeline.

[0402] Step 6:

[0403] The server generates narration using a natural language processing module, converts text information into speech, and generates narration using a translation module if multilingual support is required.

[0404] Step 7:

[0405] The server automatically generates subtitles to match the narration, creating text information to be displayed on the screen in sync with the narration.

[0406] Step 8:

[0407] The server automatically selects background music from its library that matches the content of the video and adds it to the video, selecting the best background music for each scene.

[0408] Step 9:

[0409] The terminal displays the generated video to the user and provides a preview function for the user to verify.

[0410] Step 10:

[0411] The user can review the generated video and input any necessary corrections. The user can use a designated input form to specify corrections to specific text, narration, subtitles, and animation.

[0412] Step 11:

[0413] The server executes the relevant modules again based on the user's feedback, adjusts and modifies the video content, and performs necessary processing according to the feedback content.

[0414] Step 12:

[0415] The device will present the final video file to the user and provide a download link, and the video will be exported in the specified format (e.g., MP4).

[0416] The above processing steps create a system that allows anyone to easily create high-quality training videos.

[0417] Example 1

[0418] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0419] Traditionally, creating training videos based on text-based materials required a great deal of time and effort, and often required specialized knowledge. Furthermore, creating narration, subtitles, and selecting background music all required manual work, which added to the burden. This made it difficult for anyone to easily create high-quality training videos. Furthermore, creating multilingual narration and subtitles required advanced technology and time.

[0420] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0421] In this invention, the server includes a means for importing text-based materials, a means for determining the format of the imported materials and extracting text data and image data, a means for analyzing the extracted materials using a generation AI, organizing the content by section, and brushing up the layout to make it easy to view, and a means for automatically generating narration, subtitles, and background music that matches the scenes. This enables users, even without specialized knowledge, to quickly generate training videos that are easy to view and listen to and have rich content based on their own materials.

[0422] "Text-based materials" refers to materials that contain written data such as PDFs, Word files, and text files.

[0423] "Format" refers to the data format of the material, examples of which include PDF format and Word format.

[0424] "Text data" refers to textual information extracted from materials.

[0425] "Image data" refers to image information extracted from materials.

[0426] "Generative AI" is a system that uses artificial intelligence technology to analyze and refine materials.

[0427] "Natural language processing" is a technology that allows computers to understand, analyze, and generate human language.

[0428] "Narration" is audio generated from text.

[0429] "Subtitles" are texts that appear within a video.

[0430] "BGM" refers to the background music added to a video.

[0431] "User's correction instructions" refer to feedback and requests for improvements from users regarding the content of the automatically generated video.

[0432] "Interface means" refers to the means by which a user accesses and operates the system, and specifically includes a web interface and an application.

[0433] "Multilingual support" refers to the ability to generate narration and subtitles in multiple languages.

[0434] This invention is a system that allows users to easily import text-based materials and automatically generate high-quality training videos. This system uses a server, terminal, generation AI, text extraction module, image extraction module, natural language processing module, speech synthesis module, translation module, narration generation module, background music selection module, and user interface. The specific implementation of this invention is described below.

[0435] 1. User Incorporation of Materials:

[0436] Users use their devices to upload text-based materials (PDF, Word, text files, etc.) to the system, which receives the materials and sends them to the server via an HTTP POST request.

[0437] 2. Analysis of data by the server:

[0438] The server determines the format of the received document file and extracts the text data using the appropriate text extraction module (e.g., PDF parser, Docx parser), and also uses the image extraction module to obtain the image data. Through this process, the text and image information in the document is stored in a database.

[0439] 3. Brushing up materials with generative AI:

[0440] The server-based generative AI analyzes the document's text and image data. It uses natural language processing (NLP) algorithms to identify important keywords and phrases and organize the content of each page into sections. It also improves page layout and legibility based on basic principles of visual design. If necessary, the generative AI's image generation module generates missing images and charts and adds them to the document.

[0441] 4. Video generation and additional elements:

[0442] The server generates animated videos based on the generated text and images. Based on the generated materials, it uses a natural language processing module and a speech synthesis module to generate narration from the text. If multilingual support is required, it uses a translation module to translate and generate narration and subtitles. Furthermore, it uses a background music selection module to automatically select background music that matches the content of the video and add it to the video.

[0443] 5. Check and correct the video:

[0444] Users can view the generated videos on their devices and provide feedback and corrections through the user interface. The server then adjusts and corrects the video content based on the user's corrections. This allows users to easily create high-quality training videos.

[0445] Examples:

[0446] For example, if a user uploads "product training" materials, the AI ​​will visually arrange each product's features, usage instructions, and precautions in an easy-to-understand manner, adding diagrams and animations as needed. Narration and subtitles are also automatically generated, and appropriate background music is selected.

[0447] Example prompt sentence:

[0448] Example prompt 1: "Upload the PDF version of our Product Training Materials and generate a training video highlighting key points for each section."

[0449] Example prompt 2: "Based on a Word document, extract text data and create a narrated training video with a clear layout."

[0450] As described above, by linking various modules, this system is able to automatically generate high-quality training videos and meet user needs quickly and efficiently.

[0451] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0452] Step 1:

[0453] A user uses their own device to select text-based materials (PDF, Word, text files, etc.) and upload them to the system. The device then sends the selected materials to the server. The input is the material file selected by the user, and the output is the material file sent to the server.

[0454] Specific behavior:

[0455] The user clicks the "Upload" button and selects the file they want to upload. The device then sends the selected file to the server via an HTTP POST request.

[0456] Step 2:

[0457] The server determines the format of the received document file. It then calls the appropriate text extraction module, such as a PDF parser for PDF format or a Docx parser for Word format, to extract the text data. It also uses an image extraction module to obtain image data. The input is the document file sent to the server, and the output is the extracted text data and image data.

[0458] Specific behavior:

[0459] The server checks the extension of the received file and passes the file to the appropriate parser module. The text extraction module extracts the text information and stores it in a database. The image extraction module extracts the image data and stores it in a database as well.

[0460] Step 3:

[0461] The server's generation AI analyzes the extracted text data and image data. It uses natural language processing algorithms to identify keywords and important phrases and organize the content of each page into sections. It also refines the page layout and generates and adds missing images and charts to improve visibility. The input is the extracted text data and image data, and the output is improved document data.

[0462] Specific behavior:

[0463] The generative AI analyzes the text data, extracts key keywords and phrases, rearranges the layout of each page, and incorporates newly generated images and charts into the document as needed.

[0464] Step 4:

[0465] The server generates animated videos based on the reference data improved by the generative AI. The server generates narration text using a natural language processing module and generates audio files using a speech synthesis module. If multilingual support is required, the translation module translates and generates narration and subtitles. In addition, the background music selection module automatically selects appropriate background music and adds it to the video. The input is the improved reference data, and the output is the generated animated video.

[0466] Specific behavior:

[0467] The generative AI creates animated slides based on the text and images on each page. The narration text is generated using a natural language processing module and converted into an audio file using a speech synthesis module. Appropriate music is selected from the background music library and added to the video.

[0468] Step 5:

[0469] The user checks the generated animation on the device and inputs feedback and corrections using the user interface. The server adjusts and corrects the video content based on the user's instructions. The input is the generated animation video and the user's feedback, and the output is the final video that has been corrected and adjusted.

[0470] Specific behavior:

[0471] The user opens a preview of the video on their device and inputs corrections as they are made. The server updates the video content based on the user's feedback and generates the final version.

[0472] As described above, this system processes and analyzes input data at each processing step, automatically generating high-quality training videos that meet the user's needs.

[0473] (Application example 1)

[0474] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0475] In modern business and educational settings, training and instruction using text-based materials is common. However, these materials are generally visually unappealing and lack narration or video formats, resulting in low learning effectiveness. Furthermore, for many users, video editing and adding narration, which require specialized knowledge, are difficult and time-consuming. Furthermore, creating training videos tailored to specific tasks, such as multilingual support, operation manuals, and security training, requires additional specialized skills. Therefore, there is a need for a system that allows anyone to easily generate high-quality training videos and apply them to operation guides and security training.

[0476] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0477] In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, polishing the screen layout to make them easier to see and hear, and converting them into animated videos, a means for automatically generating narration, subtitles, and background music that matches the scenes, a means for converting the narration from text to audio, and a means for automatically selecting and generating visual decoration elements to be used in generating the videos. This allows users to simply upload text-based materials and automatically generate high-quality training videos that are visually appealing and include additional elements such as multilingual narration, subtitles, and appropriate background music. The system also provides an interface that allows users to review the generated videos and make necessary corrections, enabling them to make corrections according to specific tasks, making it possible to generate videos for specific purposes such as operation guides and security training.

[0478] "Text-based materials" refers to digital documents that primarily contain text, such as PDFs, Word files, and text files.

[0479] "Generative AI" refers to artificial intelligence that generates new content based on given data.

[0480] "Analysis" is the process of analyzing data and understanding and extracting its contents.

[0481] "Screen configuration" refers to the layout that visually displays information.

[0482] "Animated video" refers to video content that has movement and change.

[0483] "Narration" refers to audio that provides explanations and commentary in videos and presentations.

[0484] "Subtitles" are elements that display textual information in videos and presentations.

[0485] A "scene" refers to a specific situation or situation within a video.

[0486] "BGM" refers to the music played in the background of a video or presentation.

[0487] "Convert to speech" refers to the process of converting text information into audio data.

[0488] "Visual decoration" refers to visual information such as images or icons used to enhance visual appeal.

[0489] "User" means any individual or entity that uses the system or service.

[0490] An "interface" is the means by which a system and a user interact with each other to perform operations and exchange information.

[0491] "Modification instructions" refer to instructions for changes that the user makes to the generated video.

[0492] "Multilingual support" refers to having the functionality to support multiple languages.

[0493] An "operation guide" is a document or video that explains how to use a particular device or system.

[0494] "Security training" refers to training designed to provide knowledge about security measures and national security.

[0495] The present invention relates to a system for automatically generating attractive training videos for use as operational guides or security training for electronic payment services. Specific embodiments of this system are described below.

[0496] System program description

[0497] This system allows users to easily import text-based materials and automatically generate engaging training videos. It mainly consists of the following processing steps:

[0498] 1. Importing materials:

[0499] The server receives text-based materials such as PDFs, Word documents, and text files uploaded from users' devices. After receiving the files, the server extracts text and image data to analyze the contents of the files.

[0500] 2. Analysis of the material:

[0501] The server uses a text extraction module (e.g., Python's PIL library) to extract text information from the materials, and an image extraction module (e.g., PIL library) to obtain image data.

[0502] 3. Brushing up with generative AI:

[0503] The server's generative AI (e.g., GPT-3 or Stable Diffusion) analyzes the content of the document and refines the screen composition to make it easier to read and listen to. It improves the page layout based on basic principles of visual design and generates appropriate images and charts to add to the document.

[0504] 4. Video generation:

[0505] It automatically generates narration, subtitles, and background music to match the scenes, and generates animated videos. The narration is converted from text to speech using a natural language processing module (e.g., pyttsx3), and images are converted to video using MoviePy.

[0506] 5. Check and correct the video:

[0507] The user checks the generated video and inputs any necessary corrections. The server readjusts and corrects the video content based on the user's feedback and provides an optimized video.

[0508] This system allows users to significantly improve their work efficiency and easily create high-quality operation guides and security training videos.

[0509] Hardware and software used

[0510] Text and image data extraction:

[0511] Software used: Python library (PIL, FPDF)

[0512] Enhanced by generative AI:

[0513] AI model used: Generative AI (e.g., GPT-3, Stable Diffusion)

[0514] Software used: NLP library (spacy), image processing library (PIL)

[0515] Video Generation:

[0516] Tools used: MoviePy, pyttsx3

[0517] User Interface:

[0518] Devices used: smartphone, PC

[0519] Specific examples

[0520] For example, if you use security training materials for an electronic payment service:

[0521] 1. The user uploads "Security Guide.pdf".

[0522] 2. The server extracts important security measures from the guide.

[0523] 3. Use the extracted information to generate a video that visually explains security measures.

[0524] 4. The video is accompanied by narration and background music, and the user watches the training video.

[0525] 5. The user gives feedback saying, "I want more emphasis on password management," and the server makes the corrections.

[0526] Example of input prompt for generative AI model

[0527] Example prompt:

[0528] "Please extract important security measures from this document and create a layout that visually explains them. The contents of the document are as follows: (Text content)"

[0529] This makes it possible for anyone to easily generate high-quality, visually appealing training videos, contributing to training and improving security awareness among users of electronic payment services.

[0530] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0531] Step 1:

[0532] The terminal uploads the user's text-based document file (PDF, Word, text file, etc.). The input is the document file that the user specifies on the terminal, and the output is that file being transferred to the server. Specifically, the file is uploaded to the server by opening a file selection dialog on the terminal and clicking the send button.

[0533] Step 2:

[0534] The server analyzes the uploaded file. After receiving the file, it determines the file format and extracts the text and image data. The input is the uploaded file, and the output is the extracted text and image data. Specifically, it uses the Python PIL library to analyze and extract the text and images.

[0535] Step 3:

[0536] The server uses a generative AI model to analyze the content and structure of the document and refine it to make it easier to read and listen to. The input is extracted text data and image data, and the output is refined text and images. Specifically, generative AI (e.g., GPT-3, Stable Diffusion) is used to identify important keywords and phrases in the document and improve the page layout based on them.

[0537] Step 4:

[0538] The server uses generative AI to generate animated videos based on the content of the materials. It animates the information on each page along a timeline and uses a natural language processing module to generate narration from the text. The input is the polished text and images, and the output is an animated video file. Specifically, MoviePy and pyttsx3 are used to create the video and audio.

[0539] Step 5:

[0540] The server automatically generates narration, subtitles, and background music to match the scene. The input is the generated animation video file, and the output is the final video file containing narration, subtitles, and background music. Background music is automatically selected from a library, and subtitles are automatically generated to match the narration.

[0541] Step 6:

[0542] The terminal displays the generated video to the user and prompts them to check it. The user checks the generated video and inputs any necessary corrections. The input is a correction instruction from the user, and the output is the correction instruction being sent to the server. Specifically, the user inputs the correction content using the correction interface and clicks the send button.

[0543] Step 7:

[0544] The server readjusts and modifies the video content based on the user's correction instructions. The input is the user's correction instructions, and the output is the final modified video file. Specifically, it uses regeneration AI to correct the specified parts, regenerates narration and subtitles, and adds necessary visual decoration elements.

[0545] This series of processing steps enables users to easily generate high-quality operation guides and security training videos.

[0546] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0547] This invention relates to a system that allows users to easily import text-based materials, which are then analyzed by a generative AI to automatically generate high-quality training videos. In particular, by combining an emotion engine, this system is characterized by recognizing user emotions and enabling more adaptive content creation that reflects that feedback. Below, the program processing of this system is explained in natural language, and specific examples are provided to illustrate the mode of carrying out the invention.

[0548] 1. User Import of Materials

[0549] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system, which then receives the materials and sends them to the server.

[0550] 2. Analysis of data by the server

[0551] The server analyzes the received file and extracts text and image data. It identifies the file format (PDF, DOCX, TXT, etc.), extracts text information using a text extraction module, and obtains image data using an image extraction module.

[0552] 3. Brushing up materials with generative AI

[0553] The server-generated AI analyzes the content and structure of the document, identifies each page into sections and identifies important keywords and phrases, improves page layout based on basic principles of visual design to improve visibility, and adds appropriate AI-generated images and charts to the document as needed.

[0554] As a concrete example, if a user uploads "product training" materials, the generating AI will visually arrange each product's features, usage instructions, precautions, etc. in an easy-to-understand manner, adding diagrams and animations.

[0555] 4. Video generation and adding additional elements

[0556] The server uses generative AI to generate animated videos based on the content of the materials. The information on each page is animated along a timeline, and a natural language processing module is used to generate narration. If multilingual support is required, a translation module is used to generate narration. Subtitles are also automatically generated to match the narration, and background music appropriate to the video content is automatically selected from a library and added to the video.

[0557] For example, in the case of "product training" materials, the generative AI will visually represent the product's features, add appropriate narration for each point, and provide subtitles and background music appropriate to the content.

[0558] 5. Emotion Analysis Using an Emotion Engine

[0559] When a user watches a video on a device, the device collects the user's facial expression and voice data. This data is sent to a server, and the server's emotion engine recognizes the user's emotions in real time. The emotion engine uses, for example, facial recognition technology and voice analysis technology to determine whether the user is interested, confused, or expressing other specific emotions.

[0560] 6. Adaptive Content Adjustment

[0561] Based on the emotion engine's judgment, the server dynamically changes some or all elements in the video (such as narration, subtitles, background music, etc.) For example, if the server determines that the user is confused, it changes the narration to make it easier to understand and adds subtitles that re-emphasize important points.

[0562] 7. Check and correct the video

[0563] The user checks the generated video and inputs any necessary corrections. The server adjusts and corrects the video content based on the user's feedback. The emotion engine also provides automatic feedback according to changes in the user's emotions, allowing the video to be optimally adjusted.

[0564] Through the above processing steps, users can easily create more intuitive and effective training videos through a system that combines an emotion recognition engine.

[0565] The processing flow will be explained below.

[0566] This invention relates to a system that takes text-based materials and automatically generates high-quality training videos using generative AI and an emotion engine. Specific processing steps are explained below.

[0567] Step 1:

[0568] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system, which then receives the materials and sends them to the server.

[0569] Step 2:

[0570] The server analyzes the received document file and extracts text and image data. Specifically, it determines the file format and extracts the text using a text extraction engine if it is a PDF, or a different text extraction engine if it is a Word document, and performs the same process on image data.

[0571] Step 3:

[0572] The server analyzes the extracted text and images using generative AI, which identifies each page and section of the document and uses natural language processing techniques to extract key keywords and phrases.

[0573] Step 4:

[0574] Server-generated AI improves page layout: the layout engine follows basic principles of visual design to optimize text position, font, color, and image placement, and adds appropriate AI-generated images and charts where necessary.

[0575] Step 5:

[0576] The server uses generative AI to generate animated videos based on the content of the document, and an engine runs to animate the information on each page along a timeline, creating visually appealing videos.

[0577] Step 6:

[0578] The server automatically generates narration using a natural language processing module, text-to-speech synthesis technology, and, if multilingual support is required, a translation module is used to generate narration in each language.

[0579] Step 7:

[0580] The server automatically generates subtitles to match the narration. The subtitle generation engine matches the narration text with the timestamps to generate subtitle data, which is then incorporated into the video.

[0581] Step 8:

[0582] The server automatically selects the best background music from the library to fit the video content, and adds it to the video using the BGM addition engine. The most suitable music for each scene is selected and played as background music.

[0583] Step 9:

[0584] When a user watches a video, the device collects facial expression and voice data from the user. It captures real-time video and audio through a camera and microphone and sends the data to a server.

[0585] Step 10:

[0586] The server's emotion engine analyzes the user's facial expressions and voice data to recognize their emotions. An emotion recognition algorithm is used to identify whether the user is interested, confused, happy, etc.

[0587] Step 11:

[0588] The server dynamically adjusts the video's narration, subtitles, and background music based on the analysis results of the emotion engine. For example, if the server determines that the user is confused, it will change the narration to make it easier to understand and highlight the subtitles and narration.

[0589] Step 12:

[0590] The device displays the generated video to the user through a preview function, and the user can check the video and provide feedback if any corrections are required.

[0591] Step 13:

[0592] The server adjusts and modifies the video based on user feedback, and then uses related modules (generative AI, emotion engine, etc.) to improve the video content.

[0593] Step 14:

[0594] The device will provide the final video file to the user in a downloadable format, exporting the video in the specified format (e.g. MP4).

[0595] Through the above processing steps, users can easily create more adaptive and effective training videos through a system that combines an emotion recognition engine.

[0596] Example 2

[0597] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0598] Conventional training video generation systems lack sufficient visual and audio quality, requiring users to spend a lot of time manually making corrections. Furthermore, they lack the ability to dynamically adjust content based on the viewer's emotions, limiting the effectiveness of the content. Furthermore, their lack of multilingual support makes them difficult to use internationally. These issues need to be addressed.

[0599] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0600] In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, polishing the screen composition to make them easier to see and hear, and converting them into animated videos, a means for automatically generating narration, subtitles, and background music to match the scenes, an emotion analysis means for collecting user facial and voice data and recognizing emotions in real time, and a means for dynamically adjusting the video content based on the emotion analysis results. This not only improves the visual and auditory quality of the materials, but also enables adaptive adjustments based on the viewer's emotions in real time, resulting in the creation of more effective training videos. Furthermore, multilingual support for narration and subtitles facilitates international use.

[0601] "Text-based materials" are electronic documents that consist primarily of textual information and are provided in formats such as PDFs, Word documents, and text files.

[0602] "Generative AI" is a software system that uses artificial intelligence techniques to analyze data and automate complex tasks.

[0603] "Narration" refers to the commentary or explanation provided as audio within a video.

[0604] "Subtitles" are narrations and audio content displayed as textual information, in a format that can be visually understood by viewers.

[0605] "BGM that matches the scene" automatically selects and adds background music that matches each scene in the video.

[0606] "User facial expression and voice data" refers to data obtained from facial expressions and voice collected while a user is watching a video.

[0607] "Real-time emotion analysis" is a process that uses facial expression recognition and voice analysis technology to instantly analyze and determine a user's emotions.

[0608] "Dynamic adjustment of video content" refers to the operation of changing and optimizing elements such as video narration, subtitles, and background music in real time based on the results of sentiment analysis.

[0609] "Multilingual" refers to the ability of a system to function in multiple languages ​​and to translate and generate narration and subtitles into those languages.

[0610] An "interface means" is a software component that provides a mechanism for a user to perform operations and inputs through interaction with the system.

[0611] This invention is a system that allows users to easily import text-based materials and automatically generates high-quality training videos using generative AI. In particular, it combines an emotion engine to recognize user emotions in real time and enable more adaptive content creation that reflects feedback.

[0612] First, the user selects text-based materials (e.g., PDF, Word, text files, etc.) from their device and uploads them to the system. The device reads the materials from local storage and sends them to the server. Specifically, the user sends "Product Training Materials.pdf" to the server using the system's upload button.

[0613] Next, the server receives the uploaded file and identifies its format. It uses the server's text extraction module to extract text information and its image extraction module to obtain image data. For example, the server analyzes "Product Training Materials.pdf" and extracts the text and images for each page.

[0614] The server's generative AI then analyzes the content and structure of the materials. Generative AI is a software system that uses artificial intelligence technology to analyze data and automate complex tasks. It divides each page into sections and identifies important keywords and phrases. It improves the page layout based on basic principles of visual design to improve visibility. If necessary, it generates additional images and diagrams and incorporates them into the materials. As a specific example, the generative AI identifies and highlights "product features," "usage instructions," and "precautions" from the main text of "product training materials."

[0615] The server then begins generating videos based on the analyzed and polished materials. Animated slides are created using text and image data, and a natural language processing module generates narration. A translation module generates narration in multiple languages ​​and automatically generates subtitles to accompany the narration. For example, a slide showing "Product Features" is animated, and a narrator generates the following: "Product A is the latest model with the best performance on the market." Appropriate background music for the video is then selected from the library and added as "BGM_Upbeat.mp3."

[0616] When a user watches a video, the device uses a camera and microphone to collect the user's facial expression and voice data. This data is then sent to a server, where the server's emotion analysis engine uses facial expression and voice analysis technology to recognize the user's emotions in real time. For example, if the device streams camera data and detects the user's smile, the server determines that the user is in an "interested" state.

[0617] Based on the results of the sentiment analysis, the server dynamically adjusts the video content. If it determines that the user is confused, it changes the narration to make it easier to understand and adds additional subtitles. For example, if the user is determined to be in a "confused" state, it changes the narration to "What's important here is the safety features of product A. Let me explain in detail," and adds a subtitle "Safety feature details" with key points highlighted.

[0618] Finally, the user reviews the generated video and inputs any necessary corrections through the interface. The server receives the user's feedback and readjusts or corrects the video content. For example, if the user inputs a correction request saying "the narration is too fast," the server presents the user with a regenerated video with the narration speed adjusted.

[0619] Examples of prompt statements

[0620] User: "I've uploaded some product training materials. I'd like to see videos that visually explain the features and usage of each product."

[0621] Generative AI: "Analyzing your document. Identifying each section and identifying keywords. Improving page layout and adding appropriate images and charts."

[0622] Emotion Engine: "We detected that your audience was confused. We'll update the narration to make it clearer and add subtitles to re-emphasize key points."

[0623] Through the above process, users can easily create intuitive and effective training videos through a system that combines an emotion analysis engine.

[0624] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0625] Step 1:

[0626] User import of materials

[0627] A user selects text-based materials (e.g., PDF, Word, text files, etc.) on their device and uploads them to the system. The device reads the materials from local storage and sends them to the server. For example, the user selects "Product Training Materials.pdf" and clicks the system's upload button. The device reads the selected file and sends an HTTP POST request to the server. The input is the material file selected by the user, and the output is the material file sent to the server.

[0628] Step 2:

[0629] Server analysis of data

[0630] The server receives the document file sent from the terminal. It determines the file format (e.g., PDF, DOCX, TXT), extracts text information using the text extraction module, and obtains image data using the image extraction module. For example, the server receives "Product Training Materials.pdf" and uses a PDF parser to extract the text and images for each page. The input is the received document file, and the output is the extracted text data and image data.

[0631] Step 3:

[0632] Brushing up materials with generative AI

[0633] The server's generation AI analyzes the content and structure of the materials. It divides each page into sections and identifies important keywords and phrases. It improves the page layout based on basic principles of visual design to improve visibility. If necessary, it generates additional images and diagrams and incorporates them into the materials. Specifically, the generation AI identifies "product features," "usage instructions," and "precautions" from the main text of the "product training materials" and highlights this information. The input is the extracted text and image data, and the output is polished material data.

[0634] Step 4:

[0635] Video generation and adding additional elements

[0636] The server begins generating videos based on the analyzed and refined materials. Animated slides are created using text and image data, and a natural language processing module is used to generate narration. A translation module generates narration in multiple languages ​​as needed. Subtitles are automatically generated to match the narration, and appropriate background music is selected from a library and added to the video. For example, a "Product Features" slide is animated, and the narrator generates, "Product A is the latest model with the best performance on the market." Additionally, background music suitable for the video is selected from a library and added as "BGM_Upbeat.mp3." The input is the refined material data, and the output is the generated animated video.

[0637] Step 5:

[0638] Emotion analysis using an emotion engine

[0639] When a user watches a video, the device uses a camera and microphone to collect the user's facial and voice data. This data is sent to a server, where the server's emotion analysis engine uses facial and voice analysis technologies to recognize the user's emotions in real time. For example, if the device streams camera data and detects a user's smile, the server determines that the user is in an "interested" state. The input is the collected facial and voice data, and the output is emotion data determined in real time.

[0640] Step 6:

[0641] Adaptive content adjustment

[0642] The server dynamically adjusts elements in the video (such as narration, subtitles, and background music) based on the results of sentiment analysis. For example, if it determines that the user is confused, it changes the narration to something more understandable and adds subtitles that re-emphasize important points. For example, it changes the narration to "What's important here are the safety features of product A. Let me explain in detail," and adds a subtitle "Safety feature details" that highlights the important points. The input is the sentiment analysis results, and the output is the dynamically adjusted video content.

[0643] Step 7:

[0644] Check and correct the video

[0645] The user reviews the generated video and inputs any necessary corrections through the interface. The server receives the user's feedback and readjusts and corrects the video content. For example, if the user inputs a correction request saying "the narration is too fast," the server presents the user with a regenerated video with the narration speed adjusted. The input is the user's correction feedback, and the output is the corrected video.

[0646] Through the above processing steps, users can easily create intuitive and effective training videos through a system that combines an emotion analysis engine.

[0647] (Application example 2)

[0648] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0649] Conventional training video creation systems have difficulty adjusting content in real time to take user emotions into account, limiting their effectiveness in improving viewer comprehension. Furthermore, the generation of multilingual narration and subtitles is inefficient because it requires manual editing due to insufficient automation.

[0650] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, brushing up the screen composition to make it easy to see and hear, and converting it into an animated video, a means for automatically generating narration, subtitles, and background music that matches the scene, and a means for analyzing the user's emotions in real time using an emotion engine and dynamically adjusting the video content. This makes it possible to automatically generate high-quality training videos that adapt to the user's emotions and also easily support multiple languages.

[0651] "Text-based materials" refers to documents that primarily contain text, such as PDFs, Word documents, and text files.

[0652] "Generative AI" refers to a system that uses artificial intelligence technology to analyze input data and generate new content.

[0653] "Brushing up" refers to improving input materials to make them visually easier to read and understand.

[0654] An "animated video" refers to a video that expresses movement by displaying multiple still images in sequence.

[0655] "Narration" refers to audio commentary that accompanies a video scene.

[0656] "Subtitles" refers to the display of video content and audio as text.

[0657] A "scene" refers to a specific situation or moment within a video.

[0658] "BGM" is an abbreviation for background music and refers to the music played in the background of a video.

[0659] "Emotion engine" refers to artificial intelligence technology that recognizes and analyzes a user's emotional state from their facial expressions and voice.

[0660] "Analyzing user emotions in real time" means determining the user's emotions from their facial expressions and voice at the moment they are watching a video.

[0661] "Dynamic adjustment of video content" refers to changing elements of a video, such as narration and subtitles, in real time based on the results of user sentiment analysis.

[0662] "Smart devices" refers to mobile information terminals with computer functions, such as smartphones and tablets.

[0663] "Collecting facial and voice data" refers to recording the user's facial movements and vocalizations using the smart device's camera and microphone.

[0664] This invention is a system that allows users to easily import text-based materials using a smart device, analyze them with a generative AI, and automatically generate high-quality training videos using an emotion engine. Specific embodiments of this system are described below.

[0665] 1. User Import of Materials

[0666] Users select text-based materials from their smart devices (such as smartphones or tablets) and upload them to the system. During this process, users press a button to open a file selection dialog, and the selected file (PDF, Word, or text file) is sent to the server. The hardware used here is a smart device, and the software is React Native and Amazon S3.

[0667] 2. Analysis of the data

[0668] The server analyzes the document files received from Amazon S3 and extracts text and image data. It identifies the file format (PDF, DOCX, TXT) and uses a specific text extraction module to extract character information. This process includes basic principles of visual design, allowing the generative AI to analyze the content and structure of the document and improve the page layout.

[0669] 3. Video generation and adding additional elements

[0670] The server uses generative AI to generate animated videos based on the content of the materials. It automatically generates narration, subtitles, and background music to match the scenes, improving visibility and comprehension. If multilingual support is required, a translation module is used to generate narration and automatically generate subtitles. The software used includes generative AI models and natural language processing technology.

[0671] 4. Emotion Analysis Using an Emotion Engine

[0672] When a user watches a video on a device, the smart device's camera and microphone are used to collect facial expressions and voice data, and emotions are analyzed in real time. The emotion engine uses facial recognition technology (OpenCV) and voice analysis technology (Google Cloud Speech-to-Text) to recognize the user's emotions.

[0673] 5. Real-time content adjustment

[0674] The server dynamically adjusts the video's narration and subtitles based on the analysis results of the emotion engine, allowing it to add additional narration explanations if the user is confused and display subtitles that emphasize important points.

[0675] 6. Check and correct the video

[0676] The user can review the generated video and input any necessary corrections. The server then adjusts the video content based on the user's feedback and provides the optimal training video.

[0677] For example, if a user uploads "product training" materials, the generative AI will visually arrange each product's features, usage instructions, and precautions in an easy-to-understand manner, automatically generating diagrams, narration, and subtitles. Furthermore, the emotion engine will analyze the user's facial expressions and voice, and if it determines that their understanding is incomplete, it will insert additional explanations into the video.

[0678] Example prompt sentence:

[0679] Please select the file you want to upload.

[0680] "Data analysis completed."

[0681] "We recognize the emotions of our users."

[0682] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0683] Step 1:

[0684] A user uses a smart device to select text-based materials and upload them to the system. Specifically, the user presses the file selection button and selects either a PDF, Word, or text file from the file selection dialog on the smart device. The selected file is sent to Amazon S3 by the device. The input is the file selected by the user, and the output is the file uploaded to the server.

[0685] Step 2:

[0686] The server receives documents uploaded to Amazon S3 and determines the file format (PDF, DOCX, TXT). It uses a text extraction module to extract text information, and if necessary, an image extraction module to obtain image data. The input is the uploaded document file, and the output is the extracted text data and image data. Specifically, it analyzes the file format and runs the extraction module corresponding to each format.

[0687] Step 3:

[0688] The server's generative AI analyzes the content and structure based on the analyzed material. It identifies important keywords and phrases and improves the page layout based on visual design. If necessary, it adds images and charts generated by the generative AI to the material. The input is the extracted text data and image data, and the output is a polished page layout and additional visual elements. Specifically, the generative AI analyzes the text and reconstructs the page layout.

[0689] Step 4:

[0690] The server uses generative AI to generate animated videos based on the content of the materials. The information on each page is animated along a timeline, and a natural language processing module is used to generate narration. Furthermore, if multilingual support is required, a translation module is used to generate narration and automatically generate subtitles. The input is the polished materials and layout information, and the output is an animated video with narration. Specific operations include natural language processing and timeline setting.

[0691] Step 5:

[0692] When a user watches a video on a device, facial expression and voice data are collected using the device's camera and microphone. This data is sent to a server in real time and analyzed by an emotion engine. The input is the user's facial expression and voice data, and the output is analyzed emotion data. Specifically, the system collects camera video and voice data in real time and performs facial recognition and voice analysis.

[0693] Step 6:

[0694] The server dynamically adjusts the video content based on the results of emotion analysis. If it determines that the user is confused, it adds supplementary narration explanations and displays subtitles that re-emphasize important points. The input is the analyzed emotion data, and the output is the adjusted video content. Specific operations include updating the narration content and subtitles.

[0695] Step 7:

[0696] The user reviews the generated video and inputs corrections as necessary. The server receives the user's feedback and readjusts the video content. The input is the user's feedback, and the output is the final corrected video. Specific operations include collecting feedback through the interface and re-editing the video.

[0697] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0698] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0699] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0700] [Third embodiment]

[0701] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0702] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0703] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0704] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0705] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0706] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0707] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0708] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0709] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0710] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0711] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0712] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0713] This invention relates to a system that allows users to easily import text-based materials and automatically generate engaging training videos. Below, we will explain the program processing of this system in natural language and provide examples to illustrate how to implement the invention.

[0714] 1. User Import of Materials

[0715] Users upload text-based materials (PDF, Word, text files, etc.) to the system from their own devices, which then receive the materials and send them to the server.

[0716] 2. Analysis of data by the server

[0717] The server analyzes the received file and extracts text and image data. It identifies the file format and uses a text extraction module to extract character information. It also uses an image extraction module to obtain image data.

[0718] 3. Brushing up materials with generative AI

[0719] The server's generated AI analyzes the content and structure of the document. It identifies each page by section and identifies important keywords and phrases. It then improves the page layout based on basic principles of visual design to improve visibility. If necessary, the AI ​​generates appropriate images and charts and adds them to the document.

[0720] As a concrete example, if a user uploads "product training" materials, the generating AI will visually arrange each product's features, usage instructions, precautions, etc. in an easy-to-understand manner, adding diagrams and animations.

[0721] 4. Video generation and adding additional elements

[0722] The server uses generative AI to generate animated videos based on the content of the materials. It animates the information on each page along a timeline and uses a natural language processing module to generate narration from the text. If multilingual support is required, a translation module is used to generate narration. Subtitles are also automatically generated to match the narration. Furthermore, background music appropriate to the video content is automatically selected from a library and added to the video.

[0723] For example, in the case of "product training" materials, the generative AI visually represents the product's features, adds appropriate narration at each point, and provides subtitles and background music appropriate to the content.

[0724] 5. Check and correct the video

[0725] Users can check the generated videos on their devices and input any necessary corrections. The server then adjusts and corrects the video content based on the user's feedback. This allows users to easily create high-quality training videos.

[0726] The above processing steps realize a system that allows anyone to quickly generate training videos based on their own materials that are easy to watch and listen to and have rich content.

[0727] The processing flow will be explained below.

[0728] Step 1:

[0729] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system. The device receives the file specified by the user and sends it to the server.

[0730] Step 2:

[0731] The server analyzes the received file and extracts text and image data. Specifically, it identifies the file format (PDF, DOCX, TXT, etc.), extracts text information using a text extraction module, and obtains image data using an image extraction module.

[0732] Step 3:

[0733] The server analyzes the extracted text and images using generative AI, which identifies each page of the document by section and identifies important keywords and phrases.

[0734] Step 4:

[0735] The server-generated AI improves the page layout, rearranging information based on basic principles of visual design to improve readability, and adding appropriate AI-generated images and charts where necessary.

[0736] Step 5:

[0737] The server uses generative AI to generate animated videos based on the content of the documents, and displays the information on each page as an animation along a timeline.

[0738] Step 6:

[0739] The server generates narration using a natural language processing module, converts text information into speech, and generates narration using a translation module if multilingual support is required.

[0740] Step 7:

[0741] The server automatically generates subtitles to match the narration, creating text information to be displayed on the screen in sync with the narration.

[0742] Step 8:

[0743] The server automatically selects background music from its library that matches the content of the video and adds it to the video, selecting the best background music for each scene.

[0744] Step 9:

[0745] The terminal displays the generated video to the user and provides a preview function for the user to verify.

[0746] Step 10:

[0747] The user can review the generated video and input any necessary corrections. The user can use a designated input form to specify corrections to specific text, narration, subtitles, and animation.

[0748] Step 11:

[0749] The server executes the relevant modules again based on the user's feedback, adjusts and modifies the video content, and performs necessary processing according to the feedback content.

[0750] Step 12:

[0751] The device will present the final video file to the user and provide a download link, and the video will be exported in the specified format (e.g., MP4).

[0752] The above processing steps create a system that allows anyone to easily create high-quality training videos.

[0753] Example 1

[0754] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0755] Traditionally, creating training videos based on text-based materials required a great deal of time and effort, and often required specialized knowledge. Furthermore, creating narration, subtitles, and selecting background music all required manual work, which added to the burden. This made it difficult for anyone to easily create high-quality training videos. Furthermore, creating multilingual narration and subtitles required advanced technology and time.

[0756] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0757] In this invention, the server includes a means for importing text-based materials, a means for determining the format of the imported materials and extracting text data and image data, a means for analyzing the extracted materials using a generation AI, organizing the content by section, and brushing up the layout to make it easy to view, and a means for automatically generating narration, subtitles, and background music that matches the scenes. This enables users, even without specialized knowledge, to quickly generate training videos that are easy to view and listen to and have rich content based on their own materials.

[0758] "Text-based materials" refers to materials that contain written data such as PDFs, Word files, and text files.

[0759] "Format" refers to the data format of the material, examples of which include PDF format and Word format.

[0760] "Text data" refers to textual information extracted from materials.

[0761] "Image data" refers to image information extracted from materials.

[0762] "Generative AI" is a system that uses artificial intelligence technology to analyze and refine materials.

[0763] "Natural language processing" is a technology that allows computers to understand, analyze, and generate human language.

[0764] "Narration" is audio generated from text.

[0765] "Subtitles" are texts that appear within a video.

[0766] "BGM" refers to the background music added to a video.

[0767] "User's correction instructions" refer to feedback and requests for improvements from users regarding the content of the automatically generated video.

[0768] "Interface means" refers to the means by which a user accesses and operates the system, and specifically includes a web interface and an application.

[0769] "Multilingual support" refers to the ability to generate narration and subtitles in multiple languages.

[0770] This invention is a system that allows users to easily import text-based materials and automatically generate high-quality training videos. This system uses a server, terminal, generation AI, text extraction module, image extraction module, natural language processing module, speech synthesis module, translation module, narration generation module, background music selection module, and user interface. The specific implementation of this invention is described below.

[0771] 1. User Incorporation of Materials:

[0772] Users use their devices to upload text-based materials (PDF, Word, text files, etc.) to the system, which receives the materials and sends them to the server via an HTTP POST request.

[0773] 2. Analysis of data by the server:

[0774] The server determines the format of the received document file and extracts the text data using the appropriate text extraction module (e.g., PDF parser, Docx parser), and also uses the image extraction module to obtain the image data. Through this process, the text and image information in the document is stored in a database.

[0775] 3. Brushing up materials with generative AI:

[0776] The server-based generative AI analyzes the document's text and image data. It uses natural language processing (NLP) algorithms to identify important keywords and phrases and organize the content of each page into sections. It also improves page layout and legibility based on basic principles of visual design. If necessary, the generative AI's image generation module generates missing images and charts and adds them to the document.

[0777] 4. Video generation and additional elements:

[0778] The server generates animated videos based on the generated text and images. Based on the generated materials, it uses a natural language processing module and a speech synthesis module to generate narration from the text. If multilingual support is required, it uses a translation module to translate and generate narration and subtitles. Furthermore, it uses a background music selection module to automatically select background music that matches the content of the video and add it to the video.

[0779] 5. Check and correct the video:

[0780] Users can view the generated videos on their devices and provide feedback and corrections through the user interface. The server then adjusts and corrects the video content based on the user's corrections. This allows users to easily create high-quality training videos.

[0781] Examples:

[0782] For example, if a user uploads "product training" materials, the AI ​​will visually arrange each product's features, usage instructions, and precautions in an easy-to-understand manner, adding diagrams and animations as needed. Narration and subtitles are also automatically generated, and appropriate background music is selected.

[0783] Example prompt sentence:

[0784] Example prompt 1: "Upload the PDF version of our Product Training Materials and generate a training video highlighting key points for each section."

[0785] Example prompt 2: "Based on a Word document, extract text data and create a narrated training video with a clear layout."

[0786] As described above, by linking various modules, this system is able to automatically generate high-quality training videos and meet user needs quickly and efficiently.

[0787] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0788] Step 1:

[0789] A user uses their own device to select text-based materials (PDF, Word, text files, etc.) and upload them to the system. The device then sends the selected materials to the server. The input is the material file selected by the user, and the output is the material file sent to the server.

[0790] Specific behavior:

[0791] The user clicks the "Upload" button and selects the file they want to upload. The device then sends the selected file to the server via an HTTP POST request.

[0792] Step 2:

[0793] The server determines the format of the received document file. It then calls the appropriate text extraction module, such as a PDF parser for PDF format or a Docx parser for Word format, to extract the text data. It also uses an image extraction module to obtain image data. The input is the document file sent to the server, and the output is the extracted text data and image data.

[0794] Specific behavior:

[0795] The server checks the extension of the received file and passes the file to the appropriate parser module. The text extraction module extracts the text information and stores it in a database. The image extraction module extracts the image data and stores it in a database as well.

[0796] Step 3:

[0797] The server's generation AI analyzes the extracted text data and image data. It uses natural language processing algorithms to identify keywords and important phrases and organize the content of each page into sections. It also refines the page layout and generates and adds missing images and charts to improve visibility. The input is the extracted text data and image data, and the output is improved document data.

[0798] Specific behavior:

[0799] The generative AI analyzes the text data, extracts key keywords and phrases, rearranges the layout of each page, and incorporates newly generated images and charts into the document as needed.

[0800] Step 4:

[0801] The server generates animated videos based on the reference data improved by the generative AI. The server generates narration text using a natural language processing module and generates audio files using a speech synthesis module. If multilingual support is required, the translation module translates and generates narration and subtitles. In addition, the background music selection module automatically selects appropriate background music and adds it to the video. The input is the improved reference data, and the output is the generated animated video.

[0802] Specific behavior:

[0803] The generative AI creates animated slides based on the text and images on each page. The narration text is generated using a natural language processing module and converted into an audio file using a speech synthesis module. Appropriate music is selected from the background music library and added to the video.

[0804] Step 5:

[0805] The user checks the generated animation on the device and inputs feedback and corrections using the user interface. The server adjusts and corrects the video content based on the user's instructions. The input is the generated animation video and the user's feedback, and the output is the final video that has been corrected and adjusted.

[0806] Specific behavior:

[0807] The user opens a preview of the video on their device and inputs corrections as they are made. The server updates the video content based on the user's feedback and generates the final version.

[0808] As described above, this system processes and analyzes input data at each processing step, automatically generating high-quality training videos that meet the user's needs.

[0809] (Application example 1)

[0810] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0811] In modern business and educational settings, training and instruction using text-based materials is common. However, these materials are generally visually unappealing and lack narration or video formats, resulting in low learning effectiveness. Furthermore, for many users, video editing and adding narration, which require specialized knowledge, are difficult and time-consuming. Furthermore, creating training videos tailored to specific tasks, such as multilingual support, operation manuals, and security training, requires additional specialized skills. Therefore, there is a need for a system that allows anyone to easily generate high-quality training videos and apply them to operation guides and security training.

[0812] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0813] In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, polishing the screen layout to make them easier to see and hear, and converting them into animated videos, a means for automatically generating narration, subtitles, and background music that matches the scenes, a means for converting the narration from text to audio, and a means for automatically selecting and generating visual decoration elements to be used in generating the videos. This allows users to simply upload text-based materials and automatically generate high-quality training videos that are visually appealing and include additional elements such as multilingual narration, subtitles, and appropriate background music. The system also provides an interface that allows users to review the generated videos and make necessary corrections, enabling them to make corrections according to specific tasks, making it possible to generate videos for specific purposes such as operation guides and security training.

[0814] "Text-based materials" refers to digital documents that primarily contain text, such as PDFs, Word files, and text files.

[0815] "Generative AI" refers to artificial intelligence that generates new content based on given data.

[0816] "Analysis" is the process of analyzing data and understanding and extracting its contents.

[0817] "Screen configuration" refers to the layout that visually displays information.

[0818] "Animated video" refers to video content that has movement and change.

[0819] "Narration" refers to audio that provides explanations and commentary in videos and presentations.

[0820] "Subtitles" are elements that display textual information in videos and presentations.

[0821] A "scene" refers to a specific situation or situation within a video.

[0822] "BGM" refers to the music played in the background of a video or presentation.

[0823] "Convert to speech" refers to the process of converting text information into audio data.

[0824] "Visual decoration" refers to visual information such as images or icons used to enhance visual appeal.

[0825] "User" means any individual or entity that uses the system or service.

[0826] An "interface" is the means by which a system and a user interact with each other to perform operations and exchange information.

[0827] "Modification instructions" refer to instructions for changes that the user makes to the generated video.

[0828] "Multilingual support" refers to having the functionality to support multiple languages.

[0829] An "operation guide" is a document or video that explains how to use a particular device or system.

[0830] "Security training" refers to training designed to provide knowledge about security measures and national security.

[0831] The present invention relates to a system for automatically generating attractive training videos for use as operational guides or security training for electronic payment services. Specific embodiments of this system are described below.

[0832] System program description

[0833] This system allows users to easily import text-based materials and automatically generate engaging training videos. It mainly consists of the following processing steps:

[0834] 1. Importing materials:

[0835] The server receives text-based materials such as PDFs, Word documents, and text files uploaded from users' devices. After receiving the files, the server extracts text and image data to analyze the contents of the files.

[0836] 2. Analysis of the material:

[0837] The server uses a text extraction module (e.g., Python's PIL library) to extract text information from the materials, and an image extraction module (e.g., PIL library) to obtain image data.

[0838] 3. Brushing up with generative AI:

[0839] The server's generative AI (e.g., GPT-3 or Stable Diffusion) analyzes the content of the document and refines the screen composition to make it easier to read and listen to. It improves the page layout based on basic principles of visual design and generates appropriate images and charts to add to the document.

[0840] 4. Video generation:

[0841] It automatically generates narration, subtitles, and background music to match the scenes, and generates animated videos. The narration is converted from text to speech using a natural language processing module (e.g., pyttsx3), and images are converted to video using MoviePy.

[0842] 5. Check and correct the video:

[0843] The user checks the generated video and inputs any necessary corrections. The server readjusts and corrects the video content based on the user's feedback and provides an optimized video.

[0844] This system allows users to significantly improve their work efficiency and easily create high-quality operation guides and security training videos.

[0845] Hardware and software used

[0846] Text and image data extraction:

[0847] Software used: Python library (PIL, FPDF)

[0848] Enhanced by generative AI:

[0849] AI model used: Generative AI (e.g., GPT-3, Stable Diffusion)

[0850] Software used: NLP library (spacy), image processing library (PIL)

[0851] Video Generation:

[0852] Tools used: MoviePy, pyttsx3

[0853] User Interface:

[0854] Devices used: smartphone, PC

[0855] Specific examples

[0856] For example, if you use security training materials for an electronic payment service:

[0857] 1. The user uploads "Security Guide.pdf".

[0858] 2. The server extracts important security measures from the guide.

[0859] 3. Use the extracted information to generate a video that visually explains security measures.

[0860] 4. The video is accompanied by narration and background music, and the user watches the training video.

[0861] 5. The user gives feedback saying, "I want more emphasis on password management," and the server makes the corrections.

[0862] Example of input prompt for generative AI model

[0863] Example prompt:

[0864] "Please extract important security measures from this document and create a layout that visually explains them. The contents of the document are as follows: (Text content)"

[0865] This makes it possible for anyone to easily generate high-quality, visually appealing training videos, contributing to training and improving security awareness among users of electronic payment services.

[0866] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0867] Step 1:

[0868] The terminal uploads the user's text-based document file (PDF, Word, text file, etc.). The input is the document file that the user specifies on the terminal, and the output is that file being transferred to the server. Specifically, the file is uploaded to the server by opening a file selection dialog on the terminal and clicking the send button.

[0869] Step 2:

[0870] The server analyzes the uploaded file. After receiving the file, it determines the file format and extracts the text and image data. The input is the uploaded file, and the output is the extracted text and image data. Specifically, it uses the Python PIL library to analyze and extract the text and images.

[0871] Step 3:

[0872] The server uses a generative AI model to analyze the content and structure of the document and refine it to make it easier to read and listen to. The input is extracted text data and image data, and the output is refined text and images. Specifically, generative AI (e.g., GPT-3, Stable Diffusion) is used to identify important keywords and phrases in the document and improve the page layout based on them.

[0873] Step 4:

[0874] The server uses generative AI to generate animated videos based on the content of the materials. It animates the information on each page along a timeline and uses a natural language processing module to generate narration from the text. The input is the polished text and images, and the output is an animated video file. Specifically, MoviePy and pyttsx3 are used to create the video and audio.

[0875] Step 5:

[0876] The server automatically generates narration, subtitles, and background music to match the scene. The input is the generated animation video file, and the output is the final video file containing narration, subtitles, and background music. Background music is automatically selected from a library, and subtitles are automatically generated to match the narration.

[0877] Step 6:

[0878] The terminal displays the generated video to the user and prompts them to check it. The user checks the generated video and inputs any necessary corrections. The input is a correction instruction from the user, and the output is the correction instruction being sent to the server. Specifically, the user inputs the correction content using the correction interface and clicks the send button.

[0879] Step 7:

[0880] The server readjusts and modifies the video content based on the user's correction instructions. The input is the user's correction instructions, and the output is the final modified video file. Specifically, it uses regeneration AI to correct the specified parts, regenerates narration and subtitles, and adds necessary visual decoration elements.

[0881] This series of processing steps enables users to easily generate high-quality operation guides and security training videos.

[0882] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0883] This invention relates to a system that allows users to easily import text-based materials, which are then analyzed by a generative AI to automatically generate high-quality training videos. In particular, by combining an emotion engine, this system is characterized by recognizing user emotions and enabling more adaptive content creation that reflects that feedback. Below, the program processing of this system is explained in natural language, and specific examples are provided to illustrate the mode of carrying out the invention.

[0884] 1. User Import of Materials

[0885] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system, which then receives the materials and sends them to the server.

[0886] 2. Analysis of data by the server

[0887] The server analyzes the received file and extracts text and image data. It identifies the file format (PDF, DOCX, TXT, etc.), extracts text information using a text extraction module, and obtains image data using an image extraction module.

[0888] 3. Brushing up materials with generative AI

[0889] The server-generated AI analyzes the content and structure of the document, identifies each page into sections and identifies important keywords and phrases, improves page layout based on basic principles of visual design to improve visibility, and adds appropriate AI-generated images and charts to the document as needed.

[0890] As a concrete example, if a user uploads "product training" materials, the generating AI will visually arrange each product's features, usage instructions, precautions, etc. in an easy-to-understand manner, adding diagrams and animations.

[0891] 4. Video generation and adding additional elements

[0892] The server uses generative AI to generate animated videos based on the content of the materials. The information on each page is animated along a timeline, and a natural language processing module is used to generate narration. If multilingual support is required, a translation module is used to generate narration. Subtitles are also automatically generated to match the narration, and background music appropriate to the video content is automatically selected from a library and added to the video.

[0893] For example, in the case of "product training" materials, the generative AI will visually represent the product's features, add appropriate narration for each point, and provide subtitles and background music appropriate to the content.

[0894] 5. Emotion Analysis Using an Emotion Engine

[0895] When a user watches a video on a device, the device collects the user's facial expression and voice data. This data is sent to a server, and the server's emotion engine recognizes the user's emotions in real time. The emotion engine uses, for example, facial recognition technology and voice analysis technology to determine whether the user is interested, confused, or expressing other specific emotions.

[0896] 6. Adaptive Content Adjustment

[0897] Based on the emotion engine's judgment, the server dynamically changes some or all elements in the video (such as narration, subtitles, background music, etc.) For example, if the server determines that the user is confused, it changes the narration to make it easier to understand and adds subtitles that re-emphasize important points.

[0898] 7. Check and correct the video

[0899] The user checks the generated video and inputs any necessary corrections. The server adjusts and corrects the video content based on the user's feedback. The emotion engine also provides automatic feedback according to changes in the user's emotions, allowing the video to be optimally adjusted.

[0900] Through the above processing steps, users can easily create more intuitive and effective training videos through a system that combines an emotion recognition engine.

[0901] The processing flow will be explained below.

[0902] This invention relates to a system that takes text-based materials and automatically generates high-quality training videos using generative AI and an emotion engine. Specific processing steps are explained below.

[0903] Step 1:

[0904] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system, which then receives the materials and sends them to the server.

[0905] Step 2:

[0906] The server analyzes the received document file and extracts text and image data. Specifically, it determines the file format and extracts the text using a text extraction engine if it is a PDF, or a different text extraction engine if it is a Word document, and performs the same process on image data.

[0907] Step 3:

[0908] The server analyzes the extracted text and images using generative AI, which identifies each page and section of the document and uses natural language processing techniques to extract key keywords and phrases.

[0909] Step 4:

[0910] Server-generated AI improves page layout: the layout engine follows basic principles of visual design to optimize text position, font, color, and image placement, and adds appropriate AI-generated images and charts where necessary.

[0911] Step 5:

[0912] The server uses generative AI to generate animated videos based on the content of the document, and an engine runs to animate the information on each page along a timeline, creating visually appealing videos.

[0913] Step 6:

[0914] The server automatically generates narration using a natural language processing module, text-to-speech synthesis technology, and, if multilingual support is required, a translation module is used to generate narration in each language.

[0915] Step 7:

[0916] The server automatically generates subtitles to match the narration. The subtitle generation engine matches the narration text with the timestamps to generate subtitle data, which is then incorporated into the video.

[0917] Step 8:

[0918] The server automatically selects the best background music from the library to fit the video content, and adds it to the video using the BGM addition engine. The most suitable music for each scene is selected and played as background music.

[0919] Step 9:

[0920] When a user watches a video, the device collects facial expression and voice data from the user. It captures real-time video and audio through a camera and microphone and sends the data to a server.

[0921] Step 10:

[0922] The server's emotion engine analyzes the user's facial expressions and voice data to recognize their emotions. An emotion recognition algorithm is used to identify whether the user is interested, confused, happy, etc.

[0923] Step 11:

[0924] The server dynamically adjusts the video's narration, subtitles, and background music based on the analysis results of the emotion engine. For example, if the server determines that the user is confused, it will change the narration to make it easier to understand and highlight the subtitles and narration.

[0925] Step 12:

[0926] The device displays the generated video to the user through a preview function, and the user can check the video and provide feedback if any corrections are required.

[0927] Step 13:

[0928] The server adjusts and modifies the video based on user feedback, and then uses related modules (generative AI, emotion engine, etc.) to improve the video content.

[0929] Step 14:

[0930] The device will provide the final video file to the user in a downloadable format, exporting the video in the specified format (e.g. MP4).

[0931] Through the above processing steps, users can easily create more adaptive and effective training videos through a system that combines an emotion recognition engine.

[0932] Example 2

[0933] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0934] Conventional training video generation systems lack sufficient visual and audio quality, requiring users to spend a lot of time manually making corrections. Furthermore, they lack the ability to dynamically adjust content based on the viewer's emotions, limiting the effectiveness of the content. Furthermore, their lack of multilingual support makes them difficult to use internationally. These issues need to be addressed.

[0935] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0936] In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, polishing the screen composition to make them easier to see and hear, and converting them into animated videos, a means for automatically generating narration, subtitles, and background music to match the scenes, an emotion analysis means for collecting user facial and voice data and recognizing emotions in real time, and a means for dynamically adjusting the video content based on the emotion analysis results. This not only improves the visual and auditory quality of the materials, but also enables adaptive adjustments based on the viewer's emotions in real time, resulting in the creation of more effective training videos. Furthermore, multilingual support for narration and subtitles facilitates international use.

[0937] "Text-based materials" are electronic documents that consist primarily of textual information and are provided in formats such as PDFs, Word documents, and text files.

[0938] "Generative AI" is a software system that uses artificial intelligence techniques to analyze data and automate complex tasks.

[0939] "Narration" refers to the commentary or explanation provided as audio within a video.

[0940] "Subtitles" are narrations and audio content displayed as textual information, in a format that can be visually understood by viewers.

[0941] "BGM that matches the scene" automatically selects and adds background music that matches each scene in the video.

[0942] "User facial expression and voice data" refers to data obtained from facial expressions and voice collected while a user is watching a video.

[0943] "Real-time emotion analysis" is a process that uses facial expression recognition and voice analysis technology to instantly analyze and determine a user's emotions.

[0944] "Dynamic adjustment of video content" refers to the operation of changing and optimizing elements such as video narration, subtitles, and background music in real time based on the results of sentiment analysis.

[0945] "Multilingual" refers to the ability of a system to function in multiple languages ​​and to translate and generate narration and subtitles into those languages.

[0946] An "interface means" is a software component that provides a mechanism for a user to perform operations and inputs through interaction with the system.

[0947] This invention is a system that allows users to easily import text-based materials and automatically generates high-quality training videos using generative AI. In particular, it combines an emotion engine to recognize user emotions in real time and enable more adaptive content creation that reflects feedback.

[0948] First, the user selects text-based materials (e.g., PDF, Word, text files, etc.) from their device and uploads them to the system. The device reads the materials from local storage and sends them to the server. Specifically, the user sends "Product Training Materials.pdf" to the server using the system's upload button.

[0949] Next, the server receives the uploaded file and identifies its format. It uses the server's text extraction module to extract text information and its image extraction module to obtain image data. For example, the server analyzes "Product Training Materials.pdf" and extracts the text and images for each page.

[0950] The server's generative AI then analyzes the content and structure of the materials. Generative AI is a software system that uses artificial intelligence technology to analyze data and automate complex tasks. It divides each page into sections and identifies important keywords and phrases. It improves the page layout based on basic principles of visual design to improve visibility. If necessary, it generates additional images and diagrams and incorporates them into the materials. As a specific example, the generative AI identifies and highlights "product features," "usage instructions," and "precautions" from the main text of "product training materials."

[0951] The server then begins generating videos based on the analyzed and polished materials. Animated slides are created using text and image data, and a natural language processing module generates narration. A translation module generates narration in multiple languages ​​and automatically generates subtitles to accompany the narration. For example, a slide showing "Product Features" is animated, and a narrator generates the following: "Product A is the latest model with the best performance on the market." Appropriate background music for the video is then selected from the library and added as "BGM_Upbeat.mp3."

[0952] When a user watches a video, the device uses a camera and microphone to collect the user's facial expression and voice data. This data is then sent to a server, where the server's emotion analysis engine uses facial expression and voice analysis technology to recognize the user's emotions in real time. For example, if the device streams camera data and detects the user's smile, the server determines that the user is in an "interested" state.

[0953] Based on the results of the sentiment analysis, the server dynamically adjusts the video content. If it determines that the user is confused, it changes the narration to make it easier to understand and adds additional subtitles. For example, if the user is determined to be in a "confused" state, it changes the narration to "What's important here is the safety features of product A. Let me explain in detail," and adds a subtitle "Safety feature details" with key points highlighted.

[0954] Finally, the user reviews the generated video and inputs any necessary corrections through the interface. The server receives the user's feedback and readjusts or corrects the video content. For example, if the user inputs a correction request saying "the narration is too fast," the server presents the user with a regenerated video with the narration speed adjusted.

[0955] Examples of prompt statements

[0956] User: "I've uploaded some product training materials. I'd like to see videos that visually explain the features and usage of each product."

[0957] Generative AI: "Analyzing your document. Identifying each section and identifying keywords. Improving page layout and adding appropriate images and charts."

[0958] Emotion Engine: "We detected that your audience was confused. We'll update the narration to make it clearer and add subtitles to re-emphasize key points."

[0959] Through the above process, users can easily create intuitive and effective training videos through a system that combines an emotion analysis engine.

[0960] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0961] Step 1:

[0962] User import of materials

[0963] A user selects text-based materials (e.g., PDF, Word, text files, etc.) on their device and uploads them to the system. The device reads the materials from local storage and sends them to the server. For example, the user selects "Product Training Materials.pdf" and clicks the system's upload button. The device reads the selected file and sends an HTTP POST request to the server. The input is the material file selected by the user, and the output is the material file sent to the server.

[0964] Step 2:

[0965] Server analysis of data

[0966] The server receives the document file sent from the terminal. It determines the file format (e.g., PDF, DOCX, TXT), extracts text information using the text extraction module, and obtains image data using the image extraction module. For example, the server receives "Product Training Materials.pdf" and uses a PDF parser to extract the text and images for each page. The input is the received document file, and the output is the extracted text data and image data.

[0967] Step 3:

[0968] Brushing up materials with generative AI

[0969] The server's generation AI analyzes the content and structure of the materials. It divides each page into sections and identifies important keywords and phrases. It improves the page layout based on basic principles of visual design to improve visibility. If necessary, it generates additional images and diagrams and incorporates them into the materials. Specifically, the generation AI identifies "product features," "usage instructions," and "precautions" from the main text of the "product training materials" and highlights this information. The input is the extracted text and image data, and the output is polished material data.

[0970] Step 4:

[0971] Video generation and adding additional elements

[0972] The server begins generating videos based on the analyzed and refined materials. Animated slides are created using text and image data, and a natural language processing module is used to generate narration. A translation module generates narration in multiple languages ​​as needed. Subtitles are automatically generated to match the narration, and appropriate background music is selected from a library and added to the video. For example, a "Product Features" slide is animated, and the narrator generates, "Product A is the latest model with the best performance on the market." Additionally, background music suitable for the video is selected from a library and added as "BGM_Upbeat.mp3." The input is the refined material data, and the output is the generated animated video.

[0973] Step 5:

[0974] Emotion analysis using an emotion engine

[0975] When a user watches a video, the device uses a camera and microphone to collect the user's facial and voice data. This data is sent to a server, where the server's emotion analysis engine uses facial and voice analysis technologies to recognize the user's emotions in real time. For example, if the device streams camera data and detects a user's smile, the server determines that the user is in an "interested" state. The input is the collected facial and voice data, and the output is emotion data determined in real time.

[0976] Step 6:

[0977] Adaptive content adjustment

[0978] The server dynamically adjusts elements in the video (such as narration, subtitles, and background music) based on the results of sentiment analysis. For example, if it determines that the user is confused, it changes the narration to something more understandable and adds subtitles that re-emphasize important points. For example, it changes the narration to "What's important here are the safety features of product A. Let me explain in detail," and adds a subtitle "Safety feature details" that highlights the important points. The input is the sentiment analysis results, and the output is the dynamically adjusted video content.

[0979] Step 7:

[0980] Check and correct the video

[0981] The user reviews the generated video and inputs any necessary corrections through the interface. The server receives the user's feedback and readjusts and corrects the video content. For example, if the user inputs a correction request saying "the narration is too fast," the server presents the user with a regenerated video with the narration speed adjusted. The input is the user's correction feedback, and the output is the corrected video.

[0982] Through the above processing steps, users can easily create intuitive and effective training videos through a system that combines an emotion analysis engine.

[0983] (Application example 2)

[0984] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0985] Conventional training video creation systems have difficulty adjusting content in real time to take user emotions into account, limiting their effectiveness in improving viewer comprehension. Furthermore, the generation of multilingual narration and subtitles is inefficient because it requires manual editing due to insufficient automation.

[0986] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, brushing up the screen composition to make it easy to see and hear, and converting it into an animated video, a means for automatically generating narration, subtitles, and background music that matches the scene, and a means for analyzing the user's emotions in real time using an emotion engine and dynamically adjusting the video content. This makes it possible to automatically generate high-quality training videos that adapt to the user's emotions and also easily support multiple languages.

[0987] "Text-based materials" refers to documents that primarily contain text, such as PDFs, Word documents, and text files.

[0988] "Generative AI" refers to a system that uses artificial intelligence technology to analyze input data and generate new content.

[0989] "Brushing up" refers to improving input materials to make them visually easier to read and understand.

[0990] An "animated video" refers to a video that expresses movement by displaying multiple still images in sequence.

[0991] "Narration" refers to audio commentary that accompanies a video scene.

[0992] "Subtitles" refers to the display of video content and audio as text.

[0993] A "scene" refers to a specific situation or moment within a video.

[0994] "BGM" is an abbreviation for background music and refers to the music played in the background of a video.

[0995] "Emotion engine" refers to artificial intelligence technology that recognizes and analyzes a user's emotional state from their facial expressions and voice.

[0996] "Analyzing user emotions in real time" means determining the user's emotions from their facial expressions and voice at the moment they are watching a video.

[0997] "Dynamic adjustment of video content" refers to changing elements of a video, such as narration and subtitles, in real time based on the results of user sentiment analysis.

[0998] "Smart devices" refers to mobile information terminals with computer functions, such as smartphones and tablets.

[0999] "Collecting facial and voice data" refers to recording the user's facial movements and vocalizations using the smart device's camera and microphone.

[1000] This invention is a system that allows users to easily import text-based materials using a smart device, analyze them with a generative AI, and automatically generate high-quality training videos using an emotion engine. Specific embodiments of this system are described below.

[1001] 1. User Import of Materials

[1002] Users select text-based materials from their smart devices (such as smartphones or tablets) and upload them to the system. During this process, users press a button to open a file selection dialog, and the selected file (PDF, Word, or text file) is sent to the server. The hardware used here is a smart device, and the software is React Native and Amazon S3.

[1003] 2. Analysis of the data

[1004] The server analyzes the document files received from Amazon S3 and extracts text and image data. It identifies the file format (PDF, DOCX, TXT) and uses a specific text extraction module to extract character information. This process includes basic principles of visual design, allowing the generative AI to analyze the content and structure of the document and improve the page layout.

[1005] 3. Video generation and adding additional elements

[1006] The server uses generative AI to generate animated videos based on the content of the materials. It automatically generates narration, subtitles, and background music to match the scenes, improving visibility and comprehension. If multilingual support is required, a translation module is used to generate narration and automatically generate subtitles. The software used includes generative AI models and natural language processing technology.

[1007] 4. Emotion Analysis Using an Emotion Engine

[1008] When a user watches a video on a device, the smart device's camera and microphone are used to collect facial expressions and voice data, and emotions are analyzed in real time. The emotion engine uses facial recognition technology (OpenCV) and voice analysis technology (Google Cloud Speech-to-Text) to recognize the user's emotions.

[1009] 5. Real-time content adjustment

[1010] The server dynamically adjusts the video's narration and subtitles based on the analysis results of the emotion engine, allowing it to add additional narration explanations if the user is confused and display subtitles that emphasize important points.

[1011] 6. Check and correct the video

[1012] The user can review the generated video and input any necessary corrections. The server then adjusts the video content based on the user's feedback and provides the optimal training video.

[1013] For example, if a user uploads "product training" materials, the generative AI will visually arrange each product's features, usage instructions, and precautions in an easy-to-understand manner, automatically generating diagrams, narration, and subtitles. Furthermore, the emotion engine will analyze the user's facial expressions and voice, and if it determines that their understanding is incomplete, it will insert additional explanations into the video.

[1014] Example prompt sentence:

[1015] Please select the file you want to upload.

[1016] "Data analysis completed."

[1017] "We recognize the emotions of our users."

[1018] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1019] Step 1:

[1020] A user uses a smart device to select text-based materials and upload them to the system. Specifically, the user presses the file selection button and selects either a PDF, Word, or text file from the file selection dialog on the smart device. The selected file is sent to Amazon S3 by the device. The input is the file selected by the user, and the output is the file uploaded to the server.

[1021] Step 2:

[1022] The server receives documents uploaded to Amazon S3 and determines the file format (PDF, DOCX, TXT). It uses a text extraction module to extract text information, and if necessary, an image extraction module to obtain image data. The input is the uploaded document file, and the output is the extracted text data and image data. Specifically, it analyzes the file format and runs the extraction module corresponding to each format.

[1023] Step 3:

[1024] The server's generative AI analyzes the content and structure based on the analyzed material. It identifies important keywords and phrases and improves the page layout based on visual design. If necessary, it adds images and charts generated by the generative AI to the material. The input is the extracted text data and image data, and the output is a polished page layout and additional visual elements. Specifically, the generative AI analyzes the text and reconstructs the page layout.

[1025] Step 4:

[1026] The server uses generative AI to generate animated videos based on the content of the materials. The information on each page is animated along a timeline, and a natural language processing module is used to generate narration. Furthermore, if multilingual support is required, a translation module is used to generate narration and automatically generate subtitles. The input is the polished materials and layout information, and the output is an animated video with narration. Specific operations include natural language processing and timeline setting.

[1027] Step 5:

[1028] When a user watches a video on a device, facial expression and voice data are collected using the device's camera and microphone. This data is sent to a server in real time and analyzed by an emotion engine. The input is the user's facial expression and voice data, and the output is analyzed emotion data. Specifically, the system collects camera video and voice data in real time and performs facial recognition and voice analysis.

[1029] Step 6:

[1030] The server dynamically adjusts the video content based on the results of emotion analysis. If it determines that the user is confused, it adds supplementary narration explanations and displays subtitles that re-emphasize important points. The input is the analyzed emotion data, and the output is the adjusted video content. Specific operations include updating the narration content and subtitles.

[1031] Step 7:

[1032] The user reviews the generated video and inputs corrections as necessary. The server receives the user's feedback and readjusts the video content. The input is the user's feedback, and the output is the final corrected video. Specific operations include collecting feedback through the interface and re-editing the video.

[1033] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1034] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1035] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1036] [Fourth embodiment]

[1037] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1038] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1039] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1040] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1041] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1042] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1043] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1044] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1045] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1046] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1047] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1048] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1049] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1050] This invention relates to a system that allows users to easily import text-based materials and automatically generate engaging training videos. Below, we will explain the program processing of this system in natural language and provide examples to illustrate how to implement the invention.

[1051] 1. User Import of Materials

[1052] Users upload text-based materials (PDF, Word, text files, etc.) to the system from their own devices, which then receive the materials and send them to the server.

[1053] 2. Analysis of data by the server

[1054] The server analyzes the received file and extracts text and image data. It identifies the file format and uses a text extraction module to extract character information. It also uses an image extraction module to obtain image data.

[1055] 3. Brushing up materials with generative AI

[1056] The server's generated AI analyzes the content and structure of the document. It identifies each page by section and identifies important keywords and phrases. It then improves the page layout based on basic principles of visual design to improve visibility. If necessary, the AI ​​generates appropriate images and charts and adds them to the document.

[1057] As a concrete example, if a user uploads "product training" materials, the generating AI will visually arrange each product's features, usage instructions, precautions, etc. in an easy-to-understand manner, adding diagrams and animations.

[1058] 4. Video generation and adding additional elements

[1059] The server uses generative AI to generate animated videos based on the content of the materials. It animates the information on each page along a timeline and uses a natural language processing module to generate narration from the text. If multilingual support is required, a translation module is used to generate narration. Subtitles are also automatically generated to match the narration. Furthermore, background music appropriate to the video content is automatically selected from a library and added to the video.

[1060] For example, in the case of "product training" materials, the generative AI visually represents the product's features, adds appropriate narration at each point, and provides subtitles and background music appropriate to the content.

[1061] 5. Check and correct the video

[1062] Users can check the generated videos on their devices and input any necessary corrections. The server then adjusts and corrects the video content based on the user's feedback. This allows users to easily create high-quality training videos.

[1063] The above processing steps realize a system that allows anyone to quickly generate training videos based on their own materials that are easy to watch and listen to and have rich content.

[1064] The processing flow will be explained below.

[1065] Step 1:

[1066] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system. The device receives the file specified by the user and sends it to the server.

[1067] Step 2:

[1068] The server analyzes the received file and extracts text and image data. Specifically, it identifies the file format (PDF, DOCX, TXT, etc.), extracts text information using a text extraction module, and obtains image data using an image extraction module.

[1069] Step 3:

[1070] The server analyzes the extracted text and images using generative AI, which identifies each page of the document by section and identifies important keywords and phrases.

[1071] Step 4:

[1072] The server-generated AI improves the page layout, rearranging information based on basic principles of visual design to improve readability, and adding appropriate AI-generated images and charts where necessary.

[1073] Step 5:

[1074] The server uses generative AI to generate animated videos based on the content of the documents, and displays the information on each page as an animation along a timeline.

[1075] Step 6:

[1076] The server generates narration using a natural language processing module, converts text information into speech, and generates narration using a translation module if multilingual support is required.

[1077] Step 7:

[1078] The server automatically generates subtitles to match the narration, creating text information to be displayed on the screen in sync with the narration.

[1079] Step 8:

[1080] The server automatically selects background music from its library that matches the content of the video and adds it to the video, selecting the best background music for each scene.

[1081] Step 9:

[1082] The terminal displays the generated video to the user and provides a preview function for the user to verify.

[1083] Step 10:

[1084] The user can review the generated video and input any necessary corrections. The user can use a designated input form to specify corrections to specific text, narration, subtitles, and animation.

[1085] Step 11:

[1086] The server executes the relevant modules again based on the user's feedback, adjusts and modifies the video content, and performs necessary processing according to the feedback content.

[1087] Step 12:

[1088] The device will present the final video file to the user and provide a download link, and the video will be exported in the specified format (e.g., MP4).

[1089] The above processing steps create a system that allows anyone to easily create high-quality training videos.

[1090] Example 1

[1091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1092] Traditionally, creating training videos based on text-based materials required a great deal of time and effort, and often required specialized knowledge. Furthermore, creating narration, subtitles, and selecting background music all required manual work, which added to the burden. This made it difficult for anyone to easily create high-quality training videos. Furthermore, creating multilingual narration and subtitles required advanced technology and time.

[1093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1094] In this invention, the server includes a means for importing text-based materials, a means for determining the format of the imported materials and extracting text data and image data, a means for analyzing the extracted materials using a generation AI, organizing the content by section, and brushing up the layout to make it easy to view, and a means for automatically generating narration, subtitles, and background music that matches the scenes. This enables users, even without specialized knowledge, to quickly generate training videos that are easy to view and listen to and have rich content based on their own materials.

[1095] "Text-based materials" refers to materials that contain written data such as PDFs, Word files, and text files.

[1096] "Format" refers to the data format of the material, examples of which include PDF format and Word format.

[1097] "Text data" refers to textual information extracted from materials.

[1098] "Image data" refers to image information extracted from materials.

[1099] "Generative AI" is a system that uses artificial intelligence technology to analyze and refine materials.

[1100] "Natural language processing" is a technology that allows computers to understand, analyze, and generate human language.

[1101] "Narration" is audio generated from text.

[1102] "Subtitles" are texts that appear within a video.

[1103] "BGM" refers to the background music added to a video.

[1104] "User's correction instructions" refer to feedback and requests for improvements from users regarding the content of the automatically generated video.

[1105] "Interface means" refers to the means by which a user accesses and operates the system, and specifically includes a web interface and an application.

[1106] "Multilingual support" refers to the ability to generate narration and subtitles in multiple languages.

[1107] This invention is a system that allows users to easily import text-based materials and automatically generate high-quality training videos. This system uses a server, terminal, generation AI, text extraction module, image extraction module, natural language processing module, speech synthesis module, translation module, narration generation module, background music selection module, and user interface. The specific implementation of this invention is described below.

[1108] 1. User Incorporation of Materials:

[1109] Users use their devices to upload text-based materials (PDF, Word, text files, etc.) to the system, which receives the materials and sends them to the server via an HTTP POST request.

[1110] 2. Analysis of data by the server:

[1111] The server determines the format of the received document file and extracts the text data using the appropriate text extraction module (e.g., PDF parser, Docx parser), and also uses the image extraction module to obtain the image data. Through this process, the text and image information in the document is stored in a database.

[1112] 3. Brushing up materials with generative AI:

[1113] The server-based generative AI analyzes the document's text and image data. It uses natural language processing (NLP) algorithms to identify important keywords and phrases and organize the content of each page into sections. It also improves page layout and legibility based on basic principles of visual design. If necessary, the generative AI's image generation module generates missing images and charts and adds them to the document.

[1114] 4. Video generation and additional elements:

[1115] The server generates animated videos based on the generated text and images. Based on the generated materials, it uses a natural language processing module and a speech synthesis module to generate narration from the text. If multilingual support is required, it uses a translation module to translate and generate narration and subtitles. Furthermore, it uses a background music selection module to automatically select background music that matches the content of the video and add it to the video.

[1116] 5. Check and correct the video:

[1117] Users can view the generated videos on their devices and provide feedback and corrections through the user interface. The server then adjusts and corrects the video content based on the user's corrections. This allows users to easily create high-quality training videos.

[1118] Examples:

[1119] For example, if a user uploads "product training" materials, the AI ​​will visually arrange each product's features, usage instructions, and precautions in an easy-to-understand manner, adding diagrams and animations as needed. Narration and subtitles are also automatically generated, and appropriate background music is selected.

[1120] Example prompt sentence:

[1121] Example prompt 1: "Upload the PDF version of our Product Training Materials and generate a training video highlighting key points for each section."

[1122] Example prompt 2: "Based on a Word document, extract text data and create a narrated training video with a clear layout."

[1123] As described above, by linking various modules, this system is able to automatically generate high-quality training videos and meet user needs quickly and efficiently.

[1124] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1125] Step 1:

[1126] A user uses their own device to select text-based materials (PDF, Word, text files, etc.) and upload them to the system. The device then sends the selected materials to the server. The input is the material file selected by the user, and the output is the material file sent to the server.

[1127] Specific behavior:

[1128] The user clicks the "Upload" button and selects the file they want to upload. The device then sends the selected file to the server via an HTTP POST request.

[1129] Step 2:

[1130] The server determines the format of the received document file. It then calls the appropriate text extraction module, such as a PDF parser for PDF format or a Docx parser for Word format, to extract the text data. It also uses an image extraction module to obtain image data. The input is the document file sent to the server, and the output is the extracted text data and image data.

[1131] Specific behavior:

[1132] The server checks the extension of the received file and passes the file to the appropriate parser module. The text extraction module extracts the text information and stores it in a database. The image extraction module extracts the image data and stores it in a database as well.

[1133] Step 3:

[1134] The server's generation AI analyzes the extracted text data and image data. It uses natural language processing algorithms to identify keywords and important phrases and organize the content of each page into sections. It also refines the page layout and generates and adds missing images and charts to improve visibility. The input is the extracted text data and image data, and the output is improved document data.

[1135] Specific behavior:

[1136] The generative AI analyzes the text data, extracts key keywords and phrases, rearranges the layout of each page, and incorporates newly generated images and charts into the document as needed.

[1137] Step 4:

[1138] The server generates animated videos based on the reference data improved by the generative AI. The server generates narration text using a natural language processing module and generates audio files using a speech synthesis module. If multilingual support is required, the translation module translates and generates narration and subtitles. In addition, the background music selection module automatically selects appropriate background music and adds it to the video. The input is the improved reference data, and the output is the generated animated video.

[1139] Specific behavior:

[1140] The generative AI creates animated slides based on the text and images on each page. The narration text is generated using a natural language processing module and converted into an audio file using a speech synthesis module. Appropriate music is selected from the background music library and added to the video.

[1141] Step 5:

[1142] The user checks the generated animation on the device and inputs feedback and corrections using the user interface. The server adjusts and corrects the video content based on the user's instructions. The input is the generated animation video and the user's feedback, and the output is the final video that has been corrected and adjusted.

[1143] Specific behavior:

[1144] The user opens a preview of the video on their device and inputs corrections as they are made. The server updates the video content based on the user's feedback and generates the final version.

[1145] As described above, this system processes and analyzes input data at each processing step, automatically generating high-quality training videos that meet the user's needs.

[1146] (Application example 1)

[1147] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1148] In modern business and educational settings, training and instruction using text-based materials is common. However, these materials are generally visually unappealing and lack narration or video formats, resulting in low learning effectiveness. Furthermore, for many users, video editing and adding narration, which require specialized knowledge, are difficult and time-consuming. Furthermore, creating training videos tailored to specific tasks, such as multilingual support, operation manuals, and security training, requires additional specialized skills. Therefore, there is a need for a system that allows anyone to easily generate high-quality training videos and apply them to operation guides and security training.

[1149] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1150] In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, polishing the screen layout to make them easier to see and hear, and converting them into animated videos, a means for automatically generating narration, subtitles, and background music that matches the scenes, a means for converting the narration from text to audio, and a means for automatically selecting and generating visual decoration elements to be used in generating the videos. This allows users to simply upload text-based materials and automatically generate high-quality training videos that are visually appealing and include additional elements such as multilingual narration, subtitles, and appropriate background music. The system also provides an interface that allows users to review the generated videos and make necessary corrections, enabling them to make corrections according to specific tasks, making it possible to generate videos for specific purposes such as operation guides and security training.

[1151] "Text-based materials" refers to digital documents that primarily contain text, such as PDFs, Word files, and text files.

[1152] "Generative AI" refers to artificial intelligence that generates new content based on given data.

[1153] "Analysis" is the process of analyzing data and understanding and extracting its contents.

[1154] "Screen configuration" refers to the layout that visually displays information.

[1155] "Animated video" refers to video content that has movement and change.

[1156] "Narration" refers to audio that provides explanations and commentary in videos and presentations.

[1157] "Subtitles" are elements that display textual information in videos and presentations.

[1158] A "scene" refers to a specific situation or situation within a video.

[1159] "BGM" refers to the music played in the background of a video or presentation.

[1160] "Convert to speech" refers to the process of converting text information into audio data.

[1161] "Visual decoration" refers to visual information such as images or icons used to enhance visual appeal.

[1162] "User" means any individual or entity that uses the system or service.

[1163] An "interface" is the means by which a system and a user interact with each other to perform operations and exchange information.

[1164] "Modification instructions" refer to instructions for changes that the user makes to the generated video.

[1165] "Multilingual support" refers to having the functionality to support multiple languages.

[1166] An "operation guide" is a document or video that explains how to use a particular device or system.

[1167] "Security training" refers to training designed to provide knowledge about security measures and national security.

[1168] The present invention relates to a system for automatically generating attractive training videos for use as operational guides or security training for electronic payment services. Specific embodiments of this system are described below.

[1169] System program description

[1170] This system allows users to easily import text-based materials and automatically generate engaging training videos. It mainly consists of the following processing steps:

[1171] 1. Importing materials:

[1172] The server receives text-based materials such as PDFs, Word documents, and text files uploaded from users' devices. After receiving the files, the server extracts text and image data to analyze the contents of the files.

[1173] 2. Analysis of the material:

[1174] The server uses a text extraction module (e.g., Python's PIL library) to extract text information from the materials, and an image extraction module (e.g., PIL library) to obtain image data.

[1175] 3. Brushing up with generative AI:

[1176] The server's generative AI (e.g., GPT-3 or Stable Diffusion) analyzes the content of the document and refines the screen composition to make it easier to read and listen to. It improves the page layout based on basic principles of visual design and generates appropriate images and charts to add to the document.

[1177] 4. Video generation:

[1178] It automatically generates narration, subtitles, and background music to match the scenes, and generates animated videos. The narration is converted from text to speech using a natural language processing module (e.g., pyttsx3), and images are converted to video using MoviePy.

[1179] 5. Check and correct the video:

[1180] The user checks the generated video and inputs any necessary corrections. The server readjusts and corrects the video content based on the user's feedback and provides an optimized video.

[1181] This system allows users to significantly improve their work efficiency and easily create high-quality operation guides and security training videos.

[1182] Hardware and software used

[1183] Text and image data extraction:

[1184] Software used: Python library (PIL, FPDF)

[1185] Enhanced by generative AI:

[1186] AI model used: Generative AI (e.g., GPT-3, Stable Diffusion)

[1187] Software used: NLP library (spacy), image processing library (PIL)

[1188] Video Generation:

[1189] Tools used: MoviePy, pyttsx3

[1190] User Interface:

[1191] Devices used: smartphone, PC

[1192] Specific examples

[1193] For example, if you use security training materials for an electronic payment service:

[1194] 1. The user uploads "Security Guide.pdf".

[1195] 2. The server extracts important security measures from the guide.

[1196] 3. Use the extracted information to generate a video that visually explains security measures.

[1197] 4. The video is accompanied by narration and background music, and the user watches the training video.

[1198] 5. The user gives feedback saying, "I want more emphasis on password management," and the server makes the corrections.

[1199] Example of input prompt for generative AI model

[1200] Example prompt:

[1201] "Please extract important security measures from this document and create a layout that visually explains them. The contents of the document are as follows: (Text content)"

[1202] This makes it possible for anyone to easily generate high-quality, visually appealing training videos, contributing to training and improving security awareness among users of electronic payment services.

[1203] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1204] Step 1:

[1205] The terminal uploads the user's text-based document file (PDF, Word, text file, etc.). The input is the document file that the user specifies on the terminal, and the output is that file being transferred to the server. Specifically, the file is uploaded to the server by opening a file selection dialog on the terminal and clicking the send button.

[1206] Step 2:

[1207] The server analyzes the uploaded file. After receiving the file, it determines the file format and extracts the text and image data. The input is the uploaded file, and the output is the extracted text and image data. Specifically, it uses the Python PIL library to analyze and extract the text and images.

[1208] Step 3:

[1209] The server uses a generative AI model to analyze the content and structure of the document and refine it to make it easier to read and listen to. The input is extracted text data and image data, and the output is refined text and images. Specifically, generative AI (e.g., GPT-3, Stable Diffusion) is used to identify important keywords and phrases in the document and improve the page layout based on them.

[1210] Step 4:

[1211] The server uses generative AI to generate animated videos based on the content of the materials. It animates the information on each page along a timeline and uses a natural language processing module to generate narration from the text. The input is the polished text and images, and the output is an animated video file. Specifically, MoviePy and pyttsx3 are used to create the video and audio.

[1212] Step 5:

[1213] The server automatically generates narration, subtitles, and background music to match the scene. The input is the generated animation video file, and the output is the final video file containing narration, subtitles, and background music. Background music is automatically selected from a library, and subtitles are automatically generated to match the narration.

[1214] Step 6:

[1215] The terminal displays the generated video to the user and prompts them to check it. The user checks the generated video and inputs any necessary corrections. The input is a correction instruction from the user, and the output is the correction instruction being sent to the server. Specifically, the user inputs the correction content using the correction interface and clicks the send button.

[1216] Step 7:

[1217] The server readjusts and modifies the video content based on the user's correction instructions. The input is the user's correction instructions, and the output is the final modified video file. Specifically, it uses regeneration AI to correct the specified parts, regenerates narration and subtitles, and adds necessary visual decoration elements.

[1218] This series of processing steps enables users to easily generate high-quality operation guides and security training videos.

[1219] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1220] This invention relates to a system that allows users to easily import text-based materials, which are then analyzed by a generative AI to automatically generate high-quality training videos. In particular, by combining an emotion engine, this system is characterized by recognizing user emotions and enabling more adaptive content creation that reflects that feedback. Below, the program processing of this system is explained in natural language, and specific examples are provided to illustrate the mode of carrying out the invention.

[1221] 1. User Import of Materials

[1222] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system, which then receives the materials and sends them to the server.

[1223] 2. Analysis of data by the server

[1224] The server analyzes the received file and extracts text and image data. It identifies the file format (PDF, DOCX, TXT, etc.), extracts text information using a text extraction module, and obtains image data using an image extraction module.

[1225] 3. Brushing up materials with generative AI

[1226] The server-generated AI analyzes the content and structure of the document, identifies each page into sections and identifies important keywords and phrases, improves page layout based on basic principles of visual design to improve visibility, and adds appropriate AI-generated images and charts to the document as needed.

[1227] As a concrete example, if a user uploads "product training" materials, the generating AI will visually arrange each product's features, usage instructions, precautions, etc. in an easy-to-understand manner, adding diagrams and animations.

[1228] 4. Video generation and adding additional elements

[1229] The server uses generative AI to generate animated videos based on the content of the materials. The information on each page is animated along a timeline, and a natural language processing module is used to generate narration. If multilingual support is required, a translation module is used to generate narration. Subtitles are also automatically generated to match the narration, and background music appropriate to the video content is automatically selected from a library and added to the video.

[1230] For example, in the case of "product training" materials, the generative AI will visually represent the product's features, add appropriate narration for each point, and provide subtitles and background music appropriate to the content.

[1231] 5. Emotion Analysis Using an Emotion Engine

[1232] When a user watches a video on a device, the device collects the user's facial expression and voice data. This data is sent to a server, and the server's emotion engine recognizes the user's emotions in real time. The emotion engine uses, for example, facial recognition technology and voice analysis technology to determine whether the user is interested, confused, or expressing other specific emotions.

[1233] 6. Adaptive Content Adjustment

[1234] Based on the emotion engine's judgment, the server dynamically changes some or all elements in the video (such as narration, subtitles, background music, etc.) For example, if the server determines that the user is confused, it changes the narration to make it easier to understand and adds subtitles that re-emphasize important points.

[1235] 7. Check and correct the video

[1236] The user checks the generated video and inputs any necessary corrections. The server adjusts and corrects the video content based on the user's feedback. The emotion engine also provides automatic feedback according to changes in the user's emotions, allowing the video to be optimally adjusted.

[1237] Through the above processing steps, users can easily create more intuitive and effective training videos through a system that combines an emotion recognition engine.

[1238] The processing flow will be explained below.

[1239] This invention relates to a system that takes text-based materials and automatically generates high-quality training videos using generative AI and an emotion engine. Specific processing steps are explained below.

[1240] Step 1:

[1241] Users select text-based materials (PDF, Word, text files, etc.) from their own devices and upload them to the system, which then receives the materials and sends them to the server.

[1242] Step 2:

[1243] The server analyzes the received document file and extracts text and image data. Specifically, it determines the file format and extracts the text using a text extraction engine if it is a PDF, or a different text extraction engine if it is a Word document, and performs the same process on image data.

[1244] Step 3:

[1245] The server analyzes the extracted text and images using generative AI, which identifies each page and section of the document and uses natural language processing techniques to extract key keywords and phrases.

[1246] Step 4:

[1247] Server-generated AI improves page layout: the layout engine follows basic principles of visual design to optimize text position, font, color, and image placement, and adds appropriate AI-generated images and charts where necessary.

[1248] Step 5:

[1249] The server uses generative AI to generate animated videos based on the content of the document, and an engine runs to animate the information on each page along a timeline, creating visually appealing videos.

[1250] Step 6:

[1251] The server automatically generates narration using a natural language processing module, text-to-speech synthesis technology, and, if multilingual support is required, a translation module is used to generate narration in each language.

[1252] Step 7:

[1253] The server automatically generates subtitles to match the narration. The subtitle generation engine matches the narration text with the timestamps to generate subtitle data, which is then incorporated into the video.

[1254] Step 8:

[1255] The server automatically selects the best background music from the library to fit the video content, and adds it to the video using the BGM addition engine. The most suitable music for each scene is selected and played as background music.

[1256] Step 9:

[1257] When a user watches a video, the device collects facial expression and voice data from the user. It captures real-time video and audio through a camera and microphone and sends the data to a server.

[1258] Step 10:

[1259] The server's emotion engine analyzes the user's facial expressions and voice data to recognize their emotions. An emotion recognition algorithm is used to identify whether the user is interested, confused, happy, etc.

[1260] Step 11:

[1261] The server dynamically adjusts the video's narration, subtitles, and background music based on the analysis results of the emotion engine. For example, if the server determines that the user is confused, it will change the narration to make it easier to understand and highlight the subtitles and narration.

[1262] Step 12:

[1263] The device displays the generated video to the user through a preview function, and the user can check the video and provide feedback if any corrections are required.

[1264] Step 13:

[1265] The server adjusts and modifies the video based on user feedback, and then uses related modules (generative AI, emotion engine, etc.) to improve the video content.

[1266] Step 14:

[1267] The device will provide the final video file to the user in a downloadable format, exporting the video in the specified format (e.g. MP4).

[1268] Through the above processing steps, users can easily create more adaptive and effective training videos through a system that combines an emotion recognition engine.

[1269] Example 2

[1270] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1271] Conventional training video generation systems lack sufficient visual and audio quality, requiring users to spend a lot of time manually making corrections. Furthermore, they lack the ability to dynamically adjust content based on the viewer's emotions, limiting the effectiveness of the content. Furthermore, their lack of multilingual support makes them difficult to use internationally. These issues need to be addressed.

[1272] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1273] In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, polishing the screen composition to make them easier to see and hear, and converting them into animated videos, a means for automatically generating narration, subtitles, and background music to match the scenes, an emotion analysis means for collecting user facial and voice data and recognizing emotions in real time, and a means for dynamically adjusting the video content based on the emotion analysis results. This not only improves the visual and auditory quality of the materials, but also enables adaptive adjustments based on the viewer's emotions in real time, resulting in the creation of more effective training videos. Furthermore, multilingual support for narration and subtitles facilitates international use.

[1274] "Text-based materials" are electronic documents that consist primarily of textual information and are provided in formats such as PDFs, Word documents, and text files.

[1275] "Generative AI" is a software system that uses artificial intelligence techniques to analyze data and automate complex tasks.

[1276] "Narration" refers to the commentary or explanation provided as audio within a video.

[1277] "Subtitles" are narrations and audio content displayed as textual information, in a format that can be visually understood by viewers.

[1278] "BGM that matches the scene" automatically selects and adds background music that matches each scene in the video.

[1279] "User facial expression and voice data" refers to data obtained from facial expressions and voice collected while a user is watching a video.

[1280] "Real-time emotion analysis" is a process that uses facial expression recognition and voice analysis technology to instantly analyze and determine a user's emotions.

[1281] "Dynamic adjustment of video content" refers to the operation of changing and optimizing elements such as video narration, subtitles, and background music in real time based on the results of sentiment analysis.

[1282] "Multilingual" refers to the ability of a system to function in multiple languages ​​and to translate and generate narration and subtitles into those languages.

[1283] An "interface means" is a software component that provides a mechanism for a user to perform operations and inputs through interaction with the system.

[1284] This invention is a system that allows users to easily import text-based materials and automatically generates high-quality training videos using generative AI. In particular, it combines an emotion engine to recognize user emotions in real time and enable more adaptive content creation that reflects feedback.

[1285] First, the user selects text-based materials (e.g., PDF, Word, text files, etc.) from their device and uploads them to the system. The device reads the materials from local storage and sends them to the server. Specifically, the user sends "Product Training Materials.pdf" to the server using the system's upload button.

[1286] Next, the server receives the uploaded file and identifies its format. It uses the server's text extraction module to extract text information and its image extraction module to obtain image data. For example, the server analyzes "Product Training Materials.pdf" and extracts the text and images for each page.

[1287] The server's generative AI then analyzes the content and structure of the materials. Generative AI is a software system that uses artificial intelligence technology to analyze data and automate complex tasks. It divides each page into sections and identifies important keywords and phrases. It improves the page layout based on basic principles of visual design to improve visibility. If necessary, it generates additional images and diagrams and incorporates them into the materials. As a specific example, the generative AI identifies and highlights "product features," "usage instructions," and "precautions" from the main text of "product training materials."

[1288] The server then begins generating videos based on the analyzed and polished materials. Animated slides are created using text and image data, and a natural language processing module generates narration. A translation module generates narration in multiple languages ​​and automatically generates subtitles to accompany the narration. For example, a slide showing "Product Features" is animated, and a narrator generates the following: "Product A is the latest model with the best performance on the market." Appropriate background music for the video is then selected from the library and added as "BGM_Upbeat.mp3."

[1289] When a user watches a video, the device uses a camera and microphone to collect the user's facial expression and voice data. This data is then sent to a server, where the server's emotion analysis engine uses facial expression and voice analysis technology to recognize the user's emotions in real time. For example, if the device streams camera data and detects the user's smile, the server determines that the user is in an "interested" state.

[1290] Based on the results of the sentiment analysis, the server dynamically adjusts the video content. If it determines that the user is confused, it changes the narration to make it easier to understand and adds additional subtitles. For example, if the user is determined to be in a "confused" state, it changes the narration to "What's important here is the safety features of product A. Let me explain in detail," and adds a subtitle "Safety feature details" with key points highlighted.

[1291] Finally, the user reviews the generated video and inputs any necessary corrections through the interface. The server receives the user's feedback and readjusts or corrects the video content. For example, if the user inputs a correction request saying "the narration is too fast," the server presents the user with a regenerated video with the narration speed adjusted.

[1292] Examples of prompt statements

[1293] User: "I've uploaded some product training materials. I'd like to see videos that visually explain the features and usage of each product."

[1294] Generative AI: "Analyzing your document. Identifying each section and identifying keywords. Improving page layout and adding appropriate images and charts."

[1295] Emotion Engine: "We detected that your audience was confused. We'll update the narration to make it clearer and add subtitles to re-emphasize key points."

[1296] Through the above process, users can easily create intuitive and effective training videos through a system that combines an emotion analysis engine.

[1297] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1298] Step 1:

[1299] User import of materials

[1300] A user selects text-based materials (e.g., PDF, Word, text files, etc.) on their device and uploads them to the system. The device reads the materials from local storage and sends them to the server. For example, the user selects "Product Training Materials.pdf" and clicks the system's upload button. The device reads the selected file and sends an HTTP POST request to the server. The input is the material file selected by the user, and the output is the material file sent to the server.

[1301] Step 2:

[1302] Server analysis of data

[1303] The server receives the document file sent from the terminal. It determines the file format (e.g., PDF, DOCX, TXT), extracts text information using the text extraction module, and obtains image data using the image extraction module. For example, the server receives "Product Training Materials.pdf" and uses a PDF parser to extract the text and images for each page. The input is the received document file, and the output is the extracted text data and image data.

[1304] Step 3:

[1305] Brushing up materials with generative AI

[1306] The server's generation AI analyzes the content and structure of the materials. It divides each page into sections and identifies important keywords and phrases. It improves the page layout based on basic principles of visual design to improve visibility. If necessary, it generates additional images and diagrams and incorporates them into the materials. Specifically, the generation AI identifies "product features," "usage instructions," and "precautions" from the main text of the "product training materials" and highlights this information. The input is the extracted text and image data, and the output is polished material data.

[1307] Step 4:

[1308] Video generation and adding additional elements

[1309] The server begins generating videos based on the analyzed and refined materials. Animated slides are created using text and image data, and a natural language processing module is used to generate narration. A translation module generates narration in multiple languages ​​as needed. Subtitles are automatically generated to match the narration, and appropriate background music is selected from a library and added to the video. For example, a "Product Features" slide is animated, and the narrator generates, "Product A is the latest model with the best performance on the market." Additionally, background music suitable for the video is selected from a library and added as "BGM_Upbeat.mp3." The input is the refined material data, and the output is the generated animated video.

[1310] Step 5:

[1311] Emotion analysis using an emotion engine

[1312] When a user watches a video, the device uses a camera and microphone to collect the user's facial and voice data. This data is sent to a server, where the server's emotion analysis engine uses facial and voice analysis technologies to recognize the user's emotions in real time. For example, if the device streams camera data and detects a user's smile, the server determines that the user is in an "interested" state. The input is the collected facial and voice data, and the output is emotion data determined in real time.

[1313] Step 6:

[1314] Adaptive content adjustment

[1315] The server dynamically adjusts elements in the video (such as narration, subtitles, and background music) based on the results of sentiment analysis. For example, if it determines that the user is confused, it changes the narration to something more understandable and adds subtitles that re-emphasize important points. For example, it changes the narration to "What's important here are the safety features of product A. Let me explain in detail," and adds a subtitle "Safety feature details" that highlights the important points. The input is the sentiment analysis results, and the output is the dynamically adjusted video content.

[1316] Step 7:

[1317] Check and correct the video

[1318] The user reviews the generated video and inputs any necessary corrections through the interface. The server receives the user's feedback and readjusts and corrects the video content. For example, if the user inputs a correction request saying "the narration is too fast," the server presents the user with a regenerated video with the narration speed adjusted. The input is the user's correction feedback, and the output is the corrected video.

[1319] Through the above processing steps, users can easily create intuitive and effective training videos through a system that combines an emotion analysis engine.

[1320] (Application example 2)

[1321] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1322] Conventional training video creation systems have difficulty adjusting content in real time to take user emotions into account, limiting their effectiveness in improving viewer comprehension. Furthermore, the generation of multilingual narration and subtitles is inefficient because it requires manual editing due to insufficient automation.

[1323] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for importing text-based materials, a means for analyzing the imported materials using a generation AI, brushing up the screen composition to make it easy to see and hear, and converting it into an animated video, a means for automatically generating narration, subtitles, and background music that matches the scene, and a means for analyzing the user's emotions in real time using an emotion engine and dynamically adjusting the video content. This makes it possible to automatically generate high-quality training videos that adapt to the user's emotions and also easily support multiple languages.

[1324] "Text-based materials" refers to documents that primarily contain text, such as PDFs, Word documents, and text files.

[1325] "Generative AI" refers to a system that uses artificial intelligence technology to analyze input data and generate new content.

[1326] "Brushing up" refers to improving input materials to make them visually easier to read and understand.

[1327] An "animated video" refers to a video that expresses movement by displaying multiple still images in sequence.

[1328] "Narration" refers to audio commentary that accompanies a video scene.

[1329] "Subtitles" refers to the display of video content and audio as text.

[1330] A "scene" refers to a specific situation or moment within a video.

[1331] "BGM" is an abbreviation for background music and refers to the music played in the background of a video.

[1332] "Emotion engine" refers to artificial intelligence technology that recognizes and analyzes a user's emotional state from their facial expressions and voice.

[1333] "Analyzing user emotions in real time" means determining the user's emotions from their facial expressions and voice at the moment they are watching a video.

[1334] "Dynamic adjustment of video content" refers to changing elements of a video, such as narration and subtitles, in real time based on the results of user sentiment analysis.

[1335] "Smart devices" refers to mobile information terminals with computer functions, such as smartphones and tablets.

[1336] "Collecting facial and voice data" refers to recording the user's facial movements and vocalizations using the smart device's camera and microphone.

[1337] This invention is a system that allows users to easily import text-based materials using a smart device, analyze them with a generative AI, and automatically generate high-quality training videos using an emotion engine. Specific embodiments of this system are described below.

[1338] 1. User Import of Materials

[1339] Users select text-based materials from their smart devices (such as smartphones or tablets) and upload them to the system. During this process, users press a button to open a file selection dialog, and the selected file (PDF, Word, or text file) is sent to the server. The hardware used here is a smart device, and the software is React Native and Amazon S3.

[1340] 2. Analysis of the data

[1341] The server analyzes the document files received from Amazon S3 and extracts text and image data. It identifies the file format (PDF, DOCX, TXT) and uses a specific text extraction module to extract character information. This process includes basic principles of visual design, allowing the generative AI to analyze the content and structure of the document and improve the page layout.

[1342] 3. Video generation and adding additional elements

[1343] The server uses generative AI to generate animated videos based on the content of the materials. It automatically generates narration, subtitles, and background music to match the scenes, improving visibility and comprehension. If multilingual support is required, a translation module is used to generate narration and automatically generate subtitles. The software used includes generative AI models and natural language processing technology.

[1344] 4. Emotion Analysis Using an Emotion Engine

[1345] When a user watches a video on a device, the smart device's camera and microphone are used to collect facial expressions and voice data, and emotions are analyzed in real time. The emotion engine uses facial recognition technology (OpenCV) and voice analysis technology (Google Cloud Speech-to-Text) to recognize the user's emotions.

[1346] 5. Real-time content adjustment

[1347] The server dynamically adjusts the video's narration and subtitles based on the analysis results of the emotion engine, allowing it to add additional narration explanations if the user is confused and display subtitles that emphasize important points.

[1348] 6. Check and correct the video

[1349] The user can review the generated video and input any necessary corrections. The server then adjusts the video content based on the user's feedback and provides the optimal training video.

[1350] For example, if a user uploads "product training" materials, the generative AI will visually arrange each product's features, usage instructions, and precautions in an easy-to-understand manner, automatically generating diagrams, narration, and subtitles. Furthermore, the emotion engine will analyze the user's facial expressions and voice, and if it determines that their understanding is incomplete, it will insert additional explanations into the video.

[1351] Example prompt sentence:

[1352] Please select the file you want to upload.

[1353] "Data analysis completed."

[1354] "We recognize the emotions of our users."

[1355] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1356] Step 1:

[1357] A user uses a smart device to select text-based materials and upload them to the system. Specifically, the user presses the file selection button and selects either a PDF, Word, or text file from the file selection dialog on the smart device. The selected file is sent to Amazon S3 by the device. The input is the file selected by the user, and the output is the file uploaded to the server.

[1358] Step 2:

[1359] The server receives documents uploaded to Amazon S3 and determines the file format (PDF, DOCX, TXT). It uses a text extraction module to extract text information, and if necessary, an image extraction module to obtain image data. The input is the uploaded document file, and the output is the extracted text data and image data. Specifically, it analyzes the file format and runs the extraction module corresponding to each format.

[1360] Step 3:

[1361] The server's generative AI analyzes the content and structure based on the analyzed material. It identifies important keywords and phrases and improves the page layout based on visual design. If necessary, it adds images and charts generated by the generative AI to the material. The input is the extracted text data and image data, and the output is a polished page layout and additional visual elements. Specifically, the generative AI analyzes the text and reconstructs the page layout.

[1362] Step 4:

[1363] The server uses generative AI to generate animated videos based on the content of the materials. The information on each page is animated along a timeline, and a natural language processing module is used to generate narration. Furthermore, if multilingual support is required, a translation module is used to generate narration and automatically generate subtitles. The input is the polished materials and layout information, and the output is an animated video with narration. Specific operations include natural language processing and timeline setting.

[1364] Step 5:

[1365] When a user watches a video on a device, facial expression and voice data are collected using the device's camera and microphone. This data is sent to a server in real time and analyzed by an emotion engine. The input is the user's facial expression and voice data, and the output is analyzed emotion data. Specifically, the system collects camera video and voice data in real time and performs facial recognition and voice analysis.

[1366] Step 6:

[1367] The server dynamically adjusts the video content based on the results of emotion analysis. If it determines that the user is confused, it adds supplementary narration explanations and displays subtitles that re-emphasize important points. The input is the analyzed emotion data, and the output is the adjusted video content. Specific operations include updating the narration content and subtitles.

[1368] Step 7:

[1369] The user reviews the generated video and inputs corrections as necessary. The server receives the user's feedback and readjusts the video content. The input is the user's feedback, and the output is the final corrected video. Specific operations include collecting feedback through the interface and re-editing the video.

[1370] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1371] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1372] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1373] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1374] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1375] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1376] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1377] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1378] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1379] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1380] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1381] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1382] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1383] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1384] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1385] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1386] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1387] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1388] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1389] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1390] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1391] The following is further disclosed regarding the above embodiment.

[1392] (Claim 1)

[1393] a means for capturing text-based material; and

[1394] A method to analyze the imported materials using generative AI, brush up the screen composition to make it easy to see and hear, and convert it into an animated video,

[1395] A means to automatically generate narration, subtitles, and background music to match the scene,

[1396] A system including:

[1397] (Claim 2)

[1398] an interface means for a user to check the automatically generated video and input any necessary corrections;

[1399] A means for adjusting and correcting the video based on a user's correction instructions;

[1400] The system of claim 1 further comprising:

[1401] (Claim 3)

[1402] 10. The system of claim 1, further comprising means for generating narration and subtitles in multiple languages.

[1403] "Example 1"

[1404] (Claim 1)

[1405] a means for capturing text-based material; and

[1406] means for determining the format of the imported material and extracting text data and image data;

[1407] The extracted materials are analyzed using generative AI, the content is organized by section, and the layout is polished to make it easier to read.

[1408] A means to automatically generate narration, subtitles, and background music to match the scene,

[1409] A system including:

[1410] (Claim 2)

[1411] an interface means for a user to check the automatically generated video and input any necessary corrections;

[1412] A means for adjusting and correcting the video based on a user's correction instructions;

[1413] The system of claim 1 further comprising:

[1414] (Claim 3)

[1415] 10. The system of claim 1, further comprising means for translating and generating narration and subtitles in multiple languages.

[1416] "Application Example 1"

[1417] (Claim 1)

[1418] a means for capturing text-based material; and

[1419] A method to analyze the imported materials using generative AI, brush up the screen composition to make it easy to see and hear, and convert it into an animated video,

[1420] A means to automatically generate narration, subtitles, and background music to match the scene,

[1421] a means of converting narration from text to audio;

[1422] means for automatically selecting and generating visual decoration elements for use in generating the video;

[1423] A system including:

[1424] (Claim 2)

[1425] an interface means for a user to check the automatically generated video and input any necessary corrections;

[1426] A means for adjusting and correcting the video based on a user's correction instructions;

[1427] A means to reselect and modify visual and audio elements suitable for the video,

[1428] 2. The system of claim 1, comprising:

[1429] (Claim 3)

[1430] A means for generating narration and subtitles in multiple languages;

[1431] A means to apply the generated videos to operation guides and security training,

[1432] 10. The system of claim 1, comprising:

[1433] "Example 2: Combining Emotion Engines"

[1434] (Claim 1)

[1435] a means for capturing text-based material; and

[1436] A method to analyze the imported materials using generative AI, brush up the screen composition to make it easy to see and hear, and convert it into an animated video,

[1437] A means to automatically generate narration, subtitles, and background music to match the scene,

[1438] An emotion analysis means that collects the user's facial expression and voice data and recognizes their emotions in real time;

[1439] means for dynamically adjusting video content based on sentiment analysis results;

[1440] A system including:

[1441] (Claim 2)

[1442] an interface means for a user to check the automatically generated video and input any necessary corrections;

[1443] A means for adjusting and correcting the video based on a user's correction instructions;

[1444] The system of claim 1 further comprising:

[1445] (Claim 3)

[1446] 10. The system of claim 1, further comprising means for generating narration and subtitles in multiple languages.

[1447] "Application example 2 when combining emotion engines"

[1448] (Claim 1)

[1449] a means for capturing text-based material; and

[1450] A method to analyze the imported materials using generative AI, brush up the screen composition to make it easy to see and hear, and convert it into an animated video,

[1451] A means to automatically generate narration, subtitles, and background music to match the scene,

[1452] A means for analyzing user emotions in real time using an emotion engine and dynamically adjusting video content;

[1453] A system including:

[1454] (Claim 2)

[1455] an interface means for a user to check the automatically generated video and input any necessary corrections;

[1456] A means for adjusting and correcting the video based on a user's correction instructions;

[1457] 10. The system of claim 1.

[1458] (Claim 3)

[1459] A means for generating narration and subtitles in multiple languages;

[1460] A means for collecting facial expression and voice data of a user using a smart device and performing emotion analysis;

[1461] 10. The system of claim 1. [Explanation of symbols]

[1462] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for capturing text-based material; and A method to analyze the imported materials using generative AI, brush up the screen composition to make it easy to see and hear, and convert it into an animated video, A means to automatically generate narration, subtitles, and background music to match the scene, A system including:

2. an interface means for a user to check the automatically generated video and input any necessary corrections; A means for adjusting and correcting the video based on a correction instruction from a user; The system of claim 1 further comprising:

3. 10. The system of claim 1, further comprising means for generating narration and subtitles in multiple languages.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A