System

The system addresses the challenge of complex manual understanding by converting manuals into easy-to-understand videos, enhancing user comprehension and reducing errors.

JP2026038085APending Publication Date: 2026-03-06SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Technical manuals and complex procedures are often difficult for general users to understand due to specialized knowledge requirements, leading to incorrect operation or configuration.

Method used

A system that converts manual files into text format, analyzes the content using generative artificial intelligence to extract important information and structure, formats it into easy-to-understand units, generates video scenarios, allows user review and correction, and provides a final video for easy understanding.

Benefits of technology

Facilitates user comprehension of complex manuals by visualizing information, reducing operational errors and improving understanding through easy-to-understand videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026038085000001_ABST
    Figure 2026038085000001_ABST
Patent Text Reader

Abstract

To provide a system capable of making a user easily understand a manual and preventing erroneous operation and setting mistake.SOLUTION: A means for a user to upload a manual file from a terminal, a means for a server to extract contents as a text format, a means for a generation artificial intelligence to analyze the extracted text and extract important information and a structure of each section, and a means for the server to divide the contents of the manual based on the analysis result of the generation artificial intelligence; A system including means for shaping into a unit that can be easily understood, means for generating a video scenario based on the shaped text by the generative artificial intelligence, means for transmitting the generated scenario to a video generation program by the server, means for generating a specific video by the user confirming the generated video and requesting correction as necessary, means for receiving the correction request by the server and instructing the generative artificial intelligence to correct the video again, and means for providing the finally completed video to the user by the server.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Currently, technical manuals and complex procedures often require a great deal of time and specialized knowledge to understand. This makes it extremely difficult for general users to understand these manuals and operate and configure them properly. This becomes even more difficult when the manuals contain a lot of technical terminology or are overly difficult to understand. This can lead to incorrect operation or configuration, resulting in inconvenience and trouble for users. To solve this problem, a means is needed to provide complex manuals in an easy-to-understand format. [Means for solving the problem]

[0005] The present invention provides a system including: a means for a user to upload a manual file from a terminal; a means for a server to receive the uploaded manual file and extract the contents in text format; a means for a generative artificial intelligence to analyze the extracted text and extract important information and structure from each section; a means for the server to divide the contents of the manual based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; a means for the generative artificial intelligence to generate a video scenario based on the formatted text; a means for the server to pass the generated scenario to a video generation program and generate a specific video; a means for the user to check the generated video and request corrections if necessary; a means for the server to receive a correction request, again instruct the generative artificial intelligence to make corrections and re-generate the video; and a means for the server to provide the final completed video to the user.By providing a system including: a means for the contents of the manual to be visually easier to understand as a video, the system makes it easier for the user to understand the manual and prevents operational errors and setting errors.

[0006] "User" refers to the end user who uses the system to upload manuals and request video confirmation and correction.

[0007] A "terminal" is a device such as a computer, smartphone, or tablet that is operated by a user.

[0008] "Server" is a central computer system that receives manuals, analyzes them, generates images, processes corrections, and provides the final images.

[0009] A "manual file" is an instruction manual or procedure document written in a format such as PDF or Word and provided by the user.

[0010] "Generative AI" refers to AI technology that analyzes text and generates video scenarios, for example, using natural language processing and machine learning.

[0011] A "video generation program" is a program that creates specific images based on a scenario created by generative artificial intelligence.

[0012] "Extract as text" is the process of extracting the desired text data from the uploaded manual file.

[0013] "Extracting important information and structure from each section" is a process of identifying particularly important information and its structure from the manual text.

[0014] A "video scenario" is a document or instruction manual that explains the content and structure of a video created by generative artificial intelligence.

[0015] "Generating concrete images" is the process of actually creating visual animations based on a video scenario.

[0016] A "modification request" is an act in which a user sends a request for improvements or changes to a generated video.

[0017] "Visually easy to understand" means that complex information is visualized in an easy-to-understand manner so that users can easily understand it. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The system based on the present invention converts complex manuals into easy-to-understand images. The system program and its processing will be explained below in natural language.

[0040] Basic system configuration

[0041] The system is based on a series of processes in which a user uploads a manual file using a terminal, and the server processes the file to generate an image. The system includes the following elements:

[0042] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[0043] 2. Server: The back-end system that receives and processes uploaded files.

[0044] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[0045] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[0046] Specific processing of the invention

[0047] The specific processing contents of the present invention will be explained below.

[0048] Manual file upload and analysis

[0049] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives and stores the file.

[0050] The server then converts the manual file into text, sometimes using OCR technology. The converted text is then passed to a generative AI that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[0051] Splitting and formatting content

[0052] Based on the results of the generative AI analysis, the server divides the manual into easy-to-understand sections. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[0053] Video scenario generation

[0054] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice prompt such as "Now plug in the power cord" followed by an animation showing the actual connection steps.

[0055] Video generation and preview

[0056] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[0057] Footage review and correction

[0058] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[0059] Provision of final footage

[0060] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[0061] Specific examples

[0062] For example, if a user has a setup manual for a home 3D printer, when the user uploads the manual, the server analyzes the contents and generates a video scenario like the one below.

[0063] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[0064] 2. Video on how to load filament: Instructions and animation on how to "Attach the filament to the spool holder and thread it through the feeder."

[0065] In this way, by visually showing specific operations, the user can easily understand the procedure and proceed with the work.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] A user accesses the web application using a terminal and uploads a manual file (PDF, Word, etc.). The upload screen provides a file selection button and a drag-and-drop function.

[0069] Step 2:

[0070] The server receives the uploaded manual file and saves it in a specified folder on the server. It detects the file type and size and stores it as a temporary file in the appropriate format.

[0071] Step 3:

[0072] The server converts the manual files stored on the server into text format. For PDF and image files, OCR technology is used to extract the text, and for Word files, the text is analyzed directly.

[0073] Step 4:

[0074] The server sends the extracted text data to a generative AI to begin analysis. The generative AI divides the text into sections and adds tags such as "headings," "procedures," and "notes."

[0075] Step 5:

[0076] Based on the analysis results received from the generative AI, the server formats the contents of the manual into appropriate units, including emphasizing important points (by bolding or color coding) and rearranging each section.

[0077] Step 6:

[0078] A generative AI system generates a video scenario based on the formatted text. This scenario includes appropriate animations and visual elements for each section. For example, the step "Connect the power cord" includes an animation of the connection action.

[0079] Step 7:

[0080] The server passes the video scenario received from the generative AI to the video generation program, which then generates the specific video. Video generation includes text-to-speech, animation creation, and the addition of effects.

[0081] Step 8:

[0082] The server temporarily saves the generated video and provides a preview link to the user, who can click on this link to view the generated video.

[0083] Step 9:

[0084] If the user checks the video and feels that corrections are necessary, they can input the specific corrections and submit a correction request. The correction request may include pointing out insufficient explanations or incorrect information.

[0085] Step 10:

[0086] The server receives the correction request and again instructs the generative AI to make the correction. The generative AI generates a new scenario based on the corrections, and the video generation program regenerates the corrected video.

[0087] Step 11:

[0088] The server stores the final video and provides the user with a final link that they can click to download and watch the final video.

[0089] This series of steps makes it possible to visualize complex manuals in a way that is easy for users to understand.

[0090] Example 1

[0091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0092] Conventional manuals are difficult to understand because they are composed only of text. In particular, procedures and precautions that require specialized knowledge cannot be accurately understood by users, which can lead to mistakes and misunderstandings. For this reason, there is a demand for manuals that are easy to understand and can be provided to users in a visual format.

[0093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0094] In this invention, the server includes: means for a user to upload a manual file from a terminal; means for the server to receive the uploaded manual file and extract the content in text format; means for a generative artificial intelligence to analyze the extracted text and extract important information and structure of each section; means for the server to divide the content of the manual based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; means for the generative artificial intelligence to generate a video scenario based on the formatted text; means for the server to pass the generated scenario to a video generation program and generate a specific video; means for the user to check the generated video and request corrections as necessary; means for the server to receive the correction request, again instruct the generative artificial intelligence to make corrections and regenerate the video; and means for the server to provide the final completed video to the user. This makes it possible to convert complex manuals into easy-to-understand videos.

[0095] The "server" is a central processing unit that receives manual files uploaded by users and processes the contents.

[0096] A "terminal" is a device that a user uses to upload a manual file.

[0097] A "manual file" is a document that describes the procedures and standards for a product or service, and includes formats such as PDF and Word.

[0098] "Text format" refers to character data extracted from a manual file.

[0099] "Generative artificial intelligence (AI)" is an artificial intelligence system that analyzes extracted text, extracts important information, and generates video scenarios.

[0100] "Analysis" is the process by which generative artificial intelligence understands the content of text data and performs tagging and extraction of important information.

[0101] A "section" is a unit in which the contents of the manual are divided into categories.

[0102] A "video scenario" is a plan in which a generative artificial intelligence uses formatted text to create the framework for a video, including visual elements and narration text.

[0103] A "video generation program" is software for generating specific video based on a video scenario.

[0104] A "modification request" is a request made by a user to request that a generated video be modified.

[0105] "Video" refers to visually presented content created by generative artificial intelligence and video generation programs.

[0106] A "preview" is a visual display means that allows a user to check the generated video.

[0107] The system based on this invention converts complex manuals into easy-to-understand images. This system is based on a series of processes in which a user uploads a manual file using a terminal, and a server processes the file to generate an image. The system program and its processing are described below.

[0108] Basic system configuration

[0109] The system includes the following elements:

[0110] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[0111] 2. Server: This is the back-end system that receives and stores uploaded files and generates videos using generative artificial intelligence (AI) and video generation programs. Tesseract OCR is used as the OCR technology.

[0112] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios. This AI engine analyzes the extracted text, tags each section, and extracts key information.

[0113] 4. Video generation program: A program for generating video based on a scenario created by generative AI. Specifically, it uses video processing tools such as FFmpeg.

[0114] Program processing

[0115] Manual file upload and analysis

[0116] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives the file and saves it in a specified directory. The server then converts the manual file into text using Tesseract OCR technology. The converted text is then passed to a generative artificial intelligence (AI) that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[0117] Specific prompt examples:

[0118] Convert the following PDF manual to text and tag each section. Example tags: Heading, Procedures, Notes.

[0119] Splitting and formatting content

[0120] The server then uses the results of the generative AI analysis to break down the manual into easy-to-understand chunks, properly formatting each section and applying text styles to highlight important steps.

[0121] Specific prompt examples:

[0122] Split and format the manual text based on the analysis results, highlighting important steps in bold.

[0123] Video scenario generation

[0124] Based on the formatted text, a generative AI system generates a video scenario for each section, including visual elements (animations, diagrams, annotations) and narrative text.

[0125] Specific prompt examples:

[0126] Please generate a video scenario based on the following text. Please create a scenario that includes visual elements and narration text.

[0127] Video generation and preview

[0128] The generated scenario is passed to a video generation program via the server, which generates specific images. The images undergo processing such as text-to-speech, animation generation, and the addition of effects, and are temporarily saved on the server. The generated images are then provided to the user as a preview link.

[0129] Footage review and correction

[0130] The user can check the generated video using the provided preview link and send correction requests as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[0131] Specific prompt examples:

[0132] Please generate a correction scenario based on the following user request:

[0133] Provision of final footage

[0134] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[0135] Example: If a user has a setup manual for a home 3D printer, the server will analyze the contents and generate specific video scenarios such as "how to connect the power" and "how to load the filament." This allows the user to easily understand the steps and proceed with the work.

[0136] The above is a specific embodiment of the system according to the present invention.

[0137] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0138] Step 1:

[0139] A user uploads a manual file from their terminal. First, the user accesses the web application and clicks the "Upload Manual File" button. Next, the user selects a manual file in PDF or Word format using the file selection dialog and starts uploading. The input is the manual file selected by the user. The output is the uploaded manual file, which is sent to the server and saved.

[0140] Step 2:

[0141] The server receives the uploaded manual file and saves it in a specified directory. Then, it converts the manual file into text format using OCR technology (Tesseract OCR). The input is the saved manual file. The output is the converted text format.

[0142] Step 3:

[0143] The server passes the converted text to a generative artificial intelligence (AI). The AI ​​analyzes the text, tags each section (headings, procedures, precautions, etc.), and extracts important information. The input is the text of the manual. The output is an analysis result in which each section is tagged and important information is extracted. For example, the AI ​​analyzes the text and divides it into sections such as "Connecting the power supply" and "Loading the filament," highlighting each step.

[0144] Example prompt for a generative AI model:

[0145] Convert the following PDF manual to text and tag each section. Example tags: Heading, Procedures, Notes.

[0146] Step 4:

[0147] Based on the analysis results of the generative AI, the server divides the manual's contents into easy-to-understand units and formats them. The server then adjusts the text style accordingly, highlighting important steps in bold. The input is the analysis results. The output is formatted and highlighted text.

[0148] Example prompt for a generative AI model:

[0149] Split and format the manual text based on the analysis results, highlighting important steps in bold.

[0150] Step 5:

[0151] A generative AI generates a video scenario based on formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narration text. The input is the formatted text. The output is a generated video scenario.

[0152] Example prompt for a generative AI model:

[0153] Please generate a video scenario based on the following text. Please create a scenario that includes visual elements and narration text.

[0154] Step 6:

[0155] The server passes the generated scenario to a video generation program, which generates a specific video. The video is processed by text-to-speech, animation generation, and effects addition. The input is the video scenario. The output is the completed video. The video is generated using tools such as FFmpeg. The generated video is temporarily stored on the server, and a preview link is provided to the user.

[0156] Step 7:

[0157] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video. The input is the user's correction request. The output is the corrected video.

[0158] Example prompt for a generative AI model:

[0159] Please generate a correction scenario based on the following user request:

[0160] Step 8:

[0161] The server provides the final completed video to the user. The server generates a link to the final video and sends it to the user, allowing the user to easily watch the video. The input is the completed video. The output is a video link provided to the user.

[0162] The above is a description of the specific operations at each step, and the data processing and data calculation based on the input and output.

[0163] (Application example 1)

[0164] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0165] Machines and equipment installed in existing factories often require complex operating manuals, making it difficult for operators to understand and apply these manuals to actual operations. Furthermore, manuals are typically provided on paper or in simple text format, making them difficult to understand visually. This increases the likelihood of operational errors and reduces production efficiency. The present invention aims to solve these problems and improve operational efficiency and accuracy by providing easy-to-understand, visual instructions for operating machines installed in factories.

[0166] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0167] In this invention, the server includes: means for a user to upload a manual file from a terminal; means for the server to receive the uploaded manual file and extract the contents in text format; means for a generative artificial intelligence to analyze the extracted text and extract important information and structure of each section; means for the server to divide the manual content based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; means for the generative artificial intelligence to generate a video scenario based on the formatted text; means for the server to pass the generated scenario to a video generation program and generate a specific video; means for the user to check the generated video and request corrections as necessary; means for the server to receive the correction request, again instruct the generative artificial intelligence to make corrections and regenerate the video; means for the server to provide the final completed video to the user; and means for the system to visualize the operating procedures of machines installed in a factory and display the generated video on a display screen to support robot operation, thereby improving the efficiency of robot operation in a factory and enabling accurate understanding of the operating procedures.

[0168] "User" means a person who uses the system to upload and review manual files.

[0169] "Terminal" refers to the device used by the user to upload the manual file.

[0170] A "manual" is a document that describes the operating procedures for a machine or device.

[0171] "File" means a digital document containing the contents of an operation manual.

[0172] "Server" refers to the computer system that receives and processes uploaded manual files.

[0173] "Extraction" refers to the process by which the server extracts the necessary text information from the uploaded file.

[0174] "Generative AI" refers to an AI engine that analyzes extracted text and generates a video scenario.

[0175] "Analysis" refers to the process by which generative artificial intelligence analyzes the content of text.

[0176] A "section" is a part or item of a manual.

[0177] "Important information" refers to operating procedures and precautions that should be particularly emphasized in the manual.

[0178] "Structure" refers to the chapter and paragraph structure of a text.

[0179] "Splitting" refers to the act of the server dividing the manual into smaller units based on the analysis results.

[0180] "Formatting" is the process of making divided text into an easy-to-understand form.

[0181] A "video scenario" refers to the direction and structure for making a video, created by generative artificial intelligence.

[0182] A "video generation program" is software for creating specific video based on a video scenario.

[0183] "Confirmation" refers to the act of the user viewing the generated video.

[0184] A "modification request" is a user request for changes to be made to the generated video.

[0185] "Regeneration" refers to the process by which the server regenerates the video based on the modification request.

[0186] "Providing" refers to the act of the server handing over the final completed video to the user.

[0187] "System" refers to the technical infrastructure that includes the above means and elements and operates as a whole.

[0188] A "factory" is a production facility where machinery and equipment are installed.

[0189] "Machinery" refers to equipment and tools used in a factory.

[0190] An "operating procedure" is a sequence of actions to ensure the correct operation of a machine or device.

[0191] "Visualization" refers to the process of creating a video to visually show operating procedures.

[0192] A "robot" is an automated device that performs tasks in a factory.

[0193] "Operation support" refers to assisting robots to perform their work efficiently.

[0194] "Display screen" means a monitor for displaying the generated image.

[0195] This invention is a system for converting complicated machine or device operating procedures into images that are easy to understand visually. A specific implementation method for this system will be described below.

[0196] Basic system configuration

[0197] The system is based on a series of processes in which a user uploads a manual file using a terminal, and the server processes the file to generate an image. The system includes the following elements:

[0198] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[0199] 2. Server: The back-end system that receives and processes uploaded files.

[0200] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[0201] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[0202] Hardware and Software Used

[0203] Hardware: Tablets and display screens attached to factory robots, and user-operated devices

[0204] software:

[0205] Python 3: Programming Language

[0206] OpenAI (registered trademark) API: Used as generative artificial intelligence

[0207] Pillow: Processing images using the Python Image Library

[0208] MoviePy: A video editing library

[0209] Data processing and calculation

[0210] Manual file upload and analysis

[0211] Users upload manual files from their devices through a web application. These manual files are often in PDF or Word format. Once uploaded, the server receives and stores the files. The server then converts the manual files into text format. This conversion may involve the use of OCR technology. The converted text is then passed to a generative artificial intelligence (AI) that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tags such as "headings," "procedures," and "notes."

[0212] Splitting and formatting content

[0213] Based on the results of the generative AI analysis, the server divides the manual into easy-to-understand sections. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[0214] Video scenario generation

[0215] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice prompt such as "Now plug in the power cord" followed by an animation showing the actual connection steps.

[0216] Video generation and preview

[0217] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[0218] Footage review and correction

[0219] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[0220] Provision of final footage

[0221] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[0222] Examples of concrete examples and prompts

[0223] For example, if a user has an operating manual for a machine installed in a factory, the server will analyze the contents and generate a video scenario like the one below when the user uploads the manual.

[0224] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[0225] 2. Video showing how to replace the filter: "Open the filter cover. Remove the old filter and insert the new filter. Close the filter cover." Instructions and animation.

[0226] Example prompt sentence:

[0227] Please convert the following operating procedures into a video scenario.

[0228] 1. How to connect the power supply:

[0229] Connect the power cord to the back of the machine.

[0230] Plug it into an outlet and press the power switch.

[0231] 2. How to replace the filter:

[0232] Open the filter cover.

[0233] Remove the old filter and insert the new one.

[0234] Close the filter cover.

[0235] In this way, the efficiency and accuracy of robot operations in factories can be improved, and operating procedures can be easily understood by users.

[0236] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0237] Step 1:

[0238] Uploading manual files

[0239] A user accesses the web application using a terminal and uploads a manual file. The uploaded manual file is a digital document in PDF or Word format.

[0240] Input: Manual file (PDF or Word format)

[0241] Output: Manual file saved on the server

[0242] Step 2:

[0243] Converting manual files to text format

[0244] The server receives the uploaded manual file and converts the content into text format, which may involve the use of OCR technology, and extracts the required text information.

[0245] Input: Manual file stored on the server

[0246] Output: Extracted text data

[0247] Step 3:

[0248] Text Analysis

[0249] The server passes the extracted text data to a generative AI model for analysis. The AI ​​analyzes the content of the text, tags it with headings, procedures, notes, etc., and extracts important information from each section.

[0250] Input: Extracted text data

[0251] Output: Analysis results (text section information, key information, tags)

[0252] Step 4:

[0253] Dividing and formatting the manual content

[0254] Based on the analysis results of the generative AI, the server divides the manual's contents into easy-to-understand units. Specifically, each section is appropriately formatted and its level of difficulty is assessed. The text style is then adjusted, with important steps being highlighted in bold.

[0255] Input: Analysis results (text section information, key information, tags)

[0256] Output: Formatted text data

[0257] Step 5:

[0258] Video scenario generation

[0259] Based on the formatted text, a generative AI system generates a video scenario for each section, including visual elements (animations, diagrams, annotations) and narrative text.

[0260] Input: Formatted text data

[0261] Output: Video scenario

[0262] Step 6:

[0263] Generating concrete images

[0264] The server passes the generated scenario to a video generation program, which generates specific images. Video generation includes text-to-speech, animation generation, and the addition of effects.

[0265] Input: Video scenario

[0266] Output: Generated video file

[0267] Step 7:

[0268] Preview and check footage

[0269] The user can check and preview the video generated on the device, check the video content, and send correction requests if necessary.

[0270] Input: Generated video file

[0271] Output: User correction request

[0272] Step 8:

[0273] Processing and regenerating correction requests

[0274] The server receives the user's correction request and again instructs the generative AI to make the correction. A new scenario based on the correction instruction is generated, and the video generation program regenerates the corrected video.

[0275] Input: Correction request, initial video scenario

[0276] Output: Corrected video file

[0277] Step 9:

[0278] Provision of final footage

[0279] The server then provides the final completed video to the user, who can then view the video on a display screen attached to the factory robot and check the operating procedures.

[0280] Input: Modified video file

[0281] Output: The final video delivered to the user

[0282] These steps will improve the efficiency and accuracy of robot operations in factories and make it easier for users to understand the operating procedures.

[0283] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0284] The system based on the present invention converts complex manuals into easy-to-understand images and also combines an emotion engine that recognizes the user's emotions. The system program and its processing will be specifically described below.

[0285] Basic system configuration

[0286] The system is based on a series of processes: a user uploads a manual file using a terminal, a server processes the file to generate a video, and an emotion engine recognizes emotions and adjusts the video when the user watches the generated video. The system includes the following elements:

[0287] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[0288] 2. Server: The back-end system that receives and processes uploaded files.

[0289] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[0290] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[0291] 5. Emotion Engine: An engine that recognizes the emotions a user is feeling while watching and adjusts the video content based on those emotions.

[0292] Specific processing of the invention

[0293] The specific processing contents of the present invention will be explained below.

[0294] Manual file upload and analysis

[0295] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives and stores the file.

[0296] The server then converts the manual file into text, sometimes using OCR technology. The converted text is then passed to a generative AI that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[0297] Splitting and formatting content

[0298] The server then divides the manual into easy-to-understand sections based on the results of the generative AI analysis. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[0299] Video scenario generation

[0300] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice command such as "Connect the power cord" followed by an animation showing the actual connection procedure.

[0301] Video generation and preview

[0302] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[0303] Footage review and correction

[0304] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[0305] Emotion recognition and video adjustment using an emotion engine

[0306] The emotion engine recognizes emotions while the user is watching the video. The emotion engine detects stress, loss of interest, and loss of attention through facial and voice recognition. For example, if the user looks anxious, the emotion engine will slow down the commentary or insert additional explanations.

[0307] If the emotion engine determines that the user's interest is declining, it will insert interactive elements (such as quizzes or multiple choice questions) into the video to attract the user's attention. This adjustment information is sent to the server, and the video generation program updates the video based on that information.

[0308] Provision of final footage

[0309] The server stores the final video optimized by the emotion engine and provides the final link to the user, who can click on it to watch and download the final emotion-adjusted video.

[0310] Specific examples

[0311] For example, if a user has a setup manual for a home 3D printer, when the user uploads the manual, the server analyzes the contents and generates a video scenario like the one below.

[0312] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[0313] 2. Video on how to load filament: Instructions and animation on how to "Attach the filament to the spool holder and thread it through the feeder."

[0314] While the user is watching the video, the emotion engine monitors their facial expressions and tone of voice, and if it detects stress or anxiety, it will provide additional explanations for the steps or adjust the speed. If interest or disengagement wanes, it will insert an interactive quiz to keep the user's attention.

[0315] Through this series of processes, complex manuals can be optimized to suit the user's emotions and presented in a visually easy-to-understand format.

[0316] The processing flow will be explained below.

[0317] Step 1:

[0318] A user accesses the web application using a terminal and uploads a manual file (PDF, Word, etc.). The upload screen provides a file selection button and a drag-and-drop function.

[0319] Step 2:

[0320] The server receives the uploaded manual file and saves it in a specified folder on the server. It detects the file type and size and stores it as a temporary file in the appropriate format.

[0321] Step 3:

[0322] The server converts the saved manual files into text format, using OCR technology for PDF and image files, and directly analyzing the text for Word files.

[0323] Step 4:

[0324] The server sends the extracted text data to a generative AI to begin analysis. The generative AI divides the text into sections and adds tags such as "headings," "procedures," and "notes."

[0325] Step 5:

[0326] Based on the analysis results received from the generative AI, the server formats the contents of the manual into appropriate units, including emphasizing important points (by bolding or color coding) and rearranging each section.

[0327] Step 6:

[0328] A generative AI system generates a video scenario based on the formatted text. This scenario includes appropriate animations and visual elements for each section. For example, the step "Connect the power cord" includes an animation of the connection action.

[0329] Step 7:

[0330] The server passes the video scenario received from the generative AI to the video generation program, which then generates the specific video. Video generation includes text-to-speech, animation creation, and the addition of effects.

[0331] Step 8:

[0332] The server temporarily saves the generated video and provides a preview link to the user, who can click on this link to view the generated video.

[0333] Step 9:

[0334] If the user checks the video and feels that corrections are necessary, they can input the specific corrections and submit a correction request. The correction request may include pointing out insufficient explanations or incorrect information.

[0335] Step 10:

[0336] The server receives the correction request and again instructs the generative AI to make the correction. The generative AI generates a new scenario based on the corrections, and the video generation program regenerates the corrected video.

[0337] Step 11:

[0338] The emotion engine recognizes emotions while the user is watching the video. The emotion engine detects stress, loss of interest, and loss of attention through facial and voice recognition. For example, if the user looks anxious, the emotion engine will slow down the commentary or insert additional explanations.

[0339] Step 12:

[0340] If the emotion engine determines that the user's interest is declining, it will insert interactive elements (such as quizzes or multiple choice questions) into the video to attract the user's attention. This adjustment information is sent to the server, and the video generation program updates the video based on that information.

[0341] Step 13:

[0342] The server stores the final video optimized by the emotion engine and provides the final link to the user, who can click on it to watch and download the final emotion-adjusted video.

[0343] Through this series of steps, complex manuals are transformed into easy-to-understand images, and the content is optimized according to the user's emotions.

[0344] Example 2

[0345] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0346] Conventional manuals and guides are mainly in text format, which not only makes the content complex and difficult to understand, but also makes it difficult to adjust to the user's emotions and level of understanding. As a result, many users are unable to understand the procedures correctly, which often leads to operational errors and wastes time. Furthermore, with conventional systems, when users request corrections, it is sometimes tedious and the system is unable to respond quickly. A new system is needed to solve these issues.

[0347] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0348] In this invention, the server includes a means for users to upload manual files from their terminals, a means for receiving the uploaded manual files and extracting their contents in text format, and a means for generative AI to analyze the extracted text and extract important information and structure for each section, thereby enabling the generation and immediate modification of images according to the user's emotions and level of understanding.

[0349] A "user" is an entity that uses the system to upload manual files, check the generated video, and request corrections.

[0350] A "terminal" is a device used by a user to upload a manual file, and includes, for example, a computer or a smartphone.

[0351] "Manual files" refer to materials uploaded by users, and include document files such as PDF and Word formats.

[0352] A "server" is a computer system that receives and processes uploaded manual files.

[0353] "Means for extracting content in text format" refers to the process of converting manual files into text data using OCR (Optical Character Recognition) technology.

[0354] "Generative AI" refers to AI technology that analyzes extracted text and extracts key information and structure from each section.

[0355] "Methods for extracting important information and structure" refers to the process of analyzing the content of text, tagging it with "headings," "procedures," "notes," etc., and extracting important information.

[0356] "Means of dividing and arranging the information into easy-to-understand units" refers to the process of appropriately arranging each section based on the analysis results of the generative artificial intelligence, and highlighting it according to its level of difficulty and importance.

[0357] "Means for generating a video scenario" refers to the process of designing a video scenario including visual elements and narrative text based on formatted text.

[0358] "Video generation program" refers to a program for generating specific video based on a generated video scenario, and includes functions such as text reading, animation generation, and adding effects.

[0359] "Preview Link" means a URL or hyperlink provided to a user to allow the user to view the generated video.

[0360] A "modification request" refers to a request submitted by a user to request modification of a generated video.

[0361] An "emotion engine" is an engine that monitors the user's facial expressions and tone of voice and adjusts the video content based on those emotions.

[0362] "Means for adjusting video content based on emotions" refers to a process for dynamically adjusting the speed and content of video explanations based on the user's emotional data detected by the emotion engine.

[0363] The system based on the present invention converts complex manuals into easy-to-understand images and combines an emotion engine that recognizes the user's emotions and adjusts the image content accordingly. This system is based on a series of processes: a user uploads a manual file using a terminal, a server processes the file to generate an image, and when the user watches the generated image, the emotion engine recognizes the emotions and adjusts the image accordingly.

[0364] Hardware and software used

[0365] Hardware

[0366] Device: The device where the user uploads the manual file, for example, a computer, tablet, or smartphone.

[0367] Server: A high-performance computer system for receiving, analyzing, storing, generating and providing images of manual files.

[0368] A camera and microphone are built into the user's device for recognition by the emotion engine.

[0369] software

[0370] Web application: An interface for users to upload files, preview footage, and request revisions. Developed using HTML5, CSS3, and JavaScript (registered trademark).

[0371] OCR technology: Software used to convert uploaded manual files into text format. For example, Tesseract OCR.

[0372] Generative Artificial Intelligence (AI) module: Algorithms for text analysis and scenario generation, incorporating natural language processing technology.

[0373] Video generator: A tool for generating video based on a scenario. Examples include FFmpeg and OpenToonz.

[0374] Emotion Engine: An engine that recognizes the user's emotions and adjusts the video content accordingly. It uses machine learning, facial recognition, and voice recognition technology.

[0375] Specific examples of implementation

[0376] Uploading manual files

[0377] A user uploads a manual file (e.g., PDF) from their device through a web application. The web application provides a file picker function that allows the user to select a file from their local disk.

[0378] File saving and text conversion

[0379] The server receives the uploaded file and stores it in a database. Then, the server converts it into text data using Tesseract OCR. For example, the text of each page of a PDF is extracted.

[0380] Text Analysis and Tagging

[0381] The server passes the converted text data to a generative AI, which analyzes the text, tags it with headings, procedures, notes, and other information, and extracts important information. NLP techniques are used to identify each section and title.

[0382] Video scenario generation

[0383] Based on the analyzed and formatted text, generative AI generates a video scenario for each section. This scenario includes visual elements (animations, diagrams, annotations) and narration text. For example, for the scenario "Connect the power cord," an animation and explanation of how to connect the cord is generated.

[0384] Video generation and preview

[0385] The server passes the scenario to a video generation program, which then generates a specific video. Based on the specific scenario, text is read aloud, animation is generated, and effects are added. The generated video file is temporarily stored on the server, and a preview link is provided to the user.

[0386] Footage review and correction

[0387] The user clicks on the provided preview link to watch the generated video. If necessary, they can submit a correction request, specifying the specific corrections to be made through a feedback form or comment function. Once the correction request is received by the server, the generative AI adjusts the scenario based on the correction instructions, and a new video is regenerated.

[0388] Emotion recognition and video adjustment

[0389] The emotion engine monitors the user's facial expressions and tone of voice via a camera and microphone while the user is watching the video. For example, if the user frowns or sighs, this is recognized. The emotion engine analyzes the user's emotional data and sends instructions to the server as needed to adjust the speed and content of the video explanation.

[0390] Prompt Sentence Examples

[0391] For example, the following prompt is passed to the generative AI:

[0392] "Generate a video scenario that shows the user how to connect a power cord. The instructions should include animation of the cord being connected and narration emphasizing key points."

[0393] The specific embodiment described above allows the user to easily understand a complex manual and to make immediate corrections and adjustments as necessary.

[0394] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0395] Step 1:

[0396] The user uploads the manual file from the terminal through the web application. The user uses the file picker function to upload the selected file from the local disk. The input is the manual file in PDF or Word format, and the output is the file uploaded to the server.

[0397] Step 2:

[0398] The server receives the uploaded manual file and stores it in a database. Then, the server converts the manual file into a text format using OCR technology (e.g., Tesseract OCR). The input is the saved file, and the output is the extracted text data. For example, the text content read from each page of a PDF is aggregated into a single text file.

[0399] Step 3:

[0400] The server inputs the converted text data into generative artificial intelligence (AI). The AI ​​analyzes the text, tags it with headings, procedures, notes, and other information, and extracts important information. The input is the extracted text data, and the output is tagged text data. For example, natural language processing techniques can be used to identify the titles and paragraphs of each section and assign them tags.

[0401] Step 4:

[0402] Based on the analysis results of the generative AI, the server divides the manual content and formats it into easy-to-understand units. Specifically, it formats each section, highlighting important procedures and notes in bold or color. The input is tagged text data, and the output is formatted text data. For example, the analyzed data can be formatted in an easy-to-read format such as HTML or PDF.

[0403] Step 5:

[0404] A generative AI system generates a video scenario based on formatted text. This scenario includes visual elements and narration text. The input is formatted text data, and the output is video scenario data (e.g., JSON format). For example, for the step "Connect the power cord," a scenario with animation and narration is designed.

[0405] Step 6:

[0406] The server passes the generated scenario to a video generation program (e.g., FFmpeg, OpenToonz), which generates the specific video. The video generation program reads the text, generates animation, and adds effects based on the scenario. The input is the video scenario data, and the output is the generated video file. For example, it synthesizes the necessary animations and narration audio based on each scenario.

[0407] Step 7:

[0408] The user clicks on the preview link provided on their device to view the generated video. If necessary, they can send a correction request. The user's feedback is sent to the server, and specific corrections are recorded in the form of comments. The input is the preview link, and the output is the user's correction request.

[0409] Step 8:

[0410] The server receives the correction request, instructs the generative AI to make the corrections again, and regenerates the video. During the regeneration process, the scenario and video content are adjusted based on user feedback. The input is the correction request and the original scenario data, and the output is a corrected video file. For example, adjustments are made such as adding detailed explanations of specific steps.

[0411] Step 9:

[0412] The emotion engine recognizes the user's emotions and monitors facial expressions and tone of voice while watching. The emotion engine uses machine learning models to detect stress, loss of interest, etc. The input is the user's real-time emotion data, and the output is emotion analysis data.

[0413] Step 10:

[0414] The server receives feedback from the emotion engine and adjusts the video content as needed. For example, if the user looks anxious, it may slow down the commentary or insert additional explanation. The input is emotion analysis data, and the output is an optimized video file.

[0415] Step 11:

[0416] The server stores the final video and provides the user with a final link that they can click to view and download the final adjusted video. The input is the optimized video file, and the output is the final preview link.

[0417] (Application example 2)

[0418] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0419] Conventional procedures and manuals for setting up and maintaining autonomous vehicles are often provided in text format, which often requires time to understand. Furthermore, users may feel anxious or stressed while performing these procedures, which increases the risk of incorrect operation. Furthermore, if users lose interest, they may skip important steps. Therefore, there is a need for visual manuals that are easy for users to understand and that adapt to their emotions.

[0420] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0421] In this invention, the server includes means for users to upload manual files from their terminals, means for the server to receive the uploaded manual files and extract the contents in text format, and means for analyzing the emotions of the user while watching the video through emotion recognition means and adjusting the video content in real time based on the analysis results. This provides a video manual that is easier for users to understand, and the video content is adjusted in real time according to the emotions, making it possible to reduce anxiety and stress in the user and maintain their interest and attention.

[0422] A "user terminal" is an electronic device that a user uses to upload manual files.

[0423] The "server" is a computer system that receives uploaded manual files, extracts the contents in text format, and generates images in cooperation with generative artificial intelligence and video generation programs.

[0424] "Generative AI" is an AI engine that analyzes extracted text, extracts important information and structure from each section, and then generates a video scenario.

[0425] A "video generation program" is software that generates specific images based on a scenario created by generative artificial intelligence.

[0426] "Emotion recognition means" is a technology that analyzes the emotions of a user while watching a video and adjusts the video content in real time based on the analysis results.

[0427] "Text format" refers to a format in which the contents of the manual file are converted into character data.

[0428] A "video scenario" is a specific instruction or guideline for video production generated by generative artificial intelligence.

[0429] A "modification request" is a request made by a user to make changes or additions to the generated video content.

[0430] "Real-time adjustment" refers to the operation of instantly adjusting the video content according to the results of the user's emotion analysis.

[0431] A "manual file" is a document file that contains instructions on how to set up and maintain the device.

[0432] The system based on this invention visualizes the setup and maintenance procedures of an autonomous vehicle, and provides a series of processes that recognize the user's emotions in real time and adjust the contents accordingly. Each element of the system and the processing based on it are specifically described below.

[0433] Program processing

[0434] The system consists of the following main components:

[0435] 1. User device: An electronic device to which the user uploads the settings and maintenance procedures for the autonomous vehicle. The device uses a smartphone (iOS, ANDROID (registered trademark)).

[0436] 2. Server: A cloud system that receives uploaded files, extracts the content into text format, and connects it to the generative AI and video generation program. The cloud server uses AWS (registered trademark), Google (registered trademark), or Microsoft (registered trademark) Azure (registered trademark).

[0437] 3. Generative Artificial Intelligence (AI): Analyzes the extracted text, extracts the key information and structure of each section, and generates a video scenario. The AI ​​engine used is OpenAI GPT-4 (registered trademark) or LLM.

[0438] 4. Video generation program: Generates specific images based on a scenario created by generative AI. The video generation tools used are Adobe After Effects API and FFmpeg.

[0439] 5. Emotion Recognition: Analyzes the emotions of users while they watch a video and adjusts the video content in real time based on the analysis results. Facial and voice recognition technologies use Microsoft Azure Face API and Affectiva SDK.

[0440] Hardware and software used

[0441] Hardware: Smartphone, cloud server.

[0442] software:

[0443] OCR tools (Google Cloud Vision, Tesseract OCR)

[0444] AI engine (OpenAI GPT-4, LLM)

[0445] Video generation tools (Adobe After Effects API, FFmpeg)

[0446] Emotion recognition engine (Microsoft Azure Face API, Affectiva SDK)

[0447] Data processing and calculation

[0448] 1. Uploading manual files: Users upload the configuration files for the autonomous vehicle from their smartphones. The file formats are PDF and Word.

[0449] 2. Text extraction: The cloud server uses OCR technology to convert the content into text format.

[0450] 3. Text analysis: Generative AI analyzes the text and extracts important information such as manual operations and points of caution.

[0451] 4. Video scenario generation: Generative AI generates an easy-to-understand video scenario based on the extracted information.

[0452] 5. Video generation: The video generation program generates specific videos based on the generated scenario. The videos include animations and narration.

[0453] 6. Emotion analysis and real-time adjustment: While the user is watching the video, the emotion recognition means analyzes the user's emotions and adjusts the video content in real time if it detects stress or loss of interest.

[0454] Specific examples

[0455] For example, a specific example of uploading a software update procedure for an autonomous vehicle is shown below.

[0456] Preparation before software update: A scenario and animation of parking the vehicle in a safe place and turning off the engine.

[0457] How to perform the update: Scenario and animation of connecting the USB drive to the vehicle and selecting Update from the Settings menu.

[0458] While the user is watching this video, emotion recognition means monitors facial expressions and vocal tone, and if the video is difficult to understand, detailed explanations are added or the video is slowed down.

[0459] Prompt Sentence Examples

[0460] An example of a prompt sentence to input to the generative AI model is as follows:

[0461] "Amazing results! Take our self-driving vehicle software update manual as an example:

[0462] Park your vehicle in a safe place and turn off the engine.

[0463] Then, plug the USB drive into your vehicle and select Updates from the Settings menu.

[0464] Convert this into a visually compelling video scenario, including any animations or text-to-speech you need."

[0465] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0466] Step 1:

[0467] User uploads manual file

[0468] Input: The user uses the terminal to select the settings and maintenance instructions for the autonomous vehicle (PDF or Word format).

[0469] Action: The user selects a manual file from the file selection dialog and clicks the upload button.

[0470] Output: The selected manual file is sent to the server.

[0471] Step 2:

[0472] The server receives the manual file and extracts the contents into text format.

[0473] Input: Manual files uploaded to the server.

[0474] How it works: The server uses an OCR tool (Google Cloud Vision, Tesseract OCR) to convert the contents of the manual file into text format.

[0475] Output: The extracted text data is passed to the generative AI.

[0476] Step 3:

[0477] Generative AI analyzes text data and extracts important information and structure

[0478] Input: Text data extracted by OCR.

[0479] How it works: Generative AI (OpenAI GPT-4, LLM) analyzes the text and extracts key information (steps, notes, etc.) and structure for each section.

[0480] Output: The parsed information and its structure is returned to the server.

[0481] Step 4:

[0482] The server formats the manual contents into easy-to-understand units based on the analysis results.

[0483] Input: Information and its structure analyzed by the generative artificial intelligence.

[0484] What it does: The server breaks up the information and formats it with text styles for better visibility (e.g., making important points bold).

[0485] Output: The formatted text data is passed to the generative artificial intelligence.

[0486] Step 5:

[0487] Generative AI generates video scenarios based on formatted text

[0488] Input: Formatted text data.

[0489] How it works: Generative AI generates a video scenario that includes visual elements (animations, diagrams, annotations) and narration text.

[0490] Output: The generated video scenario is returned to the server.

[0491] Step 6:

[0492] The server passes the video scenario to the video generation program, which generates the video.

[0493] Input: Generated video scenario.

[0494] How it works: The server uses video generation programs (Adobe After Effects API, FFmpeg) to generate specific videos (animated, with narration).

[0495] Output: The generated video data is provided to the user terminal.

[0496] Step 7:

[0497] The user checks the generated video and requests corrections if necessary.

[0498] Input: Generated video data.

[0499] How it works: The user watches the video on their device, marks any necessary corrections, and sends a correction request to the server.

[0500] Output: The modification request arrives at the server.

[0501] Step 8:

[0502] The server receives the correction request, instructs the generative AI to make the correction again, and regenerates the image.

[0503] Input: A correction request from the user.

[0504] How it works: The server passes a modification request to the generative AI, which then modifies the scenario and regenerates the video using the video generation program.

[0505] Output: The corrected video data is provided to the user terminal.

[0506] Step 9:

[0507] Through emotion recognition, the system analyzes the emotions of users while they are watching a video and adjusts the video content in real time.

[0508] Input: User's facial expressions and voice data.

[0509] How it works: The emotion recognition engine (Microsoft Azure Face API, Affectiva SDK) analyzes the user's emotions from their facial expressions and voice in real time, and adjusts the speed and content of the video if it detects stress or a loss of interest.

[0510] Output: The adjusted video content is reflected on the user's device in real time.

[0511] Step 10:

[0512] The server provides the final completed video to the user.

[0513] Input: Final adjusted video data.

[0514] How it works: The server stores the final video data and provides the final link to the user, through which the user can view and download the video.

[0515] Output: A link will be provided where users can watch and download the final video.

[0516] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0517] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0518] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0519] [Second embodiment]

[0520] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0521] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0522] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0523] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0524] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0525] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0526] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0527] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0528] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0529] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0530] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0531] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0532] The system based on the present invention converts complex manuals into easy-to-understand images. The system program and its processing will be explained below in natural language.

[0533] Basic system configuration

[0534] The system is based on a series of processes in which a user uploads a manual file using a terminal, and the server processes the file to generate an image. The system includes the following elements:

[0535] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[0536] 2. Server: The back-end system that receives and processes uploaded files.

[0537] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[0538] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[0539] Specific processing of the invention

[0540] The specific processing contents of the present invention will be explained below.

[0541] Manual file upload and analysis

[0542] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives and stores the file.

[0543] The server then converts the manual file into text, sometimes using OCR technology. The converted text is then passed to a generative AI that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[0544] Splitting and formatting content

[0545] Based on the results of the generative AI analysis, the server divides the manual into easy-to-understand sections. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[0546] Video scenario generation

[0547] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice prompt such as "Now plug in the power cord" followed by an animation showing the actual connection steps.

[0548] Video generation and preview

[0549] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[0550] Footage review and correction

[0551] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[0552] Provision of final footage

[0553] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[0554] Specific examples

[0555] For example, if a user has a setup manual for a home 3D printer, when the user uploads the manual, the server analyzes the contents and generates a video scenario like the one below.

[0556] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[0557] 2. Video on how to load filament: Instructions and animation on how to "Attach the filament to the spool holder and thread it through the feeder."

[0558] In this way, by visually showing specific operations, the user can easily understand the procedure and proceed with the work.

[0559] The processing flow will be explained below.

[0560] Step 1:

[0561] A user accesses the web application using a terminal and uploads a manual file (PDF, Word, etc.). The upload screen provides a file selection button and a drag-and-drop function.

[0562] Step 2:

[0563] The server receives the uploaded manual file and saves it in a specified folder on the server. It detects the file type and size and stores it as a temporary file in the appropriate format.

[0564] Step 3:

[0565] The server converts the manual files stored on the server into text format. For PDF and image files, OCR technology is used to extract the text, and for Word files, the text is analyzed directly.

[0566] Step 4:

[0567] The server sends the extracted text data to a generative AI to begin analysis. The generative AI divides the text into sections and adds tags such as "headings," "procedures," and "notes."

[0568] Step 5:

[0569] Based on the analysis results received from the generative AI, the server formats the contents of the manual into appropriate units, including emphasizing important points (by bolding or color coding) and rearranging each section.

[0570] Step 6:

[0571] A generative AI system generates a video scenario based on the formatted text. This scenario includes appropriate animations and visual elements for each section. For example, the step "Connect the power cord" includes an animation of the connection action.

[0572] Step 7:

[0573] The server passes the video scenario received from the generative AI to the video generation program, which then generates the specific video. Video generation includes text-to-speech, animation creation, and the addition of effects.

[0574] Step 8:

[0575] The server temporarily saves the generated video and provides a preview link to the user, who can click on this link to view the generated video.

[0576] Step 9:

[0577] If the user checks the video and feels that corrections are necessary, they can input the specific corrections and submit a correction request. The correction request may include pointing out insufficient explanations or incorrect information.

[0578] Step 10:

[0579] The server receives the correction request and again instructs the generative AI to make the correction. The generative AI generates a new scenario based on the corrections, and the video generation program regenerates the corrected video.

[0580] Step 11:

[0581] The server stores the final video and provides the user with a final link that they can click to download and watch the final video.

[0582] This series of steps makes it possible to visualize complex manuals in a way that is easy for users to understand.

[0583] Example 1

[0584] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0585] Conventional manuals are difficult to understand because they are composed only of text. In particular, procedures and precautions that require specialized knowledge cannot be accurately understood by users, which can lead to mistakes and misunderstandings. For this reason, there is a demand for manuals that are easy to understand and can be provided to users in a visual format.

[0586] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0587] In this invention, the server includes: means for a user to upload a manual file from a terminal; means for the server to receive the uploaded manual file and extract the content in text format; means for a generative artificial intelligence to analyze the extracted text and extract important information and structure of each section; means for the server to divide the content of the manual based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; means for the generative artificial intelligence to generate a video scenario based on the formatted text; means for the server to pass the generated scenario to a video generation program and generate a specific video; means for the user to check the generated video and request corrections as necessary; means for the server to receive the correction request, again instruct the generative artificial intelligence to make corrections and regenerate the video; and means for the server to provide the final completed video to the user. This makes it possible to convert complex manuals into easy-to-understand videos.

[0588] The "server" is a central processing unit that receives manual files uploaded by users and processes the contents.

[0589] A "terminal" is a device that a user uses to upload a manual file.

[0590] A "manual file" is a document that describes the procedures and standards for a product or service, and includes formats such as PDF and Word.

[0591] "Text format" refers to character data extracted from a manual file.

[0592] "Generative artificial intelligence (AI)" is an artificial intelligence system that analyzes extracted text, extracts important information, and generates video scenarios.

[0593] "Analysis" is the process by which generative artificial intelligence understands the content of text data and performs tagging and extraction of important information.

[0594] A "section" is a unit in which the contents of the manual are divided into categories.

[0595] A "video scenario" is a plan in which a generative artificial intelligence uses formatted text to create the framework for a video, including visual elements and narration text.

[0596] A "video generation program" is software for generating specific video based on a video scenario.

[0597] A "modification request" is a request made by a user to request that a generated video be modified.

[0598] "Video" refers to visually presented content created by generative artificial intelligence and video generation programs.

[0599] A "preview" is a visual display means that allows a user to check the generated video.

[0600] The system based on this invention converts complex manuals into easy-to-understand images. This system is based on a series of processes in which a user uploads a manual file using a terminal, and a server processes the file to generate an image. The system program and its processing are described below.

[0601] Basic system configuration

[0602] The system includes the following elements:

[0603] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[0604] 2. Server: This is the back-end system that receives and stores uploaded files and generates videos using generative artificial intelligence (AI) and video generation programs. Tesseract OCR is used as the OCR technology.

[0605] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios. This AI engine analyzes the extracted text, tags each section, and extracts key information.

[0606] 4. Video generation program: A program for generating video based on a scenario created by generative AI. Specifically, it uses video processing tools such as FFmpeg.

[0607] Program processing

[0608] Manual file upload and analysis

[0609] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives the file and saves it in a specified directory. The server then converts the manual file into text using Tesseract OCR technology. The converted text is then passed to a generative artificial intelligence (AI) that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[0610] Specific prompt examples:

[0611] Convert the following PDF manual to text and tag each section. Example tags: Heading, Procedures, Notes.

[0612] Splitting and formatting content

[0613] The server then uses the results of the generative AI analysis to break down the manual into easy-to-understand chunks, properly formatting each section and applying text styles to highlight important steps.

[0614] Specific prompt examples:

[0615] Split and format the manual text based on the analysis results, highlighting important steps in bold.

[0616] Video scenario generation

[0617] Based on the formatted text, a generative AI system generates a video scenario for each section, including visual elements (animations, diagrams, annotations) and narrative text.

[0618] Specific prompt examples:

[0619] Please generate a video scenario based on the following text. Please create a scenario that includes visual elements and narration text.

[0620] Video generation and preview

[0621] The generated scenario is passed to a video generation program via the server, which generates specific images. The images undergo processing such as text-to-speech, animation generation, and the addition of effects, and are temporarily saved on the server. The generated images are then provided to the user as a preview link.

[0622] Footage review and correction

[0623] The user can check the generated video using the provided preview link and send correction requests as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[0624] Specific prompt examples:

[0625] Please generate a correction scenario based on the following user request:

[0626] Provision of final footage

[0627] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[0628] Example: If a user has a setup manual for a home 3D printer, the server will analyze the contents and generate specific video scenarios such as "how to connect the power" and "how to load the filament." This allows the user to easily understand the steps and proceed with the work.

[0629] The above is a specific embodiment of the system according to the present invention.

[0630] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0631] Step 1:

[0632] A user uploads a manual file from their terminal. First, the user accesses the web application and clicks the "Upload Manual File" button. Next, the user selects a manual file in PDF or Word format using the file selection dialog and starts uploading. The input is the manual file selected by the user. The output is the uploaded manual file, which is sent to the server and saved.

[0633] Step 2:

[0634] The server receives the uploaded manual file and saves it in a specified directory. Then, it converts the manual file into text format using OCR technology (Tesseract OCR). The input is the saved manual file. The output is the converted text format.

[0635] Step 3:

[0636] The server passes the converted text to a generative artificial intelligence (AI). The AI ​​analyzes the text, tags each section (headings, procedures, precautions, etc.), and extracts important information. The input is the text of the manual. The output is an analysis result in which each section is tagged and important information is extracted. For example, the AI ​​analyzes the text and divides it into sections such as "Connecting the power supply" and "Loading the filament," highlighting each step.

[0637] Example prompt for a generative AI model:

[0638] Convert the following PDF manual to text and tag each section. Example tags: Heading, Procedures, Notes.

[0639] Step 4:

[0640] Based on the analysis results of the generative AI, the server divides the manual's contents into easy-to-understand units and formats them. The server then adjusts the text style accordingly, highlighting important steps in bold. The input is the analysis results. The output is formatted and highlighted text.

[0641] Example prompt for a generative AI model:

[0642] Split and format the manual text based on the analysis results, highlighting important steps in bold.

[0643] Step 5:

[0644] A generative AI generates a video scenario based on formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narration text. The input is the formatted text. The output is a generated video scenario.

[0645] Example prompt for a generative AI model:

[0646] Please generate a video scenario based on the following text. Please create a scenario that includes visual elements and narration text.

[0647] Step 6:

[0648] The server passes the generated scenario to a video generation program, which generates a specific video. The video is processed by text-to-speech, animation generation, and effects addition. The input is the video scenario. The output is the completed video. The video is generated using tools such as FFmpeg. The generated video is temporarily stored on the server, and a preview link is provided to the user.

[0649] Step 7:

[0650] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video. The input is the user's correction request. The output is the corrected video.

[0651] Example prompt for a generative AI model:

[0652] Please generate a correction scenario based on the following user request:

[0653] Step 8:

[0654] The server provides the final completed video to the user. The server generates a link to the final video and sends it to the user, allowing the user to easily watch the video. The input is the completed video. The output is a video link provided to the user.

[0655] The above is a description of the specific operations at each step, and the data processing and data calculation based on the input and output.

[0656] (Application example 1)

[0657] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0658] Machines and equipment installed in existing factories often require complex operating manuals, making it difficult for operators to understand and apply these manuals to actual operations. Furthermore, manuals are typically provided on paper or in simple text format, making them difficult to understand visually. This increases the likelihood of operational errors and reduces production efficiency. The present invention aims to solve these problems and improve operational efficiency and accuracy by providing easy-to-understand, visual instructions for operating machines installed in factories.

[0659] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0660] In this invention, the server includes: means for a user to upload a manual file from a terminal; means for the server to receive the uploaded manual file and extract the contents in text format; means for a generative artificial intelligence to analyze the extracted text and extract important information and structure of each section; means for the server to divide the manual content based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; means for the generative artificial intelligence to generate a video scenario based on the formatted text; means for the server to pass the generated scenario to a video generation program and generate a specific video; means for the user to check the generated video and request corrections as necessary; means for the server to receive the correction request, again instruct the generative artificial intelligence to make corrections and regenerate the video; means for the server to provide the final completed video to the user; and means for the system to visualize the operating procedures of machines installed in a factory and display the generated video on a display screen to support robot operation, thereby improving the efficiency of robot operation in a factory and enabling accurate understanding of the operating procedures.

[0661] "User" means a person who uses the system to upload and review manual files.

[0662] "Terminal" refers to the device used by the user to upload the manual file.

[0663] A "manual" is a document that describes the operating procedures for a machine or device.

[0664] "File" means a digital document containing the contents of an operation manual.

[0665] "Server" refers to the computer system that receives and processes uploaded manual files.

[0666] "Extraction" refers to the process by which the server extracts the necessary text information from the uploaded file.

[0667] "Generative AI" refers to an AI engine that analyzes extracted text and generates a video scenario.

[0668] "Analysis" refers to the process by which generative artificial intelligence analyzes the content of text.

[0669] A "section" is a part or item of a manual.

[0670] "Important information" refers to operating procedures and precautions that should be particularly emphasized in the manual.

[0671] "Structure" refers to the chapter and paragraph structure of a text.

[0672] "Splitting" refers to the act of the server dividing the manual into smaller units based on the analysis results.

[0673] "Formatting" is the process of making divided text into an easy-to-understand form.

[0674] A "video scenario" refers to the direction and structure for making a video, created by generative artificial intelligence.

[0675] A "video generation program" is software for creating specific video based on a video scenario.

[0676] "Confirmation" refers to the act of the user viewing the generated video.

[0677] A "modification request" is a user request for changes to be made to the generated video.

[0678] "Regeneration" refers to the process by which the server regenerates the video based on the modification request.

[0679] "Providing" refers to the act of the server handing over the final completed video to the user.

[0680] "System" refers to the technical infrastructure that includes the above means and elements and operates as a whole.

[0681] A "factory" is a production facility where machinery and equipment are installed.

[0682] "Machinery" refers to equipment and tools used in a factory.

[0683] An "operating procedure" is a sequence of actions to ensure the correct operation of a machine or device.

[0684] "Visualization" refers to the process of creating a video to visually show operating procedures.

[0685] A "robot" is an automated device that performs tasks in a factory.

[0686] "Operation support" refers to assisting robots to perform their work efficiently.

[0687] "Display screen" means a monitor for displaying the generated image.

[0688] This invention is a system for converting complicated machine or device operating procedures into images that are easy to understand visually. A specific implementation method for this system will be described below.

[0689] Basic system configuration

[0690] The system is based on a series of processes in which a user uploads a manual file using a terminal, and the server processes the file to generate an image. The system includes the following elements:

[0691] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[0692] 2. Server: The back-end system that receives and processes uploaded files.

[0693] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[0694] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[0695] Hardware and Software Used

[0696] Hardware: Tablets and display screens attached to factory robots, and user-operated devices

[0697] software:

[0698] Python 3: Programming Language

[0699] OpenAI API: Used as generative artificial intelligence

[0700] Pillow: Processing images using the Python Image Library

[0701] MoviePy: A video editing library

[0702] Data processing and calculation

[0703] Manual file upload and analysis

[0704] Users upload manual files from their devices through a web application. These manual files are often in PDF or Word format. Once uploaded, the server receives and stores the files. The server then converts the manual files into text format. This conversion may involve the use of OCR technology. The converted text is then passed to a generative artificial intelligence (AI) that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tags such as "headings," "procedures," and "notes."

[0705] Splitting and formatting content

[0706] Based on the results of the generative AI analysis, the server divides the manual into easy-to-understand sections. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[0707] Video scenario generation

[0708] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice prompt such as "Now plug in the power cord" followed by an animation showing the actual connection steps.

[0709] Video generation and preview

[0710] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[0711] Footage review and correction

[0712] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[0713] Provision of final footage

[0714] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[0715] Examples of concrete examples and prompts

[0716] For example, if a user has an operating manual for a machine installed in a factory, the server will analyze the contents and generate a video scenario like the one below when the user uploads the manual.

[0717] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[0718] 2. Video showing how to replace the filter: "Open the filter cover. Remove the old filter and insert the new filter. Close the filter cover." Instructions and animation.

[0719] Example prompt sentence:

[0720] Please convert the following operating procedures into a video scenario.

[0721] 1. How to connect the power supply:

[0722] Connect the power cord to the back of the machine.

[0723] Plug it into an outlet and press the power switch.

[0724] 2. How to replace the filter:

[0725] Open the filter cover.

[0726] Remove the old filter and insert the new one.

[0727] Close the filter cover.

[0728] In this way, the efficiency and accuracy of robot operations in factories can be improved, and operating procedures can be easily understood by users.

[0729] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0730] Step 1:

[0731] Uploading manual files

[0732] A user accesses the web application using a terminal and uploads a manual file. The uploaded manual file is a digital document in PDF or Word format.

[0733] Input: Manual file (PDF or Word format)

[0734] Output: Manual file saved on the server

[0735] Step 2:

[0736] Converting manual files to text format

[0737] The server receives the uploaded manual file and converts the content into text format, which may involve the use of OCR technology, and extracts the required text information.

[0738] Input: Manual file stored on the server

[0739] Output: Extracted text data

[0740] Step 3:

[0741] Text Analysis

[0742] The server passes the extracted text data to a generative AI model for analysis. The AI ​​analyzes the content of the text, tags it with headings, procedures, notes, etc., and extracts important information from each section.

[0743] Input: Extracted text data

[0744] Output: Analysis results (text section information, key information, tags)

[0745] Step 4:

[0746] Dividing and formatting the manual content

[0747] Based on the analysis results of the generative AI, the server divides the manual's contents into easy-to-understand units. Specifically, each section is appropriately formatted and its level of difficulty is assessed. The text style is then adjusted, with important steps being highlighted in bold.

[0748] Input: Analysis results (text section information, key information, tags)

[0749] Output: Formatted text data

[0750] Step 5:

[0751] Video scenario generation

[0752] Based on the formatted text, a generative AI system generates a video scenario for each section, including visual elements (animations, diagrams, annotations) and narrative text.

[0753] Input: Formatted text data

[0754] Output: Video scenario

[0755] Step 6:

[0756] Generating concrete images

[0757] The server passes the generated scenario to a video generation program, which generates specific images. Video generation includes text-to-speech, animation generation, and the addition of effects.

[0758] Input: Video scenario

[0759] Output: Generated video file

[0760] Step 7:

[0761] Preview and check footage

[0762] The user can check and preview the video generated on the device, check the video content, and send correction requests if necessary.

[0763] Input: Generated video file

[0764] Output: User correction request

[0765] Step 8:

[0766] Processing and regenerating correction requests

[0767] The server receives the user's correction request and again instructs the generative AI to make the correction. A new scenario based on the correction instruction is generated, and the video generation program regenerates the corrected video.

[0768] Input: Correction request, initial video scenario

[0769] Output: Corrected video file

[0770] Step 9:

[0771] Provision of final footage

[0772] The server then provides the final completed video to the user, who can then view the video on a display screen attached to the factory robot and check the operating procedures.

[0773] Input: Modified video file

[0774] Output: The final video delivered to the user

[0775] These steps will improve the efficiency and accuracy of robot operations in factories and make it easier for users to understand the operating procedures.

[0776] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0777] The system based on the present invention converts complex manuals into easy-to-understand images and also combines an emotion engine that recognizes the user's emotions. The system program and its processing will be specifically described below.

[0778] Basic system configuration

[0779] The system is based on a series of processes: a user uploads a manual file using a terminal, a server processes the file to generate a video, and an emotion engine recognizes emotions and adjusts the video when the user watches the generated video. The system includes the following elements:

[0780] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[0781] 2. Server: The back-end system that receives and processes uploaded files.

[0782] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[0783] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[0784] 5. Emotion Engine: An engine that recognizes the emotions a user is feeling while watching and adjusts the video content based on those emotions.

[0785] Specific processing of the invention

[0786] The specific processing contents of the present invention will be explained below.

[0787] Manual file upload and analysis

[0788] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives and stores the file.

[0789] The server then converts the manual file into text, sometimes using OCR technology. The converted text is then passed to a generative AI that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[0790] Splitting and formatting content

[0791] The server then divides the manual into easy-to-understand sections based on the results of the generative AI analysis. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[0792] Video scenario generation

[0793] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice command such as "Connect the power cord" followed by an animation showing the actual connection procedure.

[0794] Video generation and preview

[0795] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[0796] Footage review and correction

[0797] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[0798] Emotion recognition and video adjustment using an emotion engine

[0799] The emotion engine recognizes emotions while the user is watching the video. The emotion engine detects stress, loss of interest, and loss of attention through facial and voice recognition. For example, if the user looks anxious, the emotion engine will slow down the commentary or insert additional explanations.

[0800] If the emotion engine determines that the user's interest is declining, it will insert interactive elements (such as quizzes or multiple choice questions) into the video to attract the user's attention. This adjustment information is sent to the server, and the video generation program updates the video based on that information.

[0801] Provision of final footage

[0802] The server stores the final video optimized by the emotion engine and provides the final link to the user, who can click on it to watch and download the final emotion-adjusted video.

[0803] Specific examples

[0804] For example, if a user has a setup manual for a home 3D printer, when the user uploads the manual, the server analyzes the contents and generates a video scenario like the one below.

[0805] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[0806] 2. Video on how to load filament: Instructions and animation on how to "Attach the filament to the spool holder and thread it through the feeder."

[0807] While the user is watching the video, the emotion engine monitors their facial expressions and tone of voice, and if it detects stress or anxiety, it will provide additional explanations for the steps or adjust the speed. If interest or disengagement wanes, it will insert an interactive quiz to keep the user's attention.

[0808] Through this series of processes, complex manuals can be optimized to suit the user's emotions and presented in a visually easy-to-understand format.

[0809] The processing flow will be explained below.

[0810] Step 1:

[0811] A user accesses the web application using a terminal and uploads a manual file (PDF, Word, etc.). The upload screen provides a file selection button and a drag-and-drop function.

[0812] Step 2:

[0813] The server receives the uploaded manual file and saves it in a specified folder on the server. It detects the file type and size and stores it as a temporary file in the appropriate format.

[0814] Step 3:

[0815] The server converts the saved manual files into text format, using OCR technology for PDF and image files, and directly analyzing the text for Word files.

[0816] Step 4:

[0817] The server sends the extracted text data to a generative AI to begin analysis. The generative AI divides the text into sections and adds tags such as "headings," "procedures," and "notes."

[0818] Step 5:

[0819] Based on the analysis results received from the generative AI, the server formats the contents of the manual into appropriate units, including emphasizing important points (by bolding or color coding) and rearranging each section.

[0820] Step 6:

[0821] A generative AI system generates a video scenario based on the formatted text. This scenario includes appropriate animations and visual elements for each section. For example, the step "Connect the power cord" includes an animation of the connection action.

[0822] Step 7:

[0823] The server passes the video scenario received from the generative AI to the video generation program, which then generates the specific video. Video generation includes text-to-speech, animation creation, and the addition of effects.

[0824] Step 8:

[0825] The server temporarily saves the generated video and provides a preview link to the user, who can click on this link to view the generated video.

[0826] Step 9:

[0827] If the user checks the video and feels that corrections are necessary, they can input the specific corrections and submit a correction request. The correction request may include pointing out insufficient explanations or incorrect information.

[0828] Step 10:

[0829] The server receives the correction request and again instructs the generative AI to make the correction. The generative AI generates a new scenario based on the corrections, and the video generation program regenerates the corrected video.

[0830] Step 11:

[0831] The emotion engine recognizes emotions while the user is watching the video. The emotion engine detects stress, loss of interest, and loss of attention through facial and voice recognition. For example, if the user looks anxious, the emotion engine will slow down the commentary or insert additional explanations.

[0832] Step 12:

[0833] If the emotion engine determines that the user's interest is declining, it will insert interactive elements (such as quizzes or multiple choice questions) into the video to attract the user's attention. This adjustment information is sent to the server, and the video generation program updates the video based on that information.

[0834] Step 13:

[0835] The server stores the final video optimized by the emotion engine and provides the final link to the user, who can click on it to watch and download the final emotion-adjusted video.

[0836] Through this series of steps, complex manuals are transformed into easy-to-understand images, and the content is optimized according to the user's emotions.

[0837] Example 2

[0838] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0839] Conventional manuals and guides are mainly in text format, which not only makes the content complex and difficult to understand, but also makes it difficult to adjust to the user's emotions and level of understanding. As a result, many users are unable to understand the procedures correctly, which often leads to operational errors and wastes time. Furthermore, with conventional systems, when users request corrections, it is sometimes tedious and the system is unable to respond quickly. A new system is needed to solve these issues.

[0840] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0841] In this invention, the server includes a means for users to upload manual files from their terminals, a means for receiving the uploaded manual files and extracting their contents in text format, and a means for generative AI to analyze the extracted text and extract important information and structure for each section, thereby enabling the generation and immediate modification of images according to the user's emotions and level of understanding.

[0842] A "user" is an entity that uses the system to upload manual files, check the generated video, and request corrections.

[0843] A "terminal" is a device used by a user to upload a manual file, and includes, for example, a computer or a smartphone.

[0844] "Manual files" refer to materials uploaded by users, and include document files such as PDF and Word formats.

[0845] A "server" is a computer system that receives and processes uploaded manual files.

[0846] "Means for extracting content in text format" refers to the process of converting manual files into text data using OCR (Optical Character Recognition) technology.

[0847] "Generative AI" refers to AI technology that analyzes extracted text and extracts key information and structure from each section.

[0848] "Methods for extracting important information and structure" refers to the process of analyzing the content of text, tagging it with "headings," "procedures," "notes," etc., and extracting important information.

[0849] "Means of dividing and arranging the information into easy-to-understand units" refers to the process of appropriately arranging each section based on the analysis results of the generative artificial intelligence, and highlighting it according to its level of difficulty and importance.

[0850] "Means for generating a video scenario" refers to the process of designing a video scenario including visual elements and narrative text based on formatted text.

[0851] "Video generation program" refers to a program for generating specific video based on a generated video scenario, and includes functions such as text reading, animation generation, and adding effects.

[0852] "Preview Link" means a URL or hyperlink provided to a user to allow the user to view the generated video.

[0853] A "modification request" refers to a request submitted by a user to request modification of a generated video.

[0854] An "emotion engine" is an engine that monitors the user's facial expressions and tone of voice and adjusts the video content based on those emotions.

[0855] "Means for adjusting video content based on emotions" refers to a process for dynamically adjusting the speed and content of video explanations based on the user's emotional data detected by the emotion engine.

[0856] The system based on the present invention converts complex manuals into easy-to-understand images and combines an emotion engine that recognizes the user's emotions and adjusts the image content accordingly. This system is based on a series of processes: a user uploads a manual file using a terminal, a server processes the file to generate an image, and when the user watches the generated image, the emotion engine recognizes the emotions and adjusts the image accordingly.

[0857] Hardware and software used

[0858] Hardware

[0859] Device: The device where the user uploads the manual file, for example, a computer, tablet, or smartphone.

[0860] Server: A high-performance computer system for receiving, analyzing, storing, generating and providing images of manual files.

[0861] A camera and microphone are built into the user's device for recognition by the emotion engine.

[0862] software

[0863] Web application: An interface for users to upload files, preview footage, and request revisions. Developed using HTML5, CSS3, and JavaScript.

[0864] OCR technology: Software used to convert uploaded manual files into text format. For example, Tesseract OCR.

[0865] Generative Artificial Intelligence (AI) module: Algorithms for text analysis and scenario generation, incorporating natural language processing technology.

[0866] Video generator: A tool for generating video based on a scenario. Examples include FFmpeg and OpenToonz.

[0867] Emotion Engine: An engine that recognizes the user's emotions and adjusts the video content accordingly. It uses machine learning, facial recognition, and voice recognition technology.

[0868] Specific examples of implementation

[0869] Uploading manual files

[0870] A user uploads a manual file (e.g., PDF) from their device through a web application. The web application provides a file picker function that allows the user to select a file from their local disk.

[0871] File saving and text conversion

[0872] The server receives the uploaded file and stores it in a database. Then, the server converts it into text data using Tesseract OCR. For example, the text of each page of a PDF is extracted.

[0873] Text Analysis and Tagging

[0874] The server passes the converted text data to a generative AI, which analyzes the text, tags it with headings, procedures, notes, and other information, and extracts important information. NLP techniques are used to identify each section and title.

[0875] Video scenario generation

[0876] Based on the analyzed and formatted text, generative AI generates a video scenario for each section. This scenario includes visual elements (animations, diagrams, annotations) and narration text. For example, for the scenario "Connect the power cord," an animation and explanation of how to connect the cord is generated.

[0877] Video generation and preview

[0878] The server passes the scenario to a video generation program, which then generates a specific video. Based on the specific scenario, text is read aloud, animation is generated, and effects are added. The generated video file is temporarily stored on the server, and a preview link is provided to the user.

[0879] Footage review and correction

[0880] The user clicks on the provided preview link to watch the generated video. If necessary, they can submit a correction request, specifying the specific corrections to be made through a feedback form or comment function. Once the correction request is received by the server, the generative AI adjusts the scenario based on the correction instructions, and a new video is regenerated.

[0881] Emotion recognition and video adjustment

[0882] The emotion engine monitors the user's facial expressions and tone of voice via a camera and microphone while the user is watching the video. For example, if the user frowns or sighs, this is recognized. The emotion engine analyzes the user's emotional data and sends instructions to the server as needed to adjust the speed and content of the video explanation.

[0883] Prompt Sentence Examples

[0884] For example, the following prompt is passed to the generative AI:

[0885] "Generate a video scenario that shows the user how to connect a power cord. The instructions should include animation of the cord being connected and narration emphasizing key points."

[0886] The specific embodiment described above allows the user to easily understand a complex manual and to make immediate corrections and adjustments as necessary.

[0887] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0888] Step 1:

[0889] The user uploads the manual file from the terminal through the web application. The user uses the file picker function to upload the selected file from the local disk. The input is the manual file in PDF or Word format, and the output is the file uploaded to the server.

[0890] Step 2:

[0891] The server receives the uploaded manual file and stores it in a database. Then, the server converts the manual file into a text format using OCR technology (e.g., Tesseract OCR). The input is the saved file, and the output is the extracted text data. For example, the text content read from each page of a PDF is aggregated into a single text file.

[0892] Step 3:

[0893] The server inputs the converted text data into generative artificial intelligence (AI). The AI ​​analyzes the text, tags it with headings, procedures, notes, and other information, and extracts important information. The input is the extracted text data, and the output is tagged text data. For example, natural language processing techniques can be used to identify the titles and paragraphs of each section and assign them tags.

[0894] Step 4:

[0895] Based on the analysis results of the generative AI, the server divides the manual content and formats it into easy-to-understand units. Specifically, it formats each section, highlighting important procedures and notes in bold or color. The input is tagged text data, and the output is formatted text data. For example, the analyzed data can be formatted in an easy-to-read format such as HTML or PDF.

[0896] Step 5:

[0897] A generative AI system generates a video scenario based on formatted text. This scenario includes visual elements and narration text. The input is formatted text data, and the output is video scenario data (e.g., JSON format). For example, for the step "Connect the power cord," a scenario with animation and narration is designed.

[0898] Step 6:

[0899] The server passes the generated scenario to a video generation program (e.g., FFmpeg, OpenToonz), which generates the specific video. The video generation program reads the text, generates animation, and adds effects based on the scenario. The input is the video scenario data, and the output is the generated video file. For example, it synthesizes the necessary animations and narration audio based on each scenario.

[0900] Step 7:

[0901] The user clicks on the preview link provided on their device to view the generated video. If necessary, they can send a correction request. The user's feedback is sent to the server, and specific corrections are recorded in the form of comments. The input is the preview link, and the output is the user's correction request.

[0902] Step 8:

[0903] The server receives the correction request, instructs the generative AI to make the corrections again, and regenerates the video. During the regeneration process, the scenario and video content are adjusted based on user feedback. The input is the correction request and the original scenario data, and the output is a corrected video file. For example, adjustments are made such as adding detailed explanations of specific steps.

[0904] Step 9:

[0905] The emotion engine recognizes the user's emotions and monitors facial expressions and tone of voice while watching. The emotion engine uses machine learning models to detect stress, loss of interest, etc. The input is the user's real-time emotion data, and the output is emotion analysis data.

[0906] Step 10:

[0907] The server receives feedback from the emotion engine and adjusts the video content as needed. For example, if the user looks anxious, it may slow down the commentary or insert additional explanation. The input is emotion analysis data, and the output is an optimized video file.

[0908] Step 11:

[0909] The server stores the final video and provides the user with a final link that they can click to view and download the final adjusted video. The input is the optimized video file, and the output is the final preview link.

[0910] (Application example 2)

[0911] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0912] Conventional procedures and manuals for setting up and maintaining autonomous vehicles are often provided in text format, which often requires time to understand. Furthermore, users may feel anxious or stressed while performing these procedures, which increases the risk of incorrect operation. Furthermore, if users lose interest, they may skip important steps. Therefore, there is a need for visual manuals that are easy for users to understand and that adapt to their emotions.

[0913] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0914] In this invention, the server includes means for users to upload manual files from their terminals, means for the server to receive the uploaded manual files and extract the contents in text format, and means for analyzing the emotions of the user while watching the video through emotion recognition means and adjusting the video content in real time based on the analysis results. This provides a video manual that is easier for users to understand, and the video content is adjusted in real time according to the emotions, making it possible to reduce anxiety and stress in the user and maintain their interest and attention.

[0915] A "user terminal" is an electronic device that a user uses to upload manual files.

[0916] The "server" is a computer system that receives uploaded manual files, extracts the contents in text format, and generates images in cooperation with generative artificial intelligence and video generation programs.

[0917] "Generative AI" is an AI engine that analyzes extracted text, extracts important information and structure from each section, and then generates a video scenario.

[0918] A "video generation program" is software that generates specific images based on a scenario created by generative artificial intelligence.

[0919] "Emotion recognition means" is a technology that analyzes the emotions of a user while watching a video and adjusts the video content in real time based on the analysis results.

[0920] "Text format" refers to a format in which the contents of the manual file are converted into character data.

[0921] A "video scenario" is a specific instruction or guideline for video production generated by generative artificial intelligence.

[0922] A "modification request" is a request made by a user to make changes or additions to the generated video content.

[0923] "Real-time adjustment" refers to the operation of instantly adjusting the video content according to the results of the user's emotion analysis.

[0924] A "manual file" is a document file that contains instructions on how to set up and maintain the device.

[0925] The system based on this invention visualizes the setup and maintenance procedures of an autonomous vehicle, and provides a series of processes that recognize the user's emotions in real time and adjust the contents accordingly. Each element of the system and the processing based on it are specifically described below.

[0926] Program processing

[0927] The system consists of the following main components:

[0928] 1. User device: An electronic device to which the user uploads the settings and maintenance procedures for the autonomous vehicle. The device uses a smartphone (iOS, Android).

[0929] 2. Server: A cloud system that receives uploaded files, extracts the content into text format, and connects it to the generative AI and video generation program. The cloud server uses AWS, Google Cloud, or Microsoft Azure.

[0930] 3. Generative Artificial Intelligence (AI): Analyzes the extracted text, extracts key information and structure from each section, and generates a video scenario. The AI ​​engines used are OpenAI GPT-4 and LLM.

[0931] 4. Video generation program: Generates specific images based on a scenario created by generative AI. The video generation tools used are Adobe After Effects API and FFmpeg.

[0932] 5. Emotion Recognition: Analyzes the emotions of users while they watch a video and adjusts the video content in real time based on the analysis results. Facial and voice recognition technologies use Microsoft Azure Face API and Affectiva SDK.

[0933] Hardware and software used

[0934] Hardware: Smartphone, cloud server.

[0935] software:

[0936] OCR tools (Google Cloud Vision, Tesseract OCR)

[0937] AI engine (OpenAI GPT-4, LLM)

[0938] Video generation tools (Adobe After Effects API, FFmpeg)

[0939] Emotion recognition engine (Microsoft Azure Face API, Affectiva SDK)

[0940] Data processing and calculation

[0941] 1. Uploading manual files: Users upload the configuration files for the autonomous vehicle from their smartphones. The file formats are PDF and Word.

[0942] 2. Text extraction: The cloud server uses OCR technology to convert the content into text format.

[0943] 3. Text analysis: Generative AI analyzes the text and extracts important information such as manual operations and points of caution.

[0944] 4. Video scenario generation: Generative AI generates an easy-to-understand video scenario based on the extracted information.

[0945] 5. Video generation: The video generation program generates specific videos based on the generated scenario. The videos include animations and narration.

[0946] 6. Emotion analysis and real-time adjustment: While the user is watching the video, the emotion recognition means analyzes the user's emotions and adjusts the video content in real time if it detects stress or loss of interest.

[0947] Specific examples

[0948] For example, a specific example of uploading a software update procedure for an autonomous vehicle is shown below.

[0949] Preparation before software update: A scenario and animation of parking the vehicle in a safe place and turning off the engine.

[0950] How to perform the update: Scenario and animation of connecting the USB drive to the vehicle and selecting Update from the Settings menu.

[0951] While the user is watching this video, emotion recognition means monitors facial expressions and vocal tone, and if the video is difficult to understand, detailed explanations are added or the video is slowed down.

[0952] Prompt Sentence Examples

[0953] An example of a prompt sentence to input to the generative AI model is as follows:

[0954] "Amazing results! Take our self-driving vehicle software update manual as an example:

[0955] Park your vehicle in a safe place and turn off the engine.

[0956] Then, plug the USB drive into your vehicle and select Updates from the Settings menu.

[0957] Convert this into a visually compelling video scenario, including any animations or text-to-speech you need."

[0958] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0959] Step 1:

[0960] User uploads manual file

[0961] Input: The user uses the terminal to select the settings and maintenance instructions for the autonomous vehicle (PDF or Word format).

[0962] Action: The user selects a manual file from the file selection dialog and clicks the upload button.

[0963] Output: The selected manual file is sent to the server.

[0964] Step 2:

[0965] The server receives the manual file and extracts the contents into text format.

[0966] Input: Manual files uploaded to the server.

[0967] How it works: The server uses an OCR tool (Google Cloud Vision, Tesseract OCR) to convert the contents of the manual file into text format.

[0968] Output: The extracted text data is passed to the generative AI.

[0969] Step 3:

[0970] Generative AI analyzes text data and extracts important information and structure

[0971] Input: Text data extracted by OCR.

[0972] How it works: Generative AI (OpenAI GPT-4, LLM) analyzes the text and extracts key information (steps, notes, etc.) and structure for each section.

[0973] Output: The parsed information and its structure is returned to the server.

[0974] Step 4:

[0975] The server formats the manual contents into easy-to-understand units based on the analysis results.

[0976] Input: Information and its structure analyzed by the generative artificial intelligence.

[0977] What it does: The server breaks up the information and formats it with text styles for better visibility (e.g., making important points bold).

[0978] Output: The formatted text data is passed to the generative artificial intelligence.

[0979] Step 5:

[0980] Generative AI generates video scenarios based on formatted text

[0981] Input: Formatted text data.

[0982] How it works: Generative AI generates a video scenario that includes visual elements (animations, diagrams, annotations) and narration text.

[0983] Output: The generated video scenario is returned to the server.

[0984] Step 6:

[0985] The server passes the video scenario to the video generation program, which generates the video.

[0986] Input: Generated video scenario.

[0987] How it works: The server uses video generation programs (Adobe After Effects API, FFmpeg) to generate specific videos (animated, with narration).

[0988] Output: The generated video data is provided to the user terminal.

[0989] Step 7:

[0990] The user checks the generated video and requests corrections if necessary.

[0991] Input: Generated video data.

[0992] How it works: The user watches the video on their device, marks any necessary corrections, and sends a correction request to the server.

[0993] Output: The modification request arrives at the server.

[0994] Step 8:

[0995] The server receives the correction request, instructs the generative AI to make the correction again, and regenerates the image.

[0996] Input: A correction request from the user.

[0997] How it works: The server passes a modification request to the generative AI, which then modifies the scenario and regenerates the video using the video generation program.

[0998] Output: The corrected video data is provided to the user terminal.

[0999] Step 9:

[1000] Through emotion recognition, the system analyzes the emotions of users while they are watching a video and adjusts the video content in real time.

[1001] Input: User's facial expressions and voice data.

[1002] How it works: The emotion recognition engine (Microsoft Azure Face API, Affectiva SDK) analyzes the user's emotions from their facial expressions and voice in real time, and adjusts the speed and content of the video if it detects stress or a loss of interest.

[1003] Output: The adjusted video content is reflected on the user's device in real time.

[1004] Step 10:

[1005] The server provides the final completed video to the user.

[1006] Input: Final adjusted video data.

[1007] How it works: The server stores the final video data and provides the final link to the user, through which the user can view and download the video.

[1008] Output: A link will be provided where users can watch and download the final video.

[1009] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1010] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1011] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1012] [Third embodiment]

[1013] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1014] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1015] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1016] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1017] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1018] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1019] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1020] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1021] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1022] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1023] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1024] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1025] The system based on the present invention converts complex manuals into easy-to-understand images. The system program and its processing will be explained below in natural language.

[1026] Basic system configuration

[1027] The system is based on a series of processes in which a user uploads a manual file using a terminal, and the server processes the file to generate an image. The system includes the following elements:

[1028] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[1029] 2. Server: The back-end system that receives and processes uploaded files.

[1030] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[1031] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[1032] Specific processing of the invention

[1033] The specific processing contents of the present invention will be explained below.

[1034] Manual file upload and analysis

[1035] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives and stores the file.

[1036] The server then converts the manual file into text, sometimes using OCR technology. The converted text is then passed to a generative AI that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[1037] Splitting and formatting content

[1038] Based on the results of the generative AI analysis, the server divides the manual into easy-to-understand sections. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[1039] Video scenario generation

[1040] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice prompt such as "Now plug in the power cord" followed by an animation showing the actual connection steps.

[1041] Video generation and preview

[1042] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[1043] Footage review and correction

[1044] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[1045] Provision of final footage

[1046] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[1047] Specific examples

[1048] For example, if a user has a setup manual for a home 3D printer, when the user uploads the manual, the server analyzes the contents and generates a video scenario like the one below.

[1049] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[1050] 2. Video on how to load filament: Instructions and animation on how to "Attach the filament to the spool holder and thread it through the feeder."

[1051] In this way, by visually showing specific operations, the user can easily understand the procedure and proceed with the work.

[1052] The processing flow will be explained below.

[1053] Step 1:

[1054] A user accesses the web application using a terminal and uploads a manual file (PDF, Word, etc.). The upload screen provides a file selection button and a drag-and-drop function.

[1055] Step 2:

[1056] The server receives the uploaded manual file and saves it in a specified folder on the server. It detects the file type and size and stores it as a temporary file in the appropriate format.

[1057] Step 3:

[1058] The server converts the manual files stored on the server into text format. For PDF and image files, OCR technology is used to extract the text, and for Word files, the text is analyzed directly.

[1059] Step 4:

[1060] The server sends the extracted text data to a generative AI to begin analysis. The generative AI divides the text into sections and adds tags such as "headings," "procedures," and "notes."

[1061] Step 5:

[1062] Based on the analysis results received from the generative AI, the server formats the contents of the manual into appropriate units, including emphasizing important points (by bolding or color coding) and rearranging each section.

[1063] Step 6:

[1064] A generative AI system generates a video scenario based on the formatted text. This scenario includes appropriate animations and visual elements for each section. For example, the step "Connect the power cord" includes an animation of the connection action.

[1065] Step 7:

[1066] The server passes the video scenario received from the generative AI to the video generation program, which then generates the specific video. Video generation includes text-to-speech, animation creation, and the addition of effects.

[1067] Step 8:

[1068] The server temporarily saves the generated video and provides a preview link to the user, who can click on this link to view the generated video.

[1069] Step 9:

[1070] If the user checks the video and feels that corrections are necessary, they can input the specific corrections and submit a correction request. The correction request may include pointing out insufficient explanations or incorrect information.

[1071] Step 10:

[1072] The server receives the correction request and again instructs the generative AI to make the correction. The generative AI generates a new scenario based on the corrections, and the video generation program regenerates the corrected video.

[1073] Step 11:

[1074] The server stores the final video and provides the user with a final link that they can click to download and watch the final video.

[1075] This series of steps makes it possible to visualize complex manuals in a way that is easy for users to understand.

[1076] Example 1

[1077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1078] Conventional manuals are difficult to understand because they are composed only of text. In particular, procedures and precautions that require specialized knowledge cannot be accurately understood by users, which can lead to mistakes and misunderstandings. For this reason, there is a demand for manuals that are easy to understand and can be provided to users in a visual format.

[1079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1080] In this invention, the server includes: means for a user to upload a manual file from a terminal; means for the server to receive the uploaded manual file and extract the content in text format; means for a generative artificial intelligence to analyze the extracted text and extract important information and structure of each section; means for the server to divide the content of the manual based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; means for the generative artificial intelligence to generate a video scenario based on the formatted text; means for the server to pass the generated scenario to a video generation program and generate a specific video; means for the user to check the generated video and request corrections as necessary; means for the server to receive the correction request, again instruct the generative artificial intelligence to make corrections and regenerate the video; and means for the server to provide the final completed video to the user. This makes it possible to convert complex manuals into easy-to-understand videos.

[1081] The "server" is a central processing unit that receives manual files uploaded by users and processes the contents.

[1082] A "terminal" is a device that a user uses to upload a manual file.

[1083] A "manual file" is a document that describes the procedures and standards for a product or service, and includes formats such as PDF and Word.

[1084] "Text format" refers to character data extracted from a manual file.

[1085] "Generative artificial intelligence (AI)" is an artificial intelligence system that analyzes extracted text, extracts important information, and generates video scenarios.

[1086] "Analysis" is the process by which generative artificial intelligence understands the content of text data and performs tagging and extraction of important information.

[1087] A "section" is a unit in which the contents of the manual are divided into categories.

[1088] A "video scenario" is a plan in which a generative artificial intelligence uses formatted text to create the framework for a video, including visual elements and narration text.

[1089] A "video generation program" is software for generating specific video based on a video scenario.

[1090] A "modification request" is a request made by a user to request that a generated video be modified.

[1091] "Video" refers to visually presented content created by generative artificial intelligence and video generation programs.

[1092] A "preview" is a visual display means that allows a user to check the generated video.

[1093] The system based on this invention converts complex manuals into easy-to-understand images. This system is based on a series of processes in which a user uploads a manual file using a terminal, and a server processes the file to generate an image. The system program and its processing are described below.

[1094] Basic system configuration

[1095] The system includes the following elements:

[1096] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[1097] 2. Server: This is the back-end system that receives and stores uploaded files and generates videos using generative artificial intelligence (AI) and video generation programs. Tesseract OCR is used as the OCR technology.

[1098] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios. This AI engine analyzes the extracted text, tags each section, and extracts key information.

[1099] 4. Video generation program: A program for generating video based on a scenario created by generative AI. Specifically, it uses video processing tools such as FFmpeg.

[1100] Program processing

[1101] Manual file upload and analysis

[1102] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives the file and saves it in a specified directory. The server then converts the manual file into text using Tesseract OCR technology. The converted text is then passed to a generative artificial intelligence (AI) that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[1103] Specific prompt examples:

[1104] Convert the following PDF manual to text and tag each section. Example tags: Heading, Procedures, Notes.

[1105] Splitting and formatting content

[1106] The server then uses the results of the generative AI analysis to break down the manual into easy-to-understand chunks, properly formatting each section and applying text styles to highlight important steps.

[1107] Specific prompt examples:

[1108] Split and format the manual text based on the analysis results, highlighting important steps in bold.

[1109] Video scenario generation

[1110] Based on the formatted text, a generative AI system generates a video scenario for each section, including visual elements (animations, diagrams, annotations) and narrative text.

[1111] Specific prompt examples:

[1112] Please generate a video scenario based on the following text. Please create a scenario that includes visual elements and narration text.

[1113] Video generation and preview

[1114] The generated scenario is passed to a video generation program via the server, which generates specific images. The images undergo processing such as text-to-speech, animation generation, and the addition of effects, and are temporarily saved on the server. The generated images are then provided to the user as a preview link.

[1115] Footage review and correction

[1116] The user can check the generated video using the provided preview link and send correction requests as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[1117] Specific prompt examples:

[1118] Please generate a correction scenario based on the following user request:

[1119] Provision of final footage

[1120] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[1121] Example: If a user has a setup manual for a home 3D printer, the server will analyze the contents and generate specific video scenarios such as "how to connect the power" and "how to load the filament." This allows the user to easily understand the steps and proceed with the work.

[1122] The above is a specific embodiment of the system according to the present invention.

[1123] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1124] Step 1:

[1125] A user uploads a manual file from their terminal. First, the user accesses the web application and clicks the "Upload Manual File" button. Next, the user selects a manual file in PDF or Word format using the file selection dialog and starts uploading. The input is the manual file selected by the user. The output is the uploaded manual file, which is sent to the server and saved.

[1126] Step 2:

[1127] The server receives the uploaded manual file and saves it in a specified directory. Then, it converts the manual file into text format using OCR technology (Tesseract OCR). The input is the saved manual file. The output is the converted text format.

[1128] Step 3:

[1129] The server passes the converted text to a generative artificial intelligence (AI). The AI ​​analyzes the text, tags each section (headings, procedures, precautions, etc.), and extracts important information. The input is the text of the manual. The output is an analysis result in which each section is tagged and important information is extracted. For example, the AI ​​analyzes the text and divides it into sections such as "Connecting the power supply" and "Loading the filament," highlighting each step.

[1130] Example prompt for a generative AI model:

[1131] Convert the following PDF manual to text and tag each section. Example tags: Heading, Procedures, Notes.

[1132] Step 4:

[1133] Based on the analysis results of the generative AI, the server divides the manual's contents into easy-to-understand units and formats them. The server then adjusts the text style accordingly, highlighting important steps in bold. The input is the analysis results. The output is formatted and highlighted text.

[1134] Example prompt for a generative AI model:

[1135] Split and format the manual text based on the analysis results, highlighting important steps in bold.

[1136] Step 5:

[1137] A generative AI generates a video scenario based on formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narration text. The input is the formatted text. The output is a generated video scenario.

[1138] Example prompt for a generative AI model:

[1139] Please generate a video scenario based on the following text. Please create a scenario that includes visual elements and narration text.

[1140] Step 6:

[1141] The server passes the generated scenario to a video generation program, which generates a specific video. The video is processed by text-to-speech, animation generation, and effects addition. The input is the video scenario. The output is the completed video. The video is generated using tools such as FFmpeg. The generated video is temporarily stored on the server, and a preview link is provided to the user.

[1142] Step 7:

[1143] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video. The input is the user's correction request. The output is the corrected video.

[1144] Example prompt for a generative AI model:

[1145] Please generate a correction scenario based on the following user request:

[1146] Step 8:

[1147] The server provides the final completed video to the user. The server generates a link to the final video and sends it to the user, allowing the user to easily watch the video. The input is the completed video. The output is a video link provided to the user.

[1148] The above is a description of the specific operations at each step, and the data processing and data calculation based on the input and output.

[1149] (Application example 1)

[1150] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1151] Machines and equipment installed in existing factories often require complex operating manuals, making it difficult for operators to understand and apply these manuals to actual operations. Furthermore, manuals are typically provided on paper or in simple text format, making them difficult to understand visually. This increases the likelihood of operational errors and reduces production efficiency. The present invention aims to solve these problems and improve operational efficiency and accuracy by providing easy-to-understand, visual instructions for operating machines installed in factories.

[1152] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1153] In this invention, the server includes: means for a user to upload a manual file from a terminal; means for the server to receive the uploaded manual file and extract the contents in text format; means for a generative artificial intelligence to analyze the extracted text and extract important information and structure of each section; means for the server to divide the manual content based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; means for the generative artificial intelligence to generate a video scenario based on the formatted text; means for the server to pass the generated scenario to a video generation program and generate a specific video; means for the user to check the generated video and request corrections as necessary; means for the server to receive the correction request, again instruct the generative artificial intelligence to make corrections and regenerate the video; means for the server to provide the final completed video to the user; and means for the system to visualize the operating procedures of machines installed in a factory and display the generated video on a display screen to support robot operation, thereby improving the efficiency of robot operation in a factory and enabling accurate understanding of the operating procedures.

[1154] "User" means a person who uses the system to upload and review manual files.

[1155] "Terminal" refers to the device used by the user to upload the manual file.

[1156] A "manual" is a document that describes the operating procedures for a machine or device.

[1157] "File" means a digital document containing the contents of an operation manual.

[1158] "Server" refers to the computer system that receives and processes uploaded manual files.

[1159] "Extraction" refers to the process by which the server extracts the necessary text information from the uploaded file.

[1160] "Generative AI" refers to an AI engine that analyzes extracted text and generates a video scenario.

[1161] "Analysis" refers to the process by which generative artificial intelligence analyzes the content of text.

[1162] A "section" is a part or item of a manual.

[1163] "Important information" refers to operating procedures and precautions that should be particularly emphasized in the manual.

[1164] "Structure" refers to the chapter and paragraph structure of a text.

[1165] "Splitting" refers to the act of the server dividing the manual into smaller units based on the analysis results.

[1166] "Formatting" is the process of making divided text into an easy-to-understand form.

[1167] A "video scenario" refers to the direction and structure for making a video, created by generative artificial intelligence.

[1168] A "video generation program" is software for creating specific video based on a video scenario.

[1169] "Confirmation" refers to the act of the user viewing the generated video.

[1170] A "modification request" is a user request for changes to be made to the generated video.

[1171] "Regeneration" refers to the process by which the server regenerates the video based on the modification request.

[1172] "Providing" refers to the act of the server handing over the final completed video to the user.

[1173] "System" refers to the technical infrastructure that includes the above means and elements and operates as a whole.

[1174] A "factory" is a production facility where machinery and equipment are installed.

[1175] "Machinery" refers to equipment and tools used in a factory.

[1176] An "operating procedure" is a sequence of actions to ensure the correct operation of a machine or device.

[1177] "Visualization" refers to the process of creating a video to visually show operating procedures.

[1178] A "robot" is an automated device that performs tasks in a factory.

[1179] "Operation support" refers to assisting robots to perform their work efficiently.

[1180] "Display screen" means a monitor for displaying the generated image.

[1181] This invention is a system for converting complicated machine or device operating procedures into images that are easy to understand visually. A specific implementation method for this system will be described below.

[1182] Basic system configuration

[1183] The system is based on a series of processes in which a user uploads a manual file using a terminal, and the server processes the file to generate an image. The system includes the following elements:

[1184] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[1185] 2. Server: The back-end system that receives and processes uploaded files.

[1186] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[1187] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[1188] Hardware and Software Used

[1189] Hardware: Tablets and display screens attached to factory robots, and user-operated devices

[1190] software:

[1191] Python 3: Programming Language

[1192] OpenAI API: Used as generative artificial intelligence

[1193] Pillow: Processing images using the Python Image Library

[1194] MoviePy: A video editing library

[1195] Data processing and calculation

[1196] Manual file upload and analysis

[1197] Users upload manual files from their devices through a web application. These manual files are often in PDF or Word format. Once uploaded, the server receives and stores the files. The server then converts the manual files into text format. This conversion may involve the use of OCR technology. The converted text is then passed to a generative artificial intelligence (AI) that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tags such as "headings," "procedures," and "notes."

[1198] Splitting and formatting content

[1199] Based on the results of the generative AI analysis, the server divides the manual into easy-to-understand sections. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[1200] Video scenario generation

[1201] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice prompt such as "Now plug in the power cord" followed by an animation showing the actual connection steps.

[1202] Video generation and preview

[1203] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[1204] Footage review and correction

[1205] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[1206] Provision of final footage

[1207] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[1208] Examples of concrete examples and prompts

[1209] For example, if a user has an operating manual for a machine installed in a factory, the server will analyze the contents and generate a video scenario like the one below when the user uploads the manual.

[1210] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[1211] 2. Video showing how to replace the filter: "Open the filter cover. Remove the old filter and insert the new filter. Close the filter cover." Instructions and animation.

[1212] Example prompt sentence:

[1213] Please convert the following operating procedures into a video scenario.

[1214] 1. How to connect the power supply:

[1215] Connect the power cord to the back of the machine.

[1216] Plug it into an outlet and press the power switch.

[1217] 2. How to replace the filter:

[1218] Open the filter cover.

[1219] Remove the old filter and insert the new one.

[1220] Close the filter cover.

[1221] In this way, the efficiency and accuracy of robot operations in factories can be improved, and operating procedures can be easily understood by users.

[1222] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1223] Step 1:

[1224] Uploading manual files

[1225] A user accesses the web application using a terminal and uploads a manual file. The uploaded manual file is a digital document in PDF or Word format.

[1226] Input: Manual file (PDF or Word format)

[1227] Output: Manual file saved on the server

[1228] Step 2:

[1229] Converting manual files to text format

[1230] The server receives the uploaded manual file and converts the content into text format, which may involve the use of OCR technology, and extracts the required text information.

[1231] Input: Manual file stored on the server

[1232] Output: Extracted text data

[1233] Step 3:

[1234] Text Analysis

[1235] The server passes the extracted text data to a generative AI model for analysis. The AI ​​analyzes the content of the text, tags it with headings, procedures, notes, etc., and extracts important information from each section.

[1236] Input: Extracted text data

[1237] Output: Analysis results (text section information, key information, tags)

[1238] Step 4:

[1239] Dividing and formatting the manual content

[1240] Based on the analysis results of the generative AI, the server divides the manual's contents into easy-to-understand units. Specifically, each section is appropriately formatted and its level of difficulty is assessed. The text style is then adjusted, with important steps being highlighted in bold.

[1241] Input: Analysis results (text section information, key information, tags)

[1242] Output: Formatted text data

[1243] Step 5:

[1244] Video scenario generation

[1245] Based on the formatted text, a generative AI system generates a video scenario for each section, including visual elements (animations, diagrams, annotations) and narrative text.

[1246] Input: Formatted text data

[1247] Output: Video scenario

[1248] Step 6:

[1249] Generating concrete images

[1250] The server passes the generated scenario to a video generation program, which generates specific images. Video generation includes text-to-speech, animation generation, and the addition of effects.

[1251] Input: Video scenario

[1252] Output: Generated video file

[1253] Step 7:

[1254] Preview and check footage

[1255] The user can check and preview the video generated on the device, check the video content, and send correction requests if necessary.

[1256] Input: Generated video file

[1257] Output: User correction request

[1258] Step 8:

[1259] Processing and regenerating correction requests

[1260] The server receives the user's correction request and again instructs the generative AI to make the correction. A new scenario based on the correction instruction is generated, and the video generation program regenerates the corrected video.

[1261] Input: Correction request, initial video scenario

[1262] Output: Corrected video file

[1263] Step 9:

[1264] Provision of final footage

[1265] The server then provides the final completed video to the user, who can then view the video on a display screen attached to the factory robot and check the operating procedures.

[1266] Input: Modified video file

[1267] Output: The final video delivered to the user

[1268] These steps will improve the efficiency and accuracy of robot operations in factories and make it easier for users to understand the operating procedures.

[1269] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1270] The system based on the present invention converts complex manuals into easy-to-understand images and also combines an emotion engine that recognizes the user's emotions. The system program and its processing will be specifically described below.

[1271] Basic system configuration

[1272] The system is based on a series of processes: a user uploads a manual file using a terminal, a server processes the file to generate a video, and an emotion engine recognizes emotions and adjusts the video when the user watches the generated video. The system includes the following elements:

[1273] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[1274] 2. Server: The back-end system that receives and processes uploaded files.

[1275] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[1276] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[1277] 5. Emotion Engine: An engine that recognizes the emotions a user is feeling while watching and adjusts the video content based on those emotions.

[1278] Specific processing of the invention

[1279] The specific processing contents of the present invention will be explained below.

[1280] Manual file upload and analysis

[1281] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives and stores the file.

[1282] The server then converts the manual file into text, sometimes using OCR technology. The converted text is then passed to a generative AI that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[1283] Splitting and formatting content

[1284] The server then divides the manual into easy-to-understand sections based on the results of the generative AI analysis. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[1285] Video scenario generation

[1286] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice command such as "Connect the power cord" followed by an animation showing the actual connection procedure.

[1287] Video generation and preview

[1288] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[1289] Footage review and correction

[1290] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[1291] Emotion recognition and video adjustment using an emotion engine

[1292] The emotion engine recognizes emotions while the user is watching the video. The emotion engine detects stress, loss of interest, and loss of attention through facial and voice recognition. For example, if the user looks anxious, the emotion engine will slow down the commentary or insert additional explanations.

[1293] If the emotion engine determines that the user's interest is declining, it will insert interactive elements (such as quizzes or multiple choice questions) into the video to attract the user's attention. This adjustment information is sent to the server, and the video generation program updates the video based on that information.

[1294] Provision of final footage

[1295] The server stores the final video optimized by the emotion engine and provides the final link to the user, who can click on it to watch and download the final emotion-adjusted video.

[1296] Specific examples

[1297] For example, if a user has a setup manual for a home 3D printer, when the user uploads the manual, the server analyzes the contents and generates a video scenario like the one below.

[1298] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[1299] 2. Video on how to load filament: Instructions and animation on how to "Attach the filament to the spool holder and thread it through the feeder."

[1300] While the user is watching the video, the emotion engine monitors their facial expressions and tone of voice, and if it detects stress or anxiety, it will provide additional explanations for the steps or adjust the speed. If interest or disengagement wanes, it will insert an interactive quiz to keep the user's attention.

[1301] Through this series of processes, complex manuals can be optimized to suit the user's emotions and presented in a visually easy-to-understand format.

[1302] The processing flow will be explained below.

[1303] Step 1:

[1304] A user accesses the web application using a terminal and uploads a manual file (PDF, Word, etc.). The upload screen provides a file selection button and a drag-and-drop function.

[1305] Step 2:

[1306] The server receives the uploaded manual file and saves it in a specified folder on the server. It detects the file type and size and stores it as a temporary file in the appropriate format.

[1307] Step 3:

[1308] The server converts the saved manual files into text format, using OCR technology for PDF and image files, and directly analyzing the text for Word files.

[1309] Step 4:

[1310] The server sends the extracted text data to a generative AI to begin analysis. The generative AI divides the text into sections and adds tags such as "headings," "procedures," and "notes."

[1311] Step 5:

[1312] Based on the analysis results received from the generative AI, the server formats the contents of the manual into appropriate units, including emphasizing important points (by bolding or color coding) and rearranging each section.

[1313] Step 6:

[1314] A generative AI system generates a video scenario based on the formatted text. This scenario includes appropriate animations and visual elements for each section. For example, the step "Connect the power cord" includes an animation of the connection action.

[1315] Step 7:

[1316] The server passes the video scenario received from the generative AI to the video generation program, which then generates the specific video. Video generation includes text-to-speech, animation creation, and the addition of effects.

[1317] Step 8:

[1318] The server temporarily saves the generated video and provides a preview link to the user, who can click on this link to view the generated video.

[1319] Step 9:

[1320] If the user checks the video and feels that corrections are necessary, they can input the specific corrections and submit a correction request. The correction request may include pointing out insufficient explanations or incorrect information.

[1321] Step 10:

[1322] The server receives the correction request and again instructs the generative AI to make the correction. The generative AI generates a new scenario based on the corrections, and the video generation program regenerates the corrected video.

[1323] Step 11:

[1324] The emotion engine recognizes emotions while the user is watching the video. The emotion engine detects stress, loss of interest, and loss of attention through facial and voice recognition. For example, if the user looks anxious, the emotion engine will slow down the commentary or insert additional explanations.

[1325] Step 12:

[1326] If the emotion engine determines that the user's interest is declining, it will insert interactive elements (such as quizzes or multiple choice questions) into the video to attract the user's attention. This adjustment information is sent to the server, and the video generation program updates the video based on that information.

[1327] Step 13:

[1328] The server stores the final video optimized by the emotion engine and provides the final link to the user, who can click on it to watch and download the final emotion-adjusted video.

[1329] Through this series of steps, complex manuals are transformed into easy-to-understand images, and the content is optimized according to the user's emotions.

[1330] Example 2

[1331] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1332] Conventional manuals and guides are mainly in text format, which not only makes the content complex and difficult to understand, but also makes it difficult to adjust to the user's emotions and level of understanding. As a result, many users are unable to understand the procedures correctly, which often leads to operational errors and wastes time. Furthermore, with conventional systems, when users request corrections, it is sometimes tedious and the system is unable to respond quickly. A new system is needed to solve these issues.

[1333] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1334] In this invention, the server includes a means for users to upload manual files from their terminals, a means for receiving the uploaded manual files and extracting their contents in text format, and a means for generative AI to analyze the extracted text and extract important information and structure for each section, thereby enabling the generation and immediate modification of images according to the user's emotions and level of understanding.

[1335] A "user" is an entity that uses the system to upload manual files, check the generated video, and request corrections.

[1336] A "terminal" is a device used by a user to upload a manual file, and includes, for example, a computer or a smartphone.

[1337] "Manual files" refer to materials uploaded by users, and include document files such as PDF and Word formats.

[1338] A "server" is a computer system that receives and processes uploaded manual files.

[1339] "Means for extracting content in text format" refers to the process of converting manual files into text data using OCR (Optical Character Recognition) technology.

[1340] "Generative AI" refers to AI technology that analyzes extracted text and extracts key information and structure from each section.

[1341] "Methods for extracting important information and structure" refers to the process of analyzing the content of text, tagging it with "headings," "procedures," "notes," etc., and extracting important information.

[1342] "Means of dividing and arranging the information into easy-to-understand units" refers to the process of appropriately arranging each section based on the analysis results of the generative artificial intelligence, and highlighting it according to its level of difficulty and importance.

[1343] "Means for generating a video scenario" refers to the process of designing a video scenario including visual elements and narrative text based on formatted text.

[1344] "Video generation program" refers to a program for generating specific video based on a generated video scenario, and includes functions such as text reading, animation generation, and adding effects.

[1345] "Preview Link" means a URL or hyperlink provided to a user to allow the user to view the generated video.

[1346] A "modification request" refers to a request submitted by a user to request modification of a generated video.

[1347] An "emotion engine" is an engine that monitors the user's facial expressions and tone of voice and adjusts the video content based on those emotions.

[1348] "Means for adjusting video content based on emotions" refers to a process for dynamically adjusting the speed and content of video explanations based on the user's emotional data detected by the emotion engine.

[1349] The system based on the present invention converts complex manuals into easy-to-understand images and combines an emotion engine that recognizes the user's emotions and adjusts the image content accordingly. This system is based on a series of processes: a user uploads a manual file using a terminal, a server processes the file to generate an image, and when the user watches the generated image, the emotion engine recognizes the emotions and adjusts the image accordingly.

[1350] Hardware and software used

[1351] Hardware

[1352] Device: The device where the user uploads the manual file, for example, a computer, tablet, or smartphone.

[1353] Server: A high-performance computer system for receiving, analyzing, storing, generating and providing images of manual files.

[1354] A camera and microphone are built into the user's device for recognition by the emotion engine.

[1355] software

[1356] Web application: An interface for users to upload files, preview footage, and request revisions. Developed using HTML5, CSS3, and JavaScript.

[1357] OCR technology: Software used to convert uploaded manual files into text format. For example, Tesseract OCR.

[1358] Generative Artificial Intelligence (AI) module: Algorithms for text analysis and scenario generation, incorporating natural language processing technology.

[1359] Video generator: A tool for generating video based on a scenario. Examples include FFmpeg and OpenToonz.

[1360] Emotion Engine: An engine that recognizes the user's emotions and adjusts the video content accordingly. It uses machine learning, facial recognition, and voice recognition technology.

[1361] Specific examples of implementation

[1362] Uploading manual files

[1363] A user uploads a manual file (e.g., PDF) from their device through a web application. The web application provides a file picker function that allows the user to select a file from their local disk.

[1364] File saving and text conversion

[1365] The server receives the uploaded file and stores it in a database. Then, the server converts it into text data using Tesseract OCR. For example, the text of each page of a PDF is extracted.

[1366] Text Analysis and Tagging

[1367] The server passes the converted text data to a generative AI, which analyzes the text, tags it with headings, procedures, notes, and other information, and extracts important information. NLP techniques are used to identify each section and title.

[1368] Video scenario generation

[1369] Based on the analyzed and formatted text, generative AI generates a video scenario for each section. This scenario includes visual elements (animations, diagrams, annotations) and narration text. For example, for the scenario "Connect the power cord," an animation and explanation of how to connect the cord is generated.

[1370] Video generation and preview

[1371] The server passes the scenario to a video generation program, which then generates a specific video. Based on the specific scenario, text is read aloud, animation is generated, and effects are added. The generated video file is temporarily stored on the server, and a preview link is provided to the user.

[1372] Footage review and correction

[1373] The user clicks on the provided preview link to watch the generated video. If necessary, they can submit a correction request, specifying the specific corrections to be made through a feedback form or comment function. Once the correction request is received by the server, the generative AI adjusts the scenario based on the correction instructions, and a new video is regenerated.

[1374] Emotion recognition and video adjustment

[1375] The emotion engine monitors the user's facial expressions and tone of voice via a camera and microphone while the user is watching the video. For example, if the user frowns or sighs, this is recognized. The emotion engine analyzes the user's emotional data and sends instructions to the server as needed to adjust the speed and content of the video explanation.

[1376] Prompt Sentence Examples

[1377] For example, the following prompt is passed to the generative AI:

[1378] "Generate a video scenario that shows the user how to connect a power cord. The instructions should include animation of the cord being connected and narration emphasizing key points."

[1379] The specific embodiment described above allows the user to easily understand a complex manual and to make immediate corrections and adjustments as necessary.

[1380] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1381] Step 1:

[1382] The user uploads the manual file from the terminal through the web application. The user uses the file picker function to upload the selected file from the local disk. The input is the manual file in PDF or Word format, and the output is the file uploaded to the server.

[1383] Step 2:

[1384] The server receives the uploaded manual file and stores it in a database. Then, the server converts the manual file into a text format using OCR technology (e.g., Tesseract OCR). The input is the saved file, and the output is the extracted text data. For example, the text content read from each page of a PDF is aggregated into a single text file.

[1385] Step 3:

[1386] The server inputs the converted text data into generative artificial intelligence (AI). The AI ​​analyzes the text, tags it with headings, procedures, notes, and other information, and extracts important information. The input is the extracted text data, and the output is tagged text data. For example, natural language processing techniques can be used to identify the titles and paragraphs of each section and assign them tags.

[1387] Step 4:

[1388] Based on the analysis results of the generative AI, the server divides the manual content and formats it into easy-to-understand units. Specifically, it formats each section, highlighting important procedures and notes in bold or color. The input is tagged text data, and the output is formatted text data. For example, the analyzed data can be formatted in an easy-to-read format such as HTML or PDF.

[1389] Step 5:

[1390] A generative AI system generates a video scenario based on formatted text. This scenario includes visual elements and narration text. The input is formatted text data, and the output is video scenario data (e.g., JSON format). For example, for the step "Connect the power cord," a scenario with animation and narration is designed.

[1391] Step 6:

[1392] The server passes the generated scenario to a video generation program (e.g., FFmpeg, OpenToonz), which generates the specific video. The video generation program reads the text, generates animation, and adds effects based on the scenario. The input is the video scenario data, and the output is the generated video file. For example, it synthesizes the necessary animations and narration audio based on each scenario.

[1393] Step 7:

[1394] The user clicks on the preview link provided on their device to view the generated video. If necessary, they can send a correction request. The user's feedback is sent to the server, and specific corrections are recorded in the form of comments. The input is the preview link, and the output is the user's correction request.

[1395] Step 8:

[1396] The server receives the correction request, instructs the generative AI to make the corrections again, and regenerates the video. During the regeneration process, the scenario and video content are adjusted based on user feedback. The input is the correction request and the original scenario data, and the output is a corrected video file. For example, adjustments are made such as adding detailed explanations of specific steps.

[1397] Step 9:

[1398] The emotion engine recognizes the user's emotions and monitors facial expressions and tone of voice while watching. The emotion engine uses machine learning models to detect stress, loss of interest, etc. The input is the user's real-time emotion data, and the output is emotion analysis data.

[1399] Step 10:

[1400] The server receives feedback from the emotion engine and adjusts the video content as needed. For example, if the user looks anxious, it may slow down the commentary or insert additional explanation. The input is emotion analysis data, and the output is an optimized video file.

[1401] Step 11:

[1402] The server stores the final video and provides the user with a final link that they can click to view and download the final adjusted video. The input is the optimized video file, and the output is the final preview link.

[1403] (Application example 2)

[1404] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1405] Conventional procedures and manuals for setting up and maintaining autonomous vehicles are often provided in text format, which often requires time to understand. Furthermore, users may feel anxious or stressed while performing these procedures, which increases the risk of incorrect operation. Furthermore, if users lose interest, they may skip important steps. Therefore, there is a need for visual manuals that are easy for users to understand and that adapt to their emotions.

[1406] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1407] In this invention, the server includes means for users to upload manual files from their terminals, means for the server to receive the uploaded manual files and extract the contents in text format, and means for analyzing the emotions of the user while watching the video through emotion recognition means and adjusting the video content in real time based on the analysis results. This provides a video manual that is easier for users to understand, and the video content is adjusted in real time according to the emotions, making it possible to reduce anxiety and stress in the user and maintain their interest and attention.

[1408] A "user terminal" is an electronic device that a user uses to upload manual files.

[1409] The "server" is a computer system that receives uploaded manual files, extracts the contents in text format, and generates images in cooperation with generative artificial intelligence and video generation programs.

[1410] "Generative AI" is an AI engine that analyzes extracted text, extracts important information and structure from each section, and then generates a video scenario.

[1411] A "video generation program" is software that generates specific images based on a scenario created by generative artificial intelligence.

[1412] "Emotion recognition means" is a technology that analyzes the emotions of a user while watching a video and adjusts the video content in real time based on the analysis results.

[1413] "Text format" refers to a format in which the contents of the manual file are converted into character data.

[1414] A "video scenario" is a specific instruction or guideline for video production generated by generative artificial intelligence.

[1415] A "modification request" is a request made by a user to make changes or additions to the generated video content.

[1416] "Real-time adjustment" refers to the operation of instantly adjusting the video content according to the results of the user's emotion analysis.

[1417] A "manual file" is a document file that contains instructions on how to set up and maintain the device.

[1418] The system based on this invention visualizes the setup and maintenance procedures of an autonomous vehicle, and provides a series of processes that recognize the user's emotions in real time and adjust the contents accordingly. Each element of the system and the processing based on it are specifically described below.

[1419] Program processing

[1420] The system consists of the following main components:

[1421] 1. User device: An electronic device to which the user uploads the settings and maintenance procedures for the autonomous vehicle. The device uses a smartphone (iOS, Android).

[1422] 2. Server: A cloud system that receives uploaded files, extracts the content into text format, and connects it to the generative AI and video generation program. The cloud server uses AWS, Google Cloud, or Microsoft Azure.

[1423] 3. Generative Artificial Intelligence (AI): Analyzes the extracted text, extracts key information and structure from each section, and generates a video scenario. The AI ​​engines used are OpenAI GPT-4 and LLM.

[1424] 4. Video generation program: Generates specific images based on a scenario created by generative AI. The video generation tools used are Adobe After Effects API and FFmpeg.

[1425] 5. Emotion Recognition: Analyzes the emotions of users while they watch a video and adjusts the video content in real time based on the analysis results. Facial and voice recognition technologies use Microsoft Azure Face API and Affectiva SDK.

[1426] Hardware and software used

[1427] Hardware: Smartphone, cloud server.

[1428] software:

[1429] OCR tools (Google Cloud Vision, Tesseract OCR)

[1430] AI engine (OpenAI GPT-4, LLM)

[1431] Video generation tools (Adobe After Effects API, FFmpeg)

[1432] Emotion recognition engine (Microsoft Azure Face API, Affectiva SDK)

[1433] Data processing and calculation

[1434] 1. Uploading manual files: Users upload the configuration files for the autonomous vehicle from their smartphones. The file formats are PDF and Word.

[1435] 2. Text extraction: The cloud server uses OCR technology to convert the content into text format.

[1436] 3. Text analysis: Generative AI analyzes the text and extracts important information such as manual operations and points of caution.

[1437] 4. Video scenario generation: Generative AI generates an easy-to-understand video scenario based on the extracted information.

[1438] 5. Video generation: The video generation program generates specific videos based on the generated scenario. The videos include animations and narration.

[1439] 6. Emotion analysis and real-time adjustment: While the user is watching the video, the emotion recognition means analyzes the user's emotions and adjusts the video content in real time if it detects stress or loss of interest.

[1440] Specific examples

[1441] For example, a specific example of uploading a software update procedure for an autonomous vehicle is shown below.

[1442] Preparation before software update: A scenario and animation of parking the vehicle in a safe place and turning off the engine.

[1443] How to perform the update: Scenario and animation of connecting the USB drive to the vehicle and selecting Update from the Settings menu.

[1444] While the user is watching this video, emotion recognition means monitors facial expressions and vocal tone, and if the video is difficult to understand, detailed explanations are added or the video is slowed down.

[1445] Prompt Sentence Examples

[1446] An example of a prompt sentence to input to the generative AI model is as follows:

[1447] "Amazing results! Take our self-driving vehicle software update manual as an example:

[1448] Park your vehicle in a safe place and turn off the engine.

[1449] Then, plug the USB drive into your vehicle and select Updates from the Settings menu.

[1450] Convert this into a visually compelling video scenario, including any animations or text-to-speech you need."

[1451] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1452] Step 1:

[1453] User uploads manual file

[1454] Input: The user uses the terminal to select the settings and maintenance instructions for the autonomous vehicle (PDF or Word format).

[1455] Action: The user selects a manual file from the file selection dialog and clicks the upload button.

[1456] Output: The selected manual file is sent to the server.

[1457] Step 2:

[1458] The server receives the manual file and extracts the contents into text format.

[1459] Input: Manual files uploaded to the server.

[1460] How it works: The server uses an OCR tool (Google Cloud Vision, Tesseract OCR) to convert the contents of the manual file into text format.

[1461] Output: The extracted text data is passed to the generative AI.

[1462] Step 3:

[1463] Generative AI analyzes text data and extracts important information and structure

[1464] Input: Text data extracted by OCR.

[1465] How it works: Generative AI (OpenAI GPT-4, LLM) analyzes the text and extracts key information (steps, notes, etc.) and structure for each section.

[1466] Output: The parsed information and its structure is returned to the server.

[1467] Step 4:

[1468] The server formats the manual contents into easy-to-understand units based on the analysis results.

[1469] Input: Information and its structure analyzed by the generative artificial intelligence.

[1470] What it does: The server breaks up the information and formats it with text styles for better visibility (e.g., making important points bold).

[1471] Output: The formatted text data is passed to the generative artificial intelligence.

[1472] Step 5:

[1473] Generative AI generates video scenarios based on formatted text

[1474] Input: Formatted text data.

[1475] How it works: Generative AI generates a video scenario that includes visual elements (animations, diagrams, annotations) and narration text.

[1476] Output: The generated video scenario is returned to the server.

[1477] Step 6:

[1478] The server passes the video scenario to the video generation program, which generates the video.

[1479] Input: Generated video scenario.

[1480] How it works: The server uses video generation programs (Adobe After Effects API, FFmpeg) to generate specific videos (animated, with narration).

[1481] Output: The generated video data is provided to the user terminal.

[1482] Step 7:

[1483] The user checks the generated video and requests corrections if necessary.

[1484] Input: Generated video data.

[1485] How it works: The user watches the video on their device, marks any necessary corrections, and sends a correction request to the server.

[1486] Output: The modification request arrives at the server.

[1487] Step 8:

[1488] The server receives the correction request, instructs the generative AI to make the correction again, and regenerates the image.

[1489] Input: A correction request from the user.

[1490] How it works: The server passes a modification request to the generative AI, which then modifies the scenario and regenerates the video using the video generation program.

[1491] Output: The corrected video data is provided to the user terminal.

[1492] Step 9:

[1493] Through emotion recognition, the system analyzes the emotions of users while they are watching a video and adjusts the video content in real time.

[1494] Input: User's facial expressions and voice data.

[1495] How it works: The emotion recognition engine (Microsoft Azure Face API, Affectiva SDK) analyzes the user's emotions from their facial expressions and voice in real time, and adjusts the speed and content of the video if it detects stress or a loss of interest.

[1496] Output: The adjusted video content is reflected on the user's device in real time.

[1497] Step 10:

[1498] The server provides the final completed video to the user.

[1499] Input: Final adjusted video data.

[1500] How it works: The server stores the final video data and provides the final link to the user, through which the user can view and download the video.

[1501] Output: A link will be provided where users can watch and download the final video.

[1502] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1503] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1504] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1505] [Fourth embodiment]

[1506] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1507] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1508] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1509] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1510] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1511] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1512] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1513] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1514] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1515] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1516] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1517] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1518] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1519] The system based on the present invention converts complex manuals into easy-to-understand images. The system program and its processing will be explained below in natural language.

[1520] Basic system configuration

[1521] The system is based on a series of processes in which a user uploads a manual file using a terminal, and the server processes the file to generate an image. The system includes the following elements:

[1522] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[1523] 2. Server: The back-end system that receives and processes uploaded files.

[1524] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[1525] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[1526] Specific processing of the invention

[1527] The specific processing contents of the present invention will be explained below.

[1528] Manual file upload and analysis

[1529] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives and stores the file.

[1530] The server then converts the manual file into text, sometimes using OCR technology. The converted text is then passed to a generative AI that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[1531] Splitting and formatting content

[1532] Based on the results of the generative AI analysis, the server divides the manual into easy-to-understand sections. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[1533] Video scenario generation

[1534] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice prompt such as "Now plug in the power cord" followed by an animation showing the actual connection steps.

[1535] Video generation and preview

[1536] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[1537] Footage review and correction

[1538] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[1539] Provision of final footage

[1540] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[1541] Specific examples

[1542] For example, if a user has a setup manual for a home 3D printer, when the user uploads the manual, the server analyzes the contents and generates a video scenario like the one below.

[1543] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[1544] 2. Video on how to load filament: Instructions and animation on how to "Attach the filament to the spool holder and thread it through the feeder."

[1545] In this way, by visually showing specific operations, the user can easily understand the procedure and proceed with the work.

[1546] The processing flow will be explained below.

[1547] Step 1:

[1548] A user accesses the web application using a terminal and uploads a manual file (PDF, Word, etc.). The upload screen provides a file selection button and a drag-and-drop function.

[1549] Step 2:

[1550] The server receives the uploaded manual file and saves it in a specified folder on the server. It detects the file type and size and stores it as a temporary file in the appropriate format.

[1551] Step 3:

[1552] The server converts the manual files stored on the server into text format. For PDF and image files, OCR technology is used to extract the text, and for Word files, the text is analyzed directly.

[1553] Step 4:

[1554] The server sends the extracted text data to a generative AI to begin analysis. The generative AI divides the text into sections and adds tags such as "headings," "procedures," and "notes."

[1555] Step 5:

[1556] Based on the analysis results received from the generative AI, the server formats the contents of the manual into appropriate units, including emphasizing important points (by bolding or color coding) and rearranging each section.

[1557] Step 6:

[1558] A generative AI system generates a video scenario based on the formatted text. This scenario includes appropriate animations and visual elements for each section. For example, the step "Connect the power cord" includes an animation of the connection action.

[1559] Step 7:

[1560] The server passes the video scenario received from the generative AI to the video generation program, which then generates the specific video. Video generation includes text-to-speech, animation creation, and the addition of effects.

[1561] Step 8:

[1562] The server temporarily saves the generated video and provides a preview link to the user, who can click on this link to view the generated video.

[1563] Step 9:

[1564] If the user checks the video and feels that corrections are necessary, they can input the specific corrections and submit a correction request. The correction request may include pointing out insufficient explanations or incorrect information.

[1565] Step 10:

[1566] The server receives the correction request and again instructs the generative AI to make the correction. The generative AI generates a new scenario based on the corrections, and the video generation program regenerates the corrected video.

[1567] Step 11:

[1568] The server stores the final video and provides the user with a final link that they can click to download and watch the final video.

[1569] This series of steps makes it possible to visualize complex manuals in a way that is easy for users to understand.

[1570] Example 1

[1571] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1572] Conventional manuals are difficult to understand because they are composed only of text. In particular, procedures and precautions that require specialized knowledge cannot be accurately understood by users, which can lead to mistakes and misunderstandings. For this reason, there is a demand for manuals that are easy to understand and can be provided to users in a visual format.

[1573] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1574] In this invention, the server includes: means for a user to upload a manual file from a terminal; means for the server to receive the uploaded manual file and extract the content in text format; means for a generative artificial intelligence to analyze the extracted text and extract important information and structure of each section; means for the server to divide the content of the manual based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; means for the generative artificial intelligence to generate a video scenario based on the formatted text; means for the server to pass the generated scenario to a video generation program and generate a specific video; means for the user to check the generated video and request corrections as necessary; means for the server to receive the correction request, again instruct the generative artificial intelligence to make corrections and regenerate the video; and means for the server to provide the final completed video to the user. This makes it possible to convert complex manuals into easy-to-understand videos.

[1575] The "server" is a central processing unit that receives manual files uploaded by users and processes the contents.

[1576] A "terminal" is a device that a user uses to upload a manual file.

[1577] A "manual file" is a document that describes the procedures and standards for a product or service, and includes formats such as PDF and Word.

[1578] "Text format" refers to character data extracted from a manual file.

[1579] "Generative artificial intelligence (AI)" is an artificial intelligence system that analyzes extracted text, extracts important information, and generates video scenarios.

[1580] "Analysis" is the process by which generative artificial intelligence understands the content of text data and performs tagging and extraction of important information.

[1581] A "section" is a unit in which the contents of the manual are divided into categories.

[1582] A "video scenario" is a plan in which a generative artificial intelligence uses formatted text to create the framework for a video, including visual elements and narration text.

[1583] A "video generation program" is software for generating specific video based on a video scenario.

[1584] A "modification request" is a request made by a user to request that a generated video be modified.

[1585] "Video" refers to visually presented content created by generative artificial intelligence and video generation programs.

[1586] A "preview" is a visual display means that allows a user to check the generated video.

[1587] The system based on this invention converts complex manuals into easy-to-understand images. This system is based on a series of processes in which a user uploads a manual file using a terminal, and a server processes the file to generate an image. The system program and its processing are described below.

[1588] Basic system configuration

[1589] The system includes the following elements:

[1590] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[1591] 2. Server: This is the back-end system that receives and stores uploaded files and generates videos using generative artificial intelligence (AI) and video generation programs. Tesseract OCR is used as the OCR technology.

[1592] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios. This AI engine analyzes the extracted text, tags each section, and extracts key information.

[1593] 4. Video generation program: A program for generating video based on a scenario created by generative AI. Specifically, it uses video processing tools such as FFmpeg.

[1594] Program processing

[1595] Manual file upload and analysis

[1596] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives the file and saves it in a specified directory. The server then converts the manual file into text using Tesseract OCR technology. The converted text is then passed to a generative artificial intelligence (AI) that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[1597] Specific prompt examples:

[1598] Convert the following PDF manual to text and tag each section. Example tags: Heading, Procedures, Notes.

[1599] Splitting and formatting content

[1600] The server then uses the results of the generative AI analysis to break down the manual into easy-to-understand chunks, properly formatting each section and applying text styles to highlight important steps.

[1601] Specific prompt examples:

[1602] Split and format the manual text based on the analysis results, highlighting important steps in bold.

[1603] Video scenario generation

[1604] Based on the formatted text, a generative AI system generates a video scenario for each section, including visual elements (animations, diagrams, annotations) and narrative text.

[1605] Specific prompt examples:

[1606] Please generate a video scenario based on the following text. Please create a scenario that includes visual elements and narration text.

[1607] Video generation and preview

[1608] The generated scenario is passed to a video generation program via the server, which generates specific images. The images undergo processing such as text-to-speech, animation generation, and the addition of effects, and are temporarily saved on the server. The generated images are then provided to the user as a preview link.

[1609] Footage review and correction

[1610] The user can check the generated video using the provided preview link and send correction requests as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[1611] Specific prompt examples:

[1612] Please generate a correction scenario based on the following user request:

[1613] Provision of final footage

[1614] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[1615] Example: If a user has a setup manual for a home 3D printer, the server will analyze the contents and generate specific video scenarios such as "how to connect the power" and "how to load the filament." This allows the user to easily understand the steps and proceed with the work.

[1616] The above is a specific embodiment of the system according to the present invention.

[1617] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1618] Step 1:

[1619] A user uploads a manual file from their terminal. First, the user accesses the web application and clicks the "Upload Manual File" button. Next, the user selects a manual file in PDF or Word format using the file selection dialog and starts uploading. The input is the manual file selected by the user. The output is the uploaded manual file, which is sent to the server and saved.

[1620] Step 2:

[1621] The server receives the uploaded manual file and saves it in a specified directory. Then, it converts the manual file into text format using OCR technology (Tesseract OCR). The input is the saved manual file. The output is the converted text format.

[1622] Step 3:

[1623] The server passes the converted text to a generative artificial intelligence (AI). The AI ​​analyzes the text, tags each section (headings, procedures, precautions, etc.), and extracts important information. The input is the text of the manual. The output is an analysis result in which each section is tagged and important information is extracted. For example, the AI ​​analyzes the text and divides it into sections such as "Connecting the power supply" and "Loading the filament," highlighting each step.

[1624] Example prompt for a generative AI model:

[1625] Convert the following PDF manual to text and tag each section. Example tags: Heading, Procedures, Notes.

[1626] Step 4:

[1627] Based on the analysis results of the generative AI, the server divides the manual's contents into easy-to-understand units and formats them. The server then adjusts the text style accordingly, highlighting important steps in bold. The input is the analysis results. The output is formatted and highlighted text.

[1628] Example prompt for a generative AI model:

[1629] Split and format the manual text based on the analysis results, highlighting important steps in bold.

[1630] Step 5:

[1631] A generative AI generates a video scenario based on formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narration text. The input is the formatted text. The output is a generated video scenario.

[1632] Example prompt for a generative AI model:

[1633] Please generate a video scenario based on the following text. Please create a scenario that includes visual elements and narration text.

[1634] Step 6:

[1635] The server passes the generated scenario to a video generation program, which generates a specific video. The video is processed by text-to-speech, animation generation, and effects addition. The input is the video scenario. The output is the completed video. The video is generated using tools such as FFmpeg. The generated video is temporarily stored on the server, and a preview link is provided to the user.

[1636] Step 7:

[1637] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video. The input is the user's correction request. The output is the corrected video.

[1638] Example prompt for a generative AI model:

[1639] Please generate a correction scenario based on the following user request:

[1640] Step 8:

[1641] The server provides the final completed video to the user. The server generates a link to the final video and sends it to the user, allowing the user to easily watch the video. The input is the completed video. The output is a video link provided to the user.

[1642] The above is a description of the specific operations at each step, and the data processing and data calculation based on the input and output.

[1643] (Application example 1)

[1644] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1645] Machines and equipment installed in existing factories often require complex operating manuals, making it difficult for operators to understand and apply these manuals to actual operations. Furthermore, manuals are typically provided on paper or in simple text format, making them difficult to understand visually. This increases the likelihood of operational errors and reduces production efficiency. The present invention aims to solve these problems and improve operational efficiency and accuracy by providing easy-to-understand, visual instructions for operating machines installed in factories.

[1646] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1647] In this invention, the server includes: means for a user to upload a manual file from a terminal; means for the server to receive the uploaded manual file and extract the contents in text format; means for a generative artificial intelligence to analyze the extracted text and extract important information and structure of each section; means for the server to divide the manual content based on the analysis results of the generative artificial intelligence and format it into easy-to-understand units; means for the generative artificial intelligence to generate a video scenario based on the formatted text; means for the server to pass the generated scenario to a video generation program and generate a specific video; means for the user to check the generated video and request corrections as necessary; means for the server to receive the correction request, again instruct the generative artificial intelligence to make corrections and regenerate the video; means for the server to provide the final completed video to the user; and means for the system to visualize the operating procedures of machines installed in a factory and display the generated video on a display screen to support robot operation, thereby improving the efficiency of robot operation in a factory and enabling accurate understanding of the operating procedures.

[1648] "User" means a person who uses the system to upload and review manual files.

[1649] "Terminal" refers to the device used by the user to upload the manual file.

[1650] A "manual" is a document that describes the operating procedures for a machine or device.

[1651] "File" means a digital document containing the contents of an operation manual.

[1652] "Server" refers to the computer system that receives and processes uploaded manual files.

[1653] "Extraction" refers to the process by which the server extracts the necessary text information from the uploaded file.

[1654] "Generative AI" refers to an AI engine that analyzes extracted text and generates a video scenario.

[1655] "Analysis" refers to the process by which generative artificial intelligence analyzes the content of text.

[1656] A "section" is a part or item of a manual.

[1657] "Important information" refers to operating procedures and precautions that should be particularly emphasized in the manual.

[1658] "Structure" refers to the chapter and paragraph structure of a text.

[1659] "Splitting" refers to the act of the server dividing the manual into smaller units based on the analysis results.

[1660] "Formatting" is the process of making divided text into an easy-to-understand form.

[1661] A "video scenario" refers to the direction and structure for making a video, created by generative artificial intelligence.

[1662] A "video generation program" is software for creating specific video based on a video scenario.

[1663] "Confirmation" refers to the act of the user viewing the generated video.

[1664] A "modification request" is a user request for changes to be made to the generated video.

[1665] "Regeneration" refers to the process by which the server regenerates the video based on the modification request.

[1666] "Providing" refers to the act of the server handing over the final completed video to the user.

[1667] "System" refers to the technical infrastructure that includes the above means and elements and operates as a whole.

[1668] A "factory" is a production facility where machinery and equipment are installed.

[1669] "Machinery" refers to equipment and tools used in a factory.

[1670] An "operating procedure" is a sequence of actions to ensure the correct operation of a machine or device.

[1671] "Visualization" refers to the process of creating a video to visually show operating procedures.

[1672] A "robot" is an automated device that performs tasks in a factory.

[1673] "Operation support" refers to assisting robots to perform their work efficiently.

[1674] "Display screen" means a monitor for displaying the generated image.

[1675] This invention is a system for converting complicated machine or device operating procedures into images that are easy to understand visually. A specific implementation method for this system will be described below.

[1676] Basic system configuration

[1677] The system is based on a series of processes in which a user uploads a manual file using a terminal, and the server processes the file to generate an image. The system includes the following elements:

[1678] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[1679] 2. Server: The back-end system that receives and processes uploaded files.

[1680] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[1681] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[1682] Hardware and Software Used

[1683] Hardware: Tablets and display screens attached to factory robots, and user-operated devices

[1684] software:

[1685] Python 3: Programming Language

[1686] OpenAI API: Used as generative artificial intelligence

[1687] Pillow: Processing images using the Python Image Library

[1688] MoviePy: A video editing library

[1689] Data processing and calculation

[1690] Manual file upload and analysis

[1691] Users upload manual files from their devices through a web application. These manual files are often in PDF or Word format. Once uploaded, the server receives and stores the files. The server then converts the manual files into text format. This conversion may involve the use of OCR technology. The converted text is then passed to a generative artificial intelligence (AI) that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tags such as "headings," "procedures," and "notes."

[1692] Splitting and formatting content

[1693] Based on the results of the generative AI analysis, the server divides the manual into easy-to-understand sections. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[1694] Video scenario generation

[1695] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice prompt such as "Now plug in the power cord" followed by an animation showing the actual connection steps.

[1696] Video generation and preview

[1697] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[1698] Footage review and correction

[1699] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[1700] Provision of final footage

[1701] Finally, the server provides the final completed video to the user, making the complex manual visually easy to understand and allowing the user to easily understand it.

[1702] Examples of concrete examples and prompts

[1703] For example, if a user has an operating manual for a machine installed in a factory, the server will analyze the contents and generate a video scenario like the one below when the user uploads the manual.

[1704] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[1705] 2. Video showing how to replace the filter: "Open the filter cover. Remove the old filter and insert the new filter. Close the filter cover." Instructions and animation.

[1706] Example prompt sentence:

[1707] Please convert the following operating procedures into a video scenario.

[1708] 1. How to connect the power supply:

[1709] Connect the power cord to the back of the machine.

[1710] Plug it into an outlet and press the power switch.

[1711] 2. How to replace the filter:

[1712] Open the filter cover.

[1713] Remove the old filter and insert the new one.

[1714] Close the filter cover.

[1715] In this way, the efficiency and accuracy of robot operations in factories can be improved, and operating procedures can be easily understood by users.

[1716] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1717] Step 1:

[1718] Uploading manual files

[1719] A user accesses the web application using a terminal and uploads a manual file. The uploaded manual file is a digital document in PDF or Word format.

[1720] Input: Manual file (PDF or Word format)

[1721] Output: Manual file saved on the server

[1722] Step 2:

[1723] Converting manual files to text format

[1724] The server receives the uploaded manual file and converts the content into text format, which may involve the use of OCR technology, and extracts the required text information.

[1725] Input: Manual file stored on the server

[1726] Output: Extracted text data

[1727] Step 3:

[1728] Text Analysis

[1729] The server passes the extracted text data to a generative AI model for analysis. The AI ​​analyzes the content of the text, tags it with headings, procedures, notes, etc., and extracts important information from each section.

[1730] Input: Extracted text data

[1731] Output: Analysis results (text section information, key information, tags)

[1732] Step 4:

[1733] Dividing and formatting the manual content

[1734] Based on the analysis results of the generative AI, the server divides the manual's contents into easy-to-understand units. Specifically, each section is appropriately formatted and its level of difficulty is assessed. The text style is then adjusted, with important steps being highlighted in bold.

[1735] Input: Analysis results (text section information, key information, tags)

[1736] Output: Formatted text data

[1737] Step 5:

[1738] Video scenario generation

[1739] Based on the formatted text, a generative AI system generates a video scenario for each section, including visual elements (animations, diagrams, annotations) and narrative text.

[1740] Input: Formatted text data

[1741] Output: Video scenario

[1742] Step 6:

[1743] Generating concrete images

[1744] The server passes the generated scenario to a video generation program, which generates specific images. Video generation includes text-to-speech, animation generation, and the addition of effects.

[1745] Input: Video scenario

[1746] Output: Generated video file

[1747] Step 7:

[1748] Preview and check footage

[1749] The user can check and preview the video generated on the device, check the video content, and send correction requests if necessary.

[1750] Input: Generated video file

[1751] Output: User correction request

[1752] Step 8:

[1753] Processing and regenerating correction requests

[1754] The server receives the user's correction request and again instructs the generative AI to make the correction. A new scenario based on the correction instruction is generated, and the video generation program regenerates the corrected video.

[1755] Input: Correction request, initial video scenario

[1756] Output: Corrected video file

[1757] Step 9:

[1758] Provision of final footage

[1759] The server then provides the final completed video to the user, who can then view the video on a display screen attached to the factory robot and check the operating procedures.

[1760] Input: Modified video file

[1761] Output: The final video delivered to the user

[1762] These steps will improve the efficiency and accuracy of robot operations in factories and make it easier for users to understand the operating procedures.

[1763] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1764] The system based on the present invention converts complex manuals into easy-to-understand images and also combines an emotion engine that recognizes the user's emotions. The system program and its processing will be specifically described below.

[1765] Basic system configuration

[1766] The system is based on a series of processes: a user uploads a manual file using a terminal, a server processes the file to generate a video, and an emotion engine recognizes emotions and adjusts the video when the user watches the generated video. The system includes the following elements:

[1767] 1. User Interface (UI): A web application for users to upload manual files, preview the generated footage, and make correction requests.

[1768] 2. Server: The back-end system that receives and processes uploaded files.

[1769] 3. Generative Artificial Intelligence (AI): An AI engine that analyzes manual text and generates video scenarios.

[1770] 4. Video generation program: A program for generating video based on a scenario created by generative AI.

[1771] 5. Emotion Engine: An engine that recognizes the emotions a user is feeling while watching and adjusts the video content based on those emotions.

[1772] Specific processing of the invention

[1773] The specific processing contents of the present invention will be explained below.

[1774] Manual file upload and analysis

[1775] A user uploads a manual file from their device through a web application. This manual file is often in PDF or Word format. Once uploaded, the server receives and stores the file.

[1776] The server then converts the manual file into text, sometimes using OCR technology. The converted text is then passed to a generative AI that begins analyzing the content. The AI ​​analyzes each section of the manual and extracts key information, along with tagging it with headings, procedures, and notes.

[1777] Splitting and formatting content

[1778] The server then divides the manual into easy-to-understand sections based on the results of the generative AI analysis. Specifically, each section is appropriately formatted and its level of difficulty is assessed. For example, the text style is adjusted to highlight important steps in bold.

[1779] Video scenario generation

[1780] A generative AI system then generates a video scenario for each section based on the formatted text. This scenario includes visual elements (animations, diagrams, annotations) and narrative text. For example, it includes a voice command such as "Connect the power cord" followed by an animation showing the actual connection procedure.

[1781] Video generation and preview

[1782] The generated scenario is passed to a video generation program via the server, which then generates a specific video. The video undergoes processing such as text-to-speech, animation generation, and the addition of effects. The generated video is temporarily saved by the server, and a preview link is provided to the user.

[1783] Footage review and correction

[1784] The user checks the video generated on their device and sends a request for corrections as necessary. When the correction request arrives at the server, the generative AI again generates a scenario based on the correction instructions, and the video generation program regenerates the corrected video.

[1785] Emotion recognition and video adjustment using an emotion engine

[1786] The emotion engine recognizes emotions while the user is watching the video. The emotion engine detects stress, loss of interest, and loss of attention through facial and voice recognition. For example, if the user looks anxious, the emotion engine will slow down the commentary or insert additional explanations.

[1787] If the emotion engine determines that the user's interest is declining, it will insert interactive elements (such as quizzes or multiple choice questions) into the video to attract the user's attention. This adjustment information is sent to the server, and the video generation program updates the video based on that information.

[1788] Provision of final footage

[1789] The server stores the final video optimized by the emotion engine and provides the final link to the user, who can click on it to watch and download the final emotion-adjusted video.

[1790] Specific examples

[1791] For example, if a user has a setup manual for a home 3D printer, when the user uploads the manual, the server analyzes the contents and generates a video scenario like the one below.

[1792] 1. Video showing how to connect the power cord: An animation of connecting the cord is displayed along with the instruction "Connect the power cord and plug it into a power outlet."

[1793] 2. Video on how to load filament: Instructions and animation on how to "Attach the filament to the spool holder and thread it through the feeder."

[1794] While the user is watching the video, the emotion engine monitors their facial expressions and tone of voice, and if it detects stress or anxiety, it will provide additional explanations for the steps or adjust the speed. If interest or disengagement wanes, it will insert an interactive quiz to keep the user's attention.

[1795] Through this series of processes, complex manuals can be optimized to suit the user's emotions and presented in a visually easy-to-understand format.

[1796] The processing flow will be explained below.

[1797] Step 1:

[1798] A user accesses the web application using a terminal and uploads a manual file (PDF, Word, etc.). The upload screen provides a file selection button and a drag-and-drop function.

[1799] Step 2:

[1800] The server receives the uploaded manual file and saves it in a specified folder on the server. It detects the file type and size and stores it as a temporary file in the appropriate format.

[1801] Step 3:

[1802] The server converts the saved manual files into text format, using OCR technology for PDF and image files, and directly analyzing the text for Word files.

[1803] Step 4:

[1804] The server sends the extracted text data to a generative AI to begin analysis. The generative AI divides the text into sections and adds tags such as "headings," "procedures," and "notes."

[1805] Step 5:

[1806] Based on the analysis results received from the generative AI, the server formats the contents of the manual into appropriate units, including emphasizing important points (by bolding or color coding) and rearranging each section.

[1807] Step 6:

[1808] A generative AI system generates a video scenario based on the formatted text. This scenario includes appropriate animations and visual elements for each section. For example, the step "Connect the power cord" includes an animation of the connection action.

[1809] Step 7:

[1810] The server passes the video scenario received from the generative AI to the video generation program, which then generates the specific video. Video generation includes text-to-speech, animation creation, and the addition of effects.

[1811] Step 8:

[1812] The server temporarily saves the generated video and provides a preview link to the user, who can click on this link to view the generated video.

[1813] Step 9:

[1814] If the user checks the video and feels that corrections are necessary, they can input the specific corrections and submit a correction request. The correction request may include pointing out insufficient explanations or incorrect information.

[1815] Step 10:

[1816] The server receives the correction request and again instructs the generative AI to make the correction. The generative AI generates a new scenario based on the corrections, and the video generation program regenerates the corrected video.

[1817] Step 11:

[1818] The emotion engine recognizes emotions while the user is watching the video. The emotion engine detects stress, loss of interest, and loss of attention through facial and voice recognition. For example, if the user looks anxious, the emotion engine will slow down the commentary or insert additional explanations.

[1819] Step 12:

[1820] If the emotion engine determines that the user's interest is declining, it will insert interactive elements (such as quizzes or multiple choice questions) into the video to attract the user's attention. This adjustment information is sent to the server, and the video generation program updates the video based on that information.

[1821] Step 13:

[1822] The server stores the final video optimized by the emotion engine and provides the final link to the user, who can click on it to watch and download the final emotion-adjusted video.

[1823] Through this series of steps, complex manuals are transformed into easy-to-understand images, and the content is optimized according to the user's emotions.

[1824] Example 2

[1825] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1826] Conventional manuals and guides are mainly in text format, which not only makes the content complex and difficult to understand, but also makes it difficult to adjust to the user's emotions and level of understanding. As a result, many users are unable to understand the procedures correctly, which often leads to operational errors and wastes time. Furthermore, with conventional systems, when users request corrections, it is sometimes tedious and the system is unable to respond quickly. A new system is needed to solve these issues.

[1827] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1828] In this invention, the server includes a means for users to upload manual files from their terminals, a means for receiving the uploaded manual files and extracting their contents in text format, and a means for generative AI to analyze the extracted text and extract important information and structure for each section, thereby enabling the generation and immediate modification of images according to the user's emotions and level of understanding.

[1829] A "user" is an entity that uses the system to upload manual files, check the generated video, and request corrections.

[1830] A "terminal" is a device used by a user to upload a manual file, and includes, for example, a computer or a smartphone.

[1831] "Manual files" refer to materials uploaded by users, and include document files such as PDF and Word formats.

[1832] A "server" is a computer system that receives and processes uploaded manual files.

[1833] "Means for extracting content in text format" refers to the process of converting manual files into text data using OCR (Optical Character Recognition) technology.

[1834] "Generative AI" refers to AI technology that analyzes extracted text and extracts key information and structure from each section.

[1835] "Methods for extracting important information and structure" refers to the process of analyzing the content of text, tagging it with "headings," "procedures," "notes," etc., and extracting important information.

[1836] "Means of dividing and arranging the information into easy-to-understand units" refers to the process of appropriately arranging each section based on the analysis results of the generative artificial intelligence, and highlighting it according to its level of difficulty and importance.

[1837] "Means for generating a video scenario" refers to the process of designing a video scenario including visual elements and narrative text based on formatted text.

[1838] "Video generation program" refers to a program for generating specific video based on a generated video scenario, and includes functions such as text reading, animation generation, and adding effects.

[1839] "Preview Link" means a URL or hyperlink provided to a user to allow the user to view the generated video.

[1840] A "modification request" refers to a request submitted by a user to request modification of a generated video.

[1841] An "emotion engine" is an engine that monitors the user's facial expressions and tone of voice and adjusts the video content based on those emotions.

[1842] "Means for adjusting video content based on emotions" refers to a process for dynamically adjusting the speed and content of video explanations based on the user's emotional data detected by the emotion engine.

[1843] The system based on the present invention converts complex manuals into easy-to-understand images and combines an emotion engine that recognizes the user's emotions and adjusts the image content accordingly. This system is based on a series of processes: a user uploads a manual file using a terminal, a server processes the file to generate an image, and when the user watches the generated image, the emotion engine recognizes the emotions and adjusts the image accordingly.

[1844] Hardware and software used

[1845] Hardware

[1846] Device: The device where the user uploads the manual file, for example, a computer, tablet, or smartphone.

[1847] Server: A high-performance computer system for receiving, analyzing, storing, generating and providing images of manual files.

[1848] A camera and microphone are built into the user's device for recognition by the emotion engine.

[1849] software

[1850] Web application: An interface for users to upload files, preview footage, and request revisions. Developed using HTML5, CSS3, and JavaScript.

[1851] OCR technology: Software used to convert uploaded manual files into text format. For example, Tesseract OCR.

[1852] Generative Artificial Intelligence (AI) module: Algorithms for text analysis and scenario generation, incorporating natural language processing technology.

[1853] Video generator: A tool for generating video based on a scenario. Examples include FFmpeg and OpenToonz.

[1854] Emotion Engine: An engine that recognizes the user's emotions and adjusts the video content accordingly. It uses machine learning, facial recognition, and voice recognition technology.

[1855] Specific examples of implementation

[1856] Uploading manual files

[1857] A user uploads a manual file (e.g., PDF) from their device through a web application. The web application provides a file picker function that allows the user to select a file from their local disk.

[1858] File saving and text conversion

[1859] The server receives the uploaded file and stores it in a database. Then, the server converts it into text data using Tesseract OCR. For example, the text of each page of a PDF is extracted.

[1860] Text Analysis and Tagging

[1861] The server passes the converted text data to a generative AI, which analyzes the text, tags it with headings, procedures, notes, and other information, and extracts important information. NLP techniques are used to identify each section and title.

[1862] Video scenario generation

[1863] Based on the analyzed and formatted text, generative AI generates a video scenario for each section. This scenario includes visual elements (animations, diagrams, annotations) and narration text. For example, for the scenario "Connect the power cord," an animation and explanation of how to connect the cord is generated.

[1864] Video generation and preview

[1865] The server passes the scenario to a video generation program, which then generates a specific video. Based on the specific scenario, text is read aloud, animation is generated, and effects are added. The generated video file is temporarily stored on the server, and a preview link is provided to the user.

[1866] Footage review and correction

[1867] The user clicks on the provided preview link to watch the generated video. If necessary, they can submit a correction request, specifying the specific corrections to be made through a feedback form or comment function. Once the correction request is received by the server, the generative AI adjusts the scenario based on the correction instructions, and a new video is regenerated.

[1868] Emotion recognition and video adjustment

[1869] The emotion engine monitors the user's facial expressions and tone of voice via a camera and microphone while the user is watching the video. For example, if the user frowns or sighs, this is recognized. The emotion engine analyzes the user's emotional data and sends instructions to the server as needed to adjust the speed and content of the video explanation.

[1870] Prompt Sentence Examples

[1871] For example, the following prompt is passed to the generative AI:

[1872] "Generate a video scenario that shows the user how to connect a power cord. The instructions should include animation of the cord being connected and narration emphasizing key points."

[1873] The specific embodiment described above allows the user to easily understand a complex manual and to make immediate corrections and adjustments as necessary.

[1874] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1875] Step 1:

[1876] The user uploads the manual file from the terminal through the web application. The user uses the file picker function to upload the selected file from the local disk. The input is the manual file in PDF or Word format, and the output is the file uploaded to the server.

[1877] Step 2:

[1878] The server receives the uploaded manual file and stores it in a database. Then, the server converts the manual file into a text format using OCR technology (e.g., Tesseract OCR). The input is the saved file, and the output is the extracted text data. For example, the text content read from each page of a PDF is aggregated into a single text file.

[1879] Step 3:

[1880] The server inputs the converted text data into generative artificial intelligence (AI). The AI ​​analyzes the text, tags it with headings, procedures, notes, and other information, and extracts important information. The input is the extracted text data, and the output is tagged text data. For example, natural language processing techniques can be used to identify the titles and paragraphs of each section and assign them tags.

[1881] Step 4:

[1882] Based on the analysis results of the generative AI, the server divides the manual content and formats it into easy-to-understand units. Specifically, it formats each section, highlighting important procedures and notes in bold or color. The input is tagged text data, and the output is formatted text data. For example, the analyzed data can be formatted in an easy-to-read format such as HTML or PDF.

[1883] Step 5:

[1884] A generative AI system generates a video scenario based on formatted text. This scenario includes visual elements and narration text. The input is formatted text data, and the output is video scenario data (e.g., JSON format). For example, for the step "Connect the power cord," a scenario with animation and narration is designed.

[1885] Step 6:

[1886] The server passes the generated scenario to a video generation program (e.g., FFmpeg, OpenToonz), which generates the specific video. The video generation program reads the text, generates animation, and adds effects based on the scenario. The input is the video scenario data, and the output is the generated video file. For example, it synthesizes the necessary animations and narration audio based on each scenario.

[1887] Step 7:

[1888] The user clicks on the preview link provided on their device to view the generated video. If necessary, they can send a correction request. The user's feedback is sent to the server, and specific corrections are recorded in the form of comments. The input is the preview link, and the output is the user's correction request.

[1889] Step 8:

[1890] The server receives the correction request, instructs the generative AI to make the corrections again, and regenerates the video. During the regeneration process, the scenario and video content are adjusted based on user feedback. The input is the correction request and the original scenario data, and the output is a corrected video file. For example, adjustments are made such as adding detailed explanations of specific steps.

[1891] Step 9:

[1892] The emotion engine recognizes the user's emotions and monitors facial expressions and tone of voice while watching. The emotion engine uses machine learning models to detect stress, loss of interest, etc. The input is the user's real-time emotion data, and the output is emotion analysis data.

[1893] Step 10:

[1894] The server receives feedback from the emotion engine and adjusts the video content as needed. For example, if the user looks anxious, it may slow down the commentary or insert additional explanation. The input is emotion analysis data, and the output is an optimized video file.

[1895] Step 11:

[1896] The server stores the final video and provides the user with a final link that they can click to view and download the final adjusted video. The input is the optimized video file, and the output is the final preview link.

[1897] (Application example 2)

[1898] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1899] Conventional procedures and manuals for setting up and maintaining autonomous vehicles are often provided in text format, which often requires time to understand. Furthermore, users may feel anxious or stressed while performing these procedures, which increases the risk of incorrect operation. Furthermore, if users lose interest, they may skip important steps. Therefore, there is a need for visual manuals that are easy for users to understand and that adapt to their emotions.

[1900] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1901] In this invention, the server includes means for users to upload manual files from their terminals, means for the server to receive the uploaded manual files and extract the contents in text format, and means for analyzing the emotions of the user while watching the video through emotion recognition means and adjusting the video content in real time based on the analysis results. This provides a video manual that is easier for users to understand, and the video content is adjusted in real time according to the emotions, making it possible to reduce anxiety and stress in the user and maintain their interest and attention.

[1902] A "user terminal" is an electronic device that a user uses to upload manual files.

[1903] The "server" is a computer system that receives uploaded manual files, extracts the contents in text format, and generates images in cooperation with generative artificial intelligence and video generation programs.

[1904] "Generative AI" is an AI engine that analyzes extracted text, extracts important information and structure from each section, and then generates a video scenario.

[1905] A "video generation program" is software that generates specific images based on a scenario created by generative artificial intelligence.

[1906] "Emotion recognition means" is a technology that analyzes the emotions of a user while watching a video and adjusts the video content in real time based on the analysis results.

[1907] "Text format" refers to a format in which the contents of the manual file are converted into character data.

[1908] A "video scenario" is a specific instruction or guideline for video production generated by generative artificial intelligence.

[1909] A "modification request" is a request made by a user to make changes or additions to the generated video content.

[1910] "Real-time adjustment" refers to the operation of instantly adjusting the video content according to the results of the user's emotion analysis.

[1911] A "manual file" is a document file that contains instructions on how to set up and maintain the device.

[1912] The system based on this invention visualizes the setup and maintenance procedures of an autonomous vehicle, and provides a series of processes that recognize the user's emotions in real time and adjust the contents accordingly. Each element of the system and the processing based on it are specifically described below.

[1913] Program processing

[1914] The system consists of the following main components:

[1915] 1. User device: An electronic device to which the user uploads the settings and maintenance procedures for the autonomous vehicle. The device uses a smartphone (iOS, Android).

[1916] 2. Server: A cloud system that receives uploaded files, extracts the content into text format, and connects it to the generative AI and video generation program. The cloud server uses AWS, Google Cloud, or Microsoft Azure.

[1917] 3. Generative Artificial Intelligence (AI): Analyzes the extracted text, extracts key information and structure from each section, and generates a video scenario. The AI ​​engines used are OpenAI GPT-4 and LLM.

[1918] 4. Video generation program: Generates specific images based on a scenario created by generative AI. The video generation tools used are Adobe After Effects API and FFmpeg.

[1919] 5. Emotion Recognition: Analyzes the emotions of users while they watch a video and adjusts the video content in real time based on the analysis results. Facial and voice recognition technologies use Microsoft Azure Face API and Affectiva SDK.

[1920] Hardware and software used

[1921] Hardware: Smartphone, cloud server.

[1922] software:

[1923] OCR tools (Google Cloud Vision, Tesseract OCR)

[1924] AI engine (OpenAI GPT-4, LLM)

[1925] Video generation tools (Adobe After Effects API, FFmpeg)

[1926] Emotion recognition engine (Microsoft Azure Face API, Affectiva SDK)

[1927] Data processing and calculation

[1928] 1. Uploading manual files: Users upload the configuration files for the autonomous vehicle from their smartphones. The file formats are PDF and Word.

[1929] 2. Text extraction: The cloud server uses OCR technology to convert the content into text format.

[1930] 3. Text analysis: Generative AI analyzes the text and extracts important information such as manual operations and points of caution.

[1931] 4. Video scenario generation: Generative AI generates an easy-to-understand video scenario based on the extracted information.

[1932] 5. Video generation: The video generation program generates specific videos based on the generated scenario. The videos include animations and narration.

[1933] 6. Emotion analysis and real-time adjustment: While the user is watching the video, the emotion recognition means analyzes the user's emotions and adjusts the video content in real time if it detects stress or loss of interest.

[1934] Specific examples

[1935] For example, a specific example of uploading a software update procedure for an autonomous vehicle is shown below.

[1936] Preparation before software update: A scenario and animation of parking the vehicle in a safe place and turning off the engine.

[1937] How to perform the update: Scenario and animation of connecting the USB drive to the vehicle and selecting Update from the Settings menu.

[1938] While the user is watching this video, emotion recognition means monitors facial expressions and vocal tone, and if the video is difficult to understand, detailed explanations are added or the video is slowed down.

[1939] Prompt Sentence Examples

[1940] An example of a prompt sentence to input to the generative AI model is as follows:

[1941] "Amazing results! Take our self-driving vehicle software update manual as an example:

[1942] Park your vehicle in a safe place and turn off the engine.

[1943] Then, plug the USB drive into your vehicle and select Updates from the Settings menu.

[1944] Convert this into a visually compelling video scenario, including any animations or text-to-speech you need."

[1945] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1946] Step 1:

[1947] User uploads manual file

[1948] Input: The user uses the terminal to select the settings and maintenance instructions for the autonomous vehicle (PDF or Word format).

[1949] Action: The user selects a manual file from the file selection dialog and clicks the upload button.

[1950] Output: The selected manual file is sent to the server.

[1951] Step 2:

[1952] The server receives the manual file and extracts the contents into text format.

[1953] Input: Manual files uploaded to the server.

[1954] How it works: The server uses an OCR tool (Google Cloud Vision, Tesseract OCR) to convert the contents of the manual file into text format.

[1955] Output: The extracted text data is passed to the generative AI.

[1956] Step 3:

[1957] Generative AI analyzes text data and extracts important information and structure

[1958] Input: Text data extracted by OCR.

[1959] How it works: Generative AI (OpenAI GPT-4, LLM) analyzes the text and extracts key information (steps, notes, etc.) and structure for each section.

[1960] Output: The parsed information and its structure is returned to the server.

[1961] Step 4:

[1962] The server formats the manual contents into easy-to-understand units based on the analysis results.

[1963] Input: Information and its structure analyzed by the generative artificial intelligence.

[1964] What it does: The server breaks up the information and formats it with text styles for better visibility (e.g., making important points bold).

[1965] Output: The formatted text data is passed to the generative artificial intelligence.

[1966] Step 5:

[1967] Generative AI generates video scenarios based on formatted text

[1968] Input: Formatted text data.

[1969] How it works: Generative AI generates a video scenario that includes visual elements (animations, diagrams, annotations) and narration text.

[1970] Output: The generated video scenario is returned to the server.

[1971] Step 6:

[1972] The server passes the video scenario to the video generation program, which generates the video.

[1973] Input: Generated video scenario.

[1974] How it works: The server uses video generation programs (Adobe After Effects API, FFmpeg) to generate specific videos (animated, with narration).

[1975] Output: The generated video data is provided to the user terminal.

[1976] Step 7:

[1977] The user checks the generated video and requests corrections if necessary.

[1978] Input: Generated video data.

[1979] How it works: The user watches the video on their device, marks any necessary corrections, and sends a correction request to the server.

[1980] Output: The modification request arrives at the server.

[1981] Step 8:

[1982] The server receives the correction request, instructs the generative AI to make the correction again, and regenerates the image.

[1983] Input: A correction request from the user.

[1984] How it works: The server passes a modification request to the generative AI, which then modifies the scenario and regenerates the video using the video generation program.

[1985] Output: The corrected video data is provided to the user terminal.

[1986] Step 9:

[1987] Through emotion recognition, the system analyzes the emotions of users while they are watching a video and adjusts the video content in real time.

[1988] Input: User's facial expressions and voice data.

[1989] How it works: The emotion recognition engine (Microsoft Azure Face API, Affectiva SDK) analyzes the user's emotions from their facial expressions and voice in real time, and adjusts the speed and content of the video if it detects stress or a loss of interest.

[1990] Output: The adjusted video content is reflected on the user's device in real time.

[1991] Step 10:

[1992] The server provides the final completed video to the user.

[1993] Input: Final adjusted video data.

[1994] How it works: The server stores the final video data and provides the final link to the user, through which the user can view and download the video.

[1995] Output: A link will be provided where users can watch and download the final video.

[1996] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1997] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="...

Claims

1. A means for a user to upload a manual file from a terminal; A server receives the uploaded manual file and extracts the contents in text format; A generative AI analyzes the extracted text and extracts the key information and structure of each section; and The server divides the contents of the manual based on the analysis results of the generative AI and formats them into easy-to-understand units. A means for a generative artificial intelligence to generate a video scenario based on the formatted text; A means for the server to pass the generated scenario to a video generation program and generate specific video; A means for the user to check the generated video and request corrections if necessary; A means for the server to receive the correction request, instruct the generative AI to make the correction again, and regenerate the image; A means for the server to provide the final completed video to the user; A system including:

2. 2. The system of claim 1, wherein the generative artificial intelligence includes means for determining the difficulty level of the analyzed text and for formatting the text in a style that emphasizes important points.

3. 2. The system according to claim 1, wherein the video generation program includes means for reading text to speech, generating animation, and adding effects.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A