system
An AI-driven system analyzes video data to automatically generate procedure manuals, addressing inefficiencies and errors in manual methods by providing accurate, timely, and efficient manual creation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
Conventional manual creation methods for procedure manuals are time-consuming, prone to human error, and require frequent updates, leading to inefficiencies and additional costs.
A system utilizing an AI engine to analyze video data, extract important actions and objects, and generate procedure manuals in natural language, formatted for delivery in PDF, thereby automating the process and reducing manual effort.
The system significantly reduces time and effort in creating procedure manuals, ensures accuracy, and enables rapid updates by automating the generation and delivery process.
Smart Images

Figure 2026036317000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional manual creation methods are primarily manual, requiring significant time and effort. Furthermore, manual updates are required each time the situation changes, resulting in additional costs and time. Furthermore, creating manuals manually carries the risk of human error. The present invention aims to solve these problems and provide a method for creating manuals efficiently and accurately. [Means for solving the problem]
[0005] The present invention uses an AI engine that acquires video data, analyzes it, and extracts important actions and objects. It provides a system that generates a procedure manual in natural language based on the extracted data and outputs the generated procedure manual in a specified format. Furthermore, the system transmits the video data to a server via the Internet, preprocesses the received video data, and passes it to the AI analysis engine, thereby improving work efficiency and accuracy. When generating the procedure manual, the system formats the data according to a predetermined template, and outputs the generated procedure manual in PDF format for delivery to the user. This series of processes significantly reduces time and effort, enabling the rapid provision of always-up-to-date manuals.
[0006] "Video data" refers to data that includes visual information captured by a camera or recording device.
[0007] An "AI engine" is an algorithm and program that uses artificial intelligence to analyze data and extract and process specific information.
[0008] An "action" is an element that refers to a movement or operation recognized within video data.
[0009] An "object" is an element that refers to a substance or target that exists within video data.
[0010] A "procedure manual" is a document that describes in natural language the steps and operating methods for performing a specific task.
[0011] "Preprocessing" refers to the process of organizing and converting video data before it is analyzed by an AI engine.
[0012] A "template" refers to a pre-prepared format or structure used when formatting procedures or documents.
[0013] "PDF format" is an abbreviation for Portable Document Format, a standard file format for electronically representing documents.
[0014] "User" refers to the entity that uses this system to record video data and generate manuals. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] As an embodiment of the present invention, a system including the following processes is provided.
[0037] System Overview
[0038] The user records the work content with a camera, and the video data is received by the terminal and sent to the server. The server analyzes the received video data and extracts important actions and objects. Then, based on the extracted data, a procedure manual is generated in natural language and output in a specified format. A specific embodiment is shown below.
[0039] Program Overview
[0040] 1. Recording Stage
[0041] User: Record the work process with a camera. Specifically, the user records the entire process from start to finish using a camera-equipped smartphone or dedicated recording device.
[0042] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, the video data is automatically saved to the internal storage.
[0043] 2. Uploading videos
[0044] Device: Sends saved video data to the server via the Internet. The device sends saved video data to the specified server upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to inform the user of progress.
[0045] 3. Video Reception and Preprocessing
[0046] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. The server checks the consistency of the video data and performs a simple check to see if there are any missing or corrupted data.
[0047] 4. Video Analysis
[0048] Server: Sends video data to the AI analysis engine. The server divides the stored video data into frames and performs pre-processing to extract the important parts of the frames.
[0049] Server: The AI analysis engine analyzes the video data. It performs frame-by-frame object recognition, motion detection, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[0050] 5. Text Generation
[0051] Server: Generates a procedure manual in natural language based on the extracted important actions and objects. The server analyzes the action list extracted by AI and generates procedures such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[0052] Server: Formats the generated text data into a manual. Formats the text according to the manual template and divides it into sections.
[0053] 6. Manual Output
[0054] Server: Generates the finished manual in PDF or other specified formats. Converts the entire manual to PDF and adds tables, diagrams, screenshots, etc. (if needed).
[0055] Server: Sends the generated manual to the user's device. The completed manual is sent to the user as an email attachment or a dedicated download link is provided.
[0056] Specific examples
[0057] As a specific example, a situation in which an operating manual for a machine in a factory is to be created will be described.
[0058] 1. Recording Stage
[0059] User: A factory worker uses his smartphone to film how to operate a new machine.
[0060] Device: Your smartphone will save this footage to its internal storage.
[0061] 2. Uploading videos
[0062] Device: The smartphone uses Wi-Fi to upload video data to a designated server.
[0063] 3. Video Reception and Preprocessing
[0064] Server: The server receives the video data and performs preprocessing for AI analysis.
[0065] 4. Video Analysis
[0066] Server: The server sends the video data to an AI analysis engine, which recognizes and lists the actions in each frame.
[0067] 5. Text Generation
[0068] Server: Based on the extracted actions, generate a procedure manual in natural language and format it into a manual.
[0069] 6. Manual Output
[0070] Server: Generates the completed manual in PDF format and sends it to the worker via email. This allows users to easily create work manuals and quickly update them.
[0071] This significantly reduces time and effort, and enables the latest manuals to be provided promptly.
[0072] The processing flow will be explained below.
[0073] Step 1:
[0074] Users record their work with a camera. Specifically, they acquire a smartphone with a camera or a dedicated recording device and record the entire process from start to finish.
[0075] Step 2:
[0076] The device will temporarily store the recorded video data. After the recording device or smartphone has finished recording, the video data will automatically be saved to its internal storage.
[0077] Step 3:
[0078] The device sends the stored video data to the server via the Internet. The device sends the stored video data to the specified server's upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[0079] Step 4:
[0080] The server saves the received video data in the server's storage. The server organizes and saves the received data in a dedicated directory. It checks the consistency of the video data and performs a simple check to see if there are any missing or damaged data.
[0081] Step 5:
[0082] The server sends the video data to an AI analysis engine, which divides the stored video data into frames and performs preprocessing to extract the important parts of the frames.
[0083] Step 6:
[0084] The server's AI analysis engine analyzes the video data, performing frame-by-frame object recognition, motion detection, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[0085] Step 7:
[0086] The server generates a procedure manual in natural language based on the extracted important actions and objects. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[0087] Step 8:
[0088] The server formats the generated text data into a manual, formatting the text according to the manual template and dividing it into sections.
[0089] Step 9:
[0090] The server generates the finished manual in PDF or other specified format, converting the entire manual to PDF and adding tables, diagrams, screenshots, etc. (if needed).
[0091] Step 10:
[0092] The server sends the generated manual to the user's device, either by attaching the completed manual to an email or by providing a dedicated download link.
[0093] Example 1
[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0095] There has been no system that can easily record video data of work content and automatically generate procedure manuals based on that data. In particular, there is a need for an efficient and accurate process that can analyze the video data, extract important actions and objects, and generate procedure manuals. This has led to the need for a system that can reduce work hours and labor, standardize work procedures, and enable rapid update support.
[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0097] In this invention, the server includes: [means for acquiring video data;] [means for extracting important actions and objects using an AI analysis unit that analyzes the acquired video data; and] [means for generating a procedure manual in natural language based on the extracted data. This makes it possible to automatically generate a procedure manual based on the video data of the work content and output it in a specified format.
[0098] "Video data" refers to digital information in the form of moving images captured using a camera or other imaging device.
[0099] A "server" is a computing device used to receive, analyze, store, and transmit data over a network.
[0100] "AI Analysis Unit" means a collection of software and hardware that implements artificial intelligence technology used to analyze video data and identify and extract significant actions and objects.
[0101] "Critical actions" refer to actions or operations that require special attention when carrying out business procedures.
[0102] "Objects" refer to specific items or devices that exist within the video data and are required to explain business procedures.
[0103] A "procedure manual" is a document that explains business procedures in a specific and orderly manner as a series of steps.
[0104] "Means for acquiring" refers to methods and devices for collecting video data using cameras and recording devices.
[0105] "Means for generating in natural language" refers to the method and process for generating a procedure manual in a natural language that is easy for humans to understand, based on the extracted data.
[0106] A "prescribed format" is a rule that ensures that procedures are organized in a certain format and displayed in a way that is visually easy to understand.
[0107] As an embodiment of this invention, we provide a system that automatically captures and analyzes work procedures at companies or work sites and generates procedure manuals. Specifically, this system allows users to record work content with a camera, and the video data is received by a terminal and sent to a server. The details of the system are described below.
[0108] Recording Stage
[0109] User:
[0110] Work is recorded with a camera. Specifically, the user uses a camera-equipped smartphone or a dedicated recording device to record work from start to finish. For example, a factory worker uses a smartphone to record how to operate a new machine.
[0111] Device:
[0112] Recorded video data is saved to the internal storage. After recording is complete, the smartphone automatically saves the video data temporarily.
[0113] Video upload
[0114] Device:
[0115] The saved video data is sent to a server via the Internet. The device uploads the video data to the specified server via Wi-Fi or a mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[0116] Video reception and preprocessing
[0117] server:
[0118] The server stores the received video data in a dedicated directory, checks for missing or corrupted data, and checks the data for consistency. Organizes the video data into batches for analysis.
[0119] Video Analysis
[0120] server:
[0121] The server divides the stored video data into frames and sends them to the AI analysis unit. As preprocessing, the video resolution is adjusted appropriately and unnecessary frames are removed. The AI analysis unit (e.g., TENSORFLOW (registered trademark) or OpenCV) is used to perform object recognition and motion detection in each frame, and convert text and voice as needed. For example, scenes in which the user presses a button or pulls a lever are extracted and listed.
[0122] Text Generation
[0123] server:
[0124] Based on the extracted important actions and objects, a natural language instruction manual is generated. An AI model (e.g., GPT-3 (registered trademark)) is used to create steps such as "1. Turn on the power" and "2. Select the settings menu" from the listed actions. Appropriate explanations and cautions are added to each step. For example, an explanation such as "Check the power cord and be careful not to press the buttons too hard" is added.
[0125] server:
[0126] The generated text data is formatted into a manual, with each step divided into sections and visually organized using bulleted and numbered lists.
[0127] Manual Output
[0128] server:
[0129] Generate the finished manual in PDF or other format of your choice. Use tools such as Adobe Acrobat or LaTeX to add tables, charts, and screenshots to create a visually appealing manual.
[0130] server:
[0131] The generated manual is sent to the user's device. The completed manual is sent as an email attachment to the user, or a dedicated download link is provided.
[0132] Specific examples
[0133] As a specific example, we will explain a situation in which an operating manual for a new machine is being created in a factory.
[0134] User:
[0135] Factory workers use their smartphones to record themselves operating a new machine.
[0136] Device:
[0137] The smartphone saves this footage to its internal storage.
[0138] Device:
[0139] The smartphone uses Wi-Fi to upload the video data to a designated server.
[0140] server:
[0141] The server receives the video data and performs preprocessing for AI analysis.
[0142] server:
[0143] The server sends the video data to an AI analysis unit, which recognizes and lists the movements in each frame.
[0144] server:
[0145] Based on the extracted actions, a procedure manual is generated in natural language and formatted into a manual.
[0146] server:
[0147] The completed manual is generated in PDF format and sent to the worker via email, allowing users to easily create work manuals and quickly update them.
[0148] Examples of prompt statements
[0149] "Please film how to operate the new machine and create a manual."
[0150] "Please tell me the procedure for uploading video data to the server."
[0151] "Please explain in detail how to generate the procedure manual."
[0152] This will provide a system that standardizes business procedures, improves efficiency, and enables rapid update support.
[0153] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0154] Step 1:
[0155] User: Records work with a camera. The user uses a smartphone with a camera or other recording device to record the entire process from start to finish. For example, a factory worker uses a smartphone to record how to operate a new machine. The user takes care to clearly capture specific operating steps (e.g., how to press a button, pull a lever). Input: The device used is a smartphone with a camera. Output: Recorded video data.
[0156] Step 2:
[0157] Device: After recording is complete, the device saves the video data to its internal storage. At the same time, the device sends the video data to the server via Wi-Fi or the mobile network. During uploading, the device displays a progress bar to inform the user of the transfer status. Input: Recorded video data. Output: Video data sent to the server.
[0158] Step 3:
[0159] Server: The server stores the received video data in a dedicated directory in an organized manner. It first checks the data for consistency and performs a simple check to see if there are any missing or corrupted data. If the data is OK, it is broken down into frames for analysis and the resolution is adjusted if necessary. Input: Video data sent to the server. Output: Video data that has been split into frames and pre-processed.
[0160] Step 4:
[0161] Server: Sends preprocessed video data to an AI analysis unit (e.g., TensorFlow or OpenCV), which performs object recognition and action detection for each frame, and converts text and voice as needed. For example, it extracts and lists scenes in which a button is pressed or a lever is pulled. Input: Preprocessed video data for each frame. Output: A list of extracted important actions and objects.
[0162] Step 5:
[0163] Server: Based on the extracted important actions and objects, a generative AI model (e.g., GPT-3) is used to generate instructions in natural language. Detailed explanations and precautions are added for each action (e.g., pressing a button), and steps such as "1. Turn on the power" and "2. Select the settings menu" are formed. Input: A list of extracted actions and objects. Output: Instructions generated in natural language.
[0164] Step 6:
[0165] Server: Formats the generated procedure manual into a manual format and outputs it in PDF or other specified format. Tables, diagrams, and in some cases screenshots are added to the procedure manual to make it visually easy to understand. Input: Procedure manual generated in natural language. Output: Procedure manual formatted in PDF or other specified format.
[0166] Step 7:
[0167] Server: Provides the completed procedure to the user. The server sends the procedure as an email attachment or provides a dedicated download link. It may also store it in a database. Input: Procedure formatted in PDF or other specified format. Output: Procedure provided to the user.
[0168] Through this series of processes, a system is realized that can automatically generate procedure manuals from video data of work content and provide them to users quickly and efficiently, making it possible to standardize work procedures and respond to updates quickly.
[0169] (Application example 1)
[0170] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0171] In modern factories and manufacturing sites, specialized robot maintenance work is complex, and creating procedure manuals requires a great deal of time and effort. Therefore, there is a need for an effective automatic procedure manual generation system to improve the efficiency and standardization of maintenance work. However, existing manual procedure manual creation methods have issues such as difficulty in recording and analyzing work, and difficulty in quickly reflecting the latest information.
[0172] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0173] In this invention, the server includes: [means for acquiring video data using a device with a camera; [means for extracting important actions and objects using an AI analysis engine that analyzes the acquired video data; and [means for generating a procedure manual in natural language based on the extracted data.] This makes it possible [to achieve work efficiency and standardization, and to quickly reflect the latest information in the procedure manual].
[0174] "Video data" is a series of image information captured by a device with a camera, and is digital data that includes details of actions and objects.
[0175] "AI analysis engine" is a general term for software and hardware that uses artificial intelligence technology to analyze video data and extract important actions and objects.
[0176] A "procedure" is a document that provides step-by-step instructions for performing a specific task or operation.
[0177] "Natural language" is a language used by humans on a daily basis and a set of textual forms generated by computer programs.
[0178] A "generative AI model" is an artificial intelligence model that learns large amounts of data in advance to generate instruction manuals and other text data.
[0179] A "prompt" is a series of phrases or documents that are input to a generative AI model to determine the content and format of the generated text.
[0180] A "server" is a computer that receives requests from client devices over a network and provides data processing and data storage functions.
[0181] "Preprocessing" refers to the process of removing noise from video data and converting its format before analysis by the AI analysis engine.
[0182] A "data processing device" is a computer system for analyzing video data and extracting important information.
[0183] A "template" is a predefined format or style guide for formatting procedures and reports.
[0184] "Digital document format" means a document format that can be stored and displayed electronically, including PDF.
[0185] MODE FOR CARRYING OUT THE INVENTION
[0186] System Overview
[0187] As an embodiment of the present invention, we provide a system that performs the following process. In this system, a user records work content using a camera-equipped smartphone and sends the video data to a server. The server analyzes the received video data and extracts important actions and objects. Then, based on the extracted data, it generates a procedure manual in natural language and outputs it in a specified format.
[0188] Explanation of program processing
[0189] Hardware and software used
[0190] Smartphone: A smart device used as a recording device.
[0191] Server: A computer used for data analysis that receives requests from client devices over the network and processes the data.
[0192] OpenCV: A library for recording and processing video.
[0193] Requests library: Used to send video data to the server.
[0194] AI analysis engine: An engine for analyzing video data and extracting important actions and objects.
[0195] PDF generation library: A library for converting instructions into PDF format (e.g., ReportLab).
[0196] Processing flow
[0197] 1. Recording: The user records the robot's maintenance work using a smartphone. The recorded data is saved in the smartphone's internal storage.
[0198] 2. Video Upload: Once the recording is complete, the device (smartphone) sends the saved video data to the server via Wi-Fi or mobile network. This data transmission is performed using the Requests library.
[0199] 3. Data processing: The server receives the transmitted video data and performs preprocessing, which includes noise reduction and format conversion. The AI analysis engine then divides the video data into frames and extracts important actions and objects.
[0200] 4. Procedure generation: A procedure manual is generated in natural language based on the data extracted by the AI analysis engine. A generative AI model is used here. A prompt sentence for generating the procedure manual is also created. By inputting this prompt sentence into the generative AI model, automatically generated text can be obtained.
[0201] 5. Output of procedure manual: The generated procedure manual is output in PDF format using a PDF generation library, and the output PDF is provided to the user.
[0202] Specific examples
[0203] For example, let's say a maintenance task of adding lubricant to the joints of a robot is recorded in a factory. This recorded data is sent to a server, which then uses an analysis engine to analyze the video and extract important actions (such as "remove the cover," "prepare the lubricant," and "apply lubricant to the specified location"). Based on the extracted actions, the AI model automatically generates a procedure manual and provides it to the engineer in PDF format.
[0204] Prompt Sentence Examples
[0205] Video data URL: {video_url}
[0206] Analysis output:
[0207] Object Recognition
[0208] Action Detection
[0209] Procedure generation
[0210] The above is an embodiment of the present invention. This system makes it possible to improve the efficiency and standardize maintenance work, and to quickly update the latest information in the procedure manual.
[0211] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0212] Program processing flow
[0213] Step 1:
[0214] Recording
[0215] The user records the robot's maintenance work using a smartphone. At this time, a camera-equipped device is used to capture the entire process from start to finish. The recorded data (video data) is saved in the smartphone's internal storage.
[0216] Input: Maintenance work video
[0217] Output: Video data stored in the internal storage
[0218] Step 2:
[0219] Video upload
[0220] Once recording is complete, the device (smartphone) sends the saved video data to the server via the Internet. The device uses the Requests library to send the video data to the specified server's upload endpoint. At this time, the file transfer status is displayed to inform the user of the progress.
[0221] Input: Video data stored in the internal storage
[0222] Output: Video data sent to the server
[0223] Step 3:
[0224] Data reception and preprocessing
[0225] The server receives the transmitted video data and stores it in storage. Upon receiving the data, it performs pre-processing such as noise removal and format conversion. This process converts the data into a format that can be smoothly processed by the analysis engine.
[0226] Input: Video data sent to the server
[0227] Output: Pre-processed video data
[0228] Step 4:
[0229] Video Analysis
[0230] The preprocessed video data is passed to the server's AI analysis engine, which divides the video data into frames and extracts important actions and objects from each frame. Here, it performs tasks such as object recognition and action detection for each frame, and creates a list of important actions.
[0231] Input: Preprocessed video data
[0232] Output: Data listing important actions and objects
[0233] Step 5:
[0234] Procedure generation
[0235] Based on the data listing important actions and objects, the server uses a generative AI model to generate instructions in natural language. A prompt is entered into the generative AI model, which then automatically generates text. An example of a prompt might be something like "Video data URL: {video_url}\nAnalysis output:\n- Object recognition\n- Action detection\n- Instruction manual generation."
[0236] Input: Data listing important actions and objects, prompts
[0237] Output: Natural language instructions
[0238] Step 6:
[0239] Output of procedure manual
[0240] The server converts the generated instructions into PDF format using a PDF generation library, and the generated PDF is sent to the user's device or a download link is provided.
[0241] Input: Natural language instructions
[0242] Output: Instructions in PDF format
[0243] By following the above steps, a procedure manual can be automatically generated from a series of maintenance tasks and provided efficiently.
[0244] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0245] As an embodiment of the present invention, we provide a system that includes the following processing. In addition to acquiring and analyzing video data, this system recognizes the user's emotions and reflects them in the generation of procedure manuals, thereby providing more personalized manuals.
[0246] System Overview
[0247] The user records their work with a camera, and the video data is received by the device and sent to the server. The server analyzes the received video data and extracts important actions and objects. It then generates a procedure manual in natural language based on the extracted data and outputs it in a specified format. The system also incorporates an emotion engine that recognizes the user's emotions and reflects the results in the generation of the procedure manual.
[0248] Program Overview
[0249] 1. Recording Stage
[0250] User: Record work content with a camera. Specifically, the user acquires a smartphone with a camera or a dedicated recording device and records the process of work from start to finish.
[0251] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, the video data is automatically saved to the internal storage.
[0252] 2. Uploading videos
[0253] Device: Sends saved video data to the server via the Internet. The device sends saved video data to the specified server upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to inform the user of progress.
[0254] 3. Video Reception and Preprocessing
[0255] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. The server checks the consistency of the video data and performs a simple check to see if there are any missing or corrupted data.
[0256] 4. Video Analysis
[0257] Server: Sends video data to the AI analysis engine. The server divides the stored video data into frames and performs pre-processing to extract the important parts of the frames.
[0258] Server: The AI analytics engine analyzes the video data. It performs frame-by-frame object recognition, motion detection, emotion analysis, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly fashion.
[0259] 5. Emotion analysis
[0260] Server: Analyzes the user's facial expressions and voice in the video data and uses an emotion engine to recognize emotions. Evaluates the user's stress level and satisfaction level and reflects that data in the analysis results.
[0261] Server: Adjust the content and tone of instructions based on the perceived emotion, for example adding more detailed explanations or additional information where the user is feeling stressed.
[0262] 6. Text Generation
[0263] Server: Generates instructions in natural language based on the extracted important actions and user emotion data. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[0264] Server: Formats the generated text data into a manual. Formats the text according to the manual template and divides it into sections.
[0265] 7. Manual Output
[0266] Server: Generates the finished manual in PDF or other specified formats. Converts the entire manual to PDF and adds tables, diagrams, screenshots, etc. (if needed).
[0267] Server: Sends the generated manual to the user's device. The completed manual is sent to the user as an email attachment or a dedicated download link is provided.
[0268] 8. Feedback and Learning
[0269] Server: Records the user's emotion recognition results and reflects them in the next procedure manual generation. Based on the recorded emotion data, the recommended procedure manual content is personalized.
[0270] Users: Provide feedback after using the manual, so the system continually learns and improves the quality of the procedure manual.
[0271] Specific examples
[0272] As a specific example, we will explain a situation in which an operating manual for a new machine is being created in a factory.
[0273] 1. Recording Stage
[0274] User: A factory worker uses his smartphone to film how to operate a new machine.
[0275] Device: Your smartphone will save this footage to its internal storage.
[0276] 2. Uploading videos
[0277] Device: The smartphone uses Wi-Fi to upload the video data to a designated server.
[0278] 3. Video Reception and Preprocessing
[0279] Server: The server receives the video data and performs preprocessing for AI analysis.
[0280] 4. Video Analysis
[0281] Server: The server sends the video data to an AI analysis engine, which recognizes and lists the actions in each frame.
[0282] 5. Emotion analysis
[0283] Server: Recognizes emotions from the user's facial expressions and voice and adjusts the contents of the instruction manual.
[0284] 6. Text Generation
[0285] Server: Generates instructions in natural language based on the extracted action and emotion data and formats them into a manual.
[0286] 7. Manual Output
[0287] Server: Generates the completed manual in PDF format and sends it to the worker's email address.
[0288] 8. Feedback and Learning
[0289] Server: The system continues to learn based on user feedback and improves the quality of the instructions.
[0290] This will significantly reduce time and effort, and enable the rapid provision of up-to-date manuals that take user feelings into consideration.
[0291] The processing flow will be explained below.
[0292] Step 1:
[0293] Users record their work with a camera. Specifically, they record the entire process from start to finish using a camera-equipped smartphone or a dedicated recording device.
[0294] Step 2:
[0295] The device will temporarily store the recorded video data. After the recording device or smartphone has finished recording, the video data will automatically be saved to its internal storage.
[0296] Step 3:
[0297] The device sends the stored video data to the server via the Internet. The device sends the stored video data to the specified server's upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[0298] Step 4:
[0299] The server saves the received video data in the server's storage. The server organizes and saves the received data in a dedicated directory. It checks the consistency of the video data and performs a simple check to see if there are any missing or damaged data.
[0300] Step 5:
[0301] The server sends the video data to an AI analysis engine, which divides the stored video data into frames and performs preprocessing to extract the important parts of the frames.
[0302] Step 6:
[0303] The server's AI analysis engine analyzes the video data, performing frame-by-frame object recognition, motion detection, emotion analysis, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[0304] Step 7:
[0305] The server uses an emotion engine to analyze the user's facial expressions and voice in the video data and recognize emotions, such as whether the user is confused or happy.
[0306] Step 8:
[0307] The server adjusts the content and tone of the instructions based on the perceived emotion, for example adding detailed explanations or additional information if it detects that the user is stressed.
[0308] Step 9:
[0309] The server generates a procedure manual in natural language based on the extracted important actions and user emotion data. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[0310] Step 10:
[0311] The server formats the generated text data into a manual, formatting the text according to the manual template and dividing it into sections.
[0312] Step 11:
[0313] The server generates the finished manual in PDF or other specified format, converting the entire manual to PDF and adding tables, diagrams, screenshots, etc. (if needed).
[0314] Step 12:
[0315] The server sends the generated manual to the user's device, either by attaching the completed manual to an email or by providing a dedicated download link.
[0316] Step 13:
[0317] The server records the user's emotion recognition results and reflects them in the next procedure manual generation. The content of the procedure manual is personalized using the user's past emotion data.
[0318] Step 14:
[0319] The server continues to learn from feedback after the manual is used, and by combining user feedback with emotional data, the quality of the manual is improved.
[0320] Example 2
[0321] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0322] Conventional procedure manual generation systems often create procedures without considering the user's emotions, resulting in a lack of means to identify and address parts that users find emotionally difficult. Furthermore, feedback is not effectively incorporated into the procedure manual generation process, resulting in inconsistent quality. This can result in procedures that are not suited to actual work.
[0323] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0324] In this invention, the server includes: [means for providing an emotion engine that analyzes emotions from the user's facial expressions and voice; [means for generating a procedure manual in natural language based on the extracted data and emotion data; and [means for collecting feedback from the user and allowing the system to learn.] This makes it possible [to provide information that takes the user's emotions into consideration, thereby improving the quality and usability of the procedure manual].
[0325] "Video data" refers to video files in which users record their work.
[0326] "AI engine" refers to an artificial intelligence algorithm that analyzes video data and extracts important actions and objects.
[0327] "Important actions" refer to the work steps and operations required to generate a procedure manual.
[0328] An "object" refers to an object that is recognized in video data.
[0329] An "emotion engine" refers to a system that recognizes and analyzes emotions from the user's facial expressions and voice in video data.
[0330] A "procedure manual" refers to a document that describes the steps a user takes to perform a specific task.
[0331] "Predetermined format" refers to the predetermined manner in which the generated procedure manual is displayed or printed.
[0332] "Terminal" refers to a device for acquiring and transmitting video data.
[0333] "Feedback" refers to evaluations and opinions about the system provided by users.
[0334] "Preprocessing" refers to the process of organizing and checking consistency of video data before analyzing it.
[0335] "PDF format" refers to the Portable Document Format, an electronic document format.
[0336] "Download link" refers to the URL for obtaining a file via the Internet.
[0337] MODE FOR CARRYING OUT THE INVENTION
[0338] The following hardware and software are primarily used in the embodiment of this invention. A user records work content using a smartphone or dedicated recording device, and the terminal temporarily stores this video data. The terminal then transmits the video data to a server, which analyzes the data using an AI engine and an emotion engine. A procedure manual is then generated in natural language based on the extracted information and output in a specified format. The procedure manual is output in PDF format and provided to the user. The system also collects user feedback and continues to learn.
[0339] Hardware and software used
[0340] Hardware
[0341] A smartphone or dedicated recording device
[0342] server
[0343] software
[0344] AI analysis engine (analysis of video data)
[0345] Emotion engine (user emotion analysis)
[0346] PDF generation software (generating instruction manuals)
[0347] Internet connection software (data transmission and reception)
[0348] Data processing and calculation flow
[0349] 1. User: Records work using a smartphone or dedicated recording device. When a user wants to record how to operate a new machine, they can record all the operating procedures at once.
[0350] 2. Device: Temporarily stores the recorded video data. When the recording is finished, the device automatically saves the data to the internal storage and notifies the user of the progress.
[0351] 3. Device: The video data is sent to the server via the Internet. The data is uploaded to the specified server via Wi-Fi or mobile network. The upload progress is displayed on the screen.
[0352] 4. Server: Stores the received video data in storage and checks the data consistency. It performs a simple check to see if there are any missing or corrupted data.
[0353] 5. Server: The video data is divided into frames, and preprocessing such as object recognition and motion detection is performed using an AI analysis engine.
[0354] 6. Server: The AI analysis engine analyzes the video data, identifies and lists important actions and procedures, such as "pressing a button" or "pulling a lever."
[0355] 7. Server: The emotion engine recognizes emotions from the user's facial expressions and voice, evaluates the data, determines the user's stress level and satisfaction, and reflects this in the operating instructions.
[0356] 8. Server: Generates a natural language instruction manual based on the extracted important actions and emotion data. The generated instruction manual includes specific steps such as "1. Turn on the power" and "2. Select the settings menu."
[0357] 9. Server: The generated text data is organized into a specified manual format, categorizing it into sections and adding tables and charts as needed.
[0358] 10. Server: The completed procedure manual is output in PDF format and sent to the user's device. The procedure manual is attached to an email or a download link is provided.
[0359] 11. Server: Collects user feedback and the system continues to learn. Based on the feedback, it will be reflected in the next generation of the procedure manual.
[0360] Specific examples
[0361] When creating an operating manual for a new machine in a factory, a user uses a smartphone to film how to operate the new machine in the factory. The video data is saved by the device and uploaded to a server via the Internet. The server receives the video data, analyzes the data using an AI analysis engine and an emotion engine, and generates a manual in natural language. This manual is created in PDF format and sent to the user. Feedback provided by the user after using the manual is used to help generate manuals in the future.
[0362] Prompt Sentence Examples
[0363] "I recorded the machine's operating procedures on my smartphone. Please generate an operating manual using this video data. Please list important actions and customize the manual with consideration for the user's emotions."
[0364] By inputting this prompt into the generative AI model, the system can efficiently generate instructions.
[0365] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0366] Step 1:
[0367] Recording Stage
[0368] User: Record work with a camera. Users record the process from start to finish using a smartphone or dedicated recording device. For example, record the process of learning how to operate a new machine. During this process, it is recommended to adjust the camera angle and lighting.
[0369] Input: User operation during work
[0370] Output: Recorded video data
[0371] Step 2:
[0372] Data storage
[0373] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, it automatically saves the video data to its internal storage. It notifies the user that the save is complete. The saved data is usually in MP4 or MOV format.
[0374] Input: Recorded video data
[0375] Output: Video data stored in the internal storage
[0376] Step 3:
[0377] Video upload
[0378] Device: Video data is sent to the server via the internet. Data is uploaded to the specified server via Wi-Fi or mobile network. The transfer progress is displayed on the screen and the user is notified when the transfer is complete. Data is encrypted during transfer as a security measure.
[0379] Input: Video data stored in the internal storage
[0380] Output: Video data sent to the server
[0381] Step 4:
[0382] Video reception and preprocessing
[0383] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. After saving, the data is checked for consistency and a simple check is performed to see if there are any missing or corrupted data. This preprocessing also includes checking the frame rate and renaming the data.
[0384] Input: Video data sent to the server
[0385] Output: Pre-processed video data
[0386] Step 5:
[0387] Video Analysis
[0388] Server: Sends video data to the AI analysis engine. The server divides the data into frames and performs preprocessing to extract important parts. It identifies important objects and actions in each frame.
[0389] Input: Preprocessed video data
[0390] Output: Object and movement data for each frame analyzed by the AI analysis engine
[0391] Server: The AI analysis engine analyzes the video data, identifies and lists important actions and procedures, for example, recognizing operations such as "pressing a button" or "pulling a lever."
[0392] Input: Object and movement data per frame
[0393] Output: A list of extracted important actions and steps
[0394] Step 6:
[0395] Emotion analysis
[0396] Server: Analyzes the user's facial expressions and voice in the video data and uses an emotion engine to recognize emotions. The server evaluates the user's stress level and satisfaction level and reflects this data in the analysis results. Analyzing subtle changes in facial expressions and tone of voice in particular enables more accurate emotion recognition.
[0397] Input: User's facial expressions and voice in video data
[0398] Output: User emotion data
[0399] Server: Adjust the content and tone of instructions based on the perceived emotion, for example adding more detailed explanations or additional information where the user is feeling stressed.
[0400] Input: User sentiment data, list of important actions and steps
[0401] Output: An outline of the procedure manual with the sentiment data reflected
[0402] Step 7:
[0403] Text Generation
[0404] Server: Generates instructions in natural language based on the extracted important actions and emotion data. AI analyzes the action list and generates specific steps such as "1. Turn on the power" and "2. Select the settings menu." Appropriate explanations and cautions are added to each step.
[0405] Input: Instructions outline with sentiment data reflected
[0406] Output: Natural language generated instructions
[0407] Server: Prepares the generated text data into a predetermined manual format. Formats the text according to a template, categorizes it into sections, and adds tables and charts as needed.
[0408] Input: Natural language generated instructions
[0409] Output: Formatted instructions
[0410] Step 8:
[0411] Manual Output
[0412] Server: Prints the completed manual in PDF format and sends it to the user's device. The instructions are attached to an email or a download link is provided. Adds tables and diagrams as needed (if necessary).
[0413] Input: Formatted instructions
[0414] Output: Instructions in PDF format provided to users
[0415] Step 9:
[0416] Feedback and Learning
[0417] Server: Collects feedback from users and the system continues to learn. Based on the feedback, it reflects it in the next generation of the procedure manual. Specifically, it analyzes evaluation points and comments and collects data to improve the procedure manual.
[0418] Input: User feedback
[0419] Output: Learning data that will be useful for the next procedure generation
[0420] (Application example 2)
[0421] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0422] Generating work instructions is important for improving the work efficiency of factory robots. However, conventional methods generate uniform instructions, making it difficult to provide personalized instructions that take into account the emotions and situations of individual users. Furthermore, analyzing work content without considering the user's emotions can lead to the risk of overlooking areas that cause stress to the user or areas for improvement.
[0423] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data, means for extracting important actions and objects using an AI engine that analyzes the acquired video data, means for using an emotion analysis engine that recognizes the user's emotions and reflects the results in generating a procedure manual, means for generating a procedure manual in natural language based on the extracted data and emotion data, and means for outputting the generated procedure manual in a predetermined format and providing it to the user. This makes it possible to provide a personalized procedure manual that takes into account the user's work situation and emotions.
[0424] "Video data" refers to a series of image frames captured by a camera, and is the digital data that is the subject of analysis.
[0425] "AI Engine" is a software platform for analyzing and extracting data using artificial intelligence technology.
[0426] "Critical actions" refer to actions or behaviors that are particularly essential in work procedures or processes.
[0427] "Object" refers to an object, such as a person or object, that is recognized in video data.
[0428] An "emotion analysis engine" is a software platform that recognizes emotions from a user's facial expressions and voice in video data and extracts that emotional state as data.
[0429] A "procedure manual" refers to a document that provides detailed instructions for a specific task or operation.
[0430] A "natural language" is a language that humans use on a daily basis, and is a written expression used to describe specific processes or instructions.
[0431] "Format" refers to a template or form that defines the structure and format of a document or data.
[0432] A "server" is a computer system that stores data, processes data, and provides services over a network.
[0433] "Analysis" is the process of breaking down data into detail to reveal its meaning and structure.
[0434] This invention is a system for automatically generating personalized operation manuals that incorporate the emotions of users by utilizing factory robots. The system of the present invention captures video of the user working, analyzes it using an AI engine, and generates operation manuals based on the extracted data and emotion analysis. Specifically, the system performs the following processes.
[0435] System Overview
[0436] 1. Recording Stage
[0437] User: Users working in the factory record their work using cameras attached to the robots, and through this process, video data is acquired in real time.
[0438] Hardware: Video data is temporarily stored in the robot's internal storage.
[0439] 2. Uploading videos
[0440] Server: The robot sends video data to the server via an internet connection, either wireless LAN or wired data communication.
[0441] 3. Video Reception and Preprocessing
[0442] Server: The received video data is stored in the server's storage. Next, as preprocessing, the data consistency is checked and processing such as noise removal and frame division is performed.
[0443] 4. Video Analysis
[0444] Server: An AI engine using OpenCV and TensorFlow analyzes the video data and extracts important actions and objects, using algorithms for object recognition and motion detection.
[0445] 5. Emotion analysis
[0446] Server: Using an emotion analysis engine such as Microsoft® Azure® Cognitive Services, the server recognizes emotions from the user's facial expressions and voice, enabling it to evaluate whether the user is feeling stressed.
[0447] 6. Text Generation
[0448] Server: Using a generative AI model such as GPT-4 (registered trademark), the server generates a procedure manual in natural language based on the extracted key actions and emotion data. The procedure manual is formatted in a predetermined format.
[0449] 7. Manual Output
[0450] Server: Generates the completed procedure manual in PDF format and provides it to the user by displaying it on the robot's display or by sending it to the user's email address.
[0451] 8. Feedback and Learning
[0452] Server: Collects feedback from users after using the procedure and uses that data to improve the analysis model for the next run, allowing the system to continuously learn and improve the quality of the procedure.
[0453] Specific examples
[0454] For example, when inspecting products in a factory, a robot can record the work, analyze the video, and compile a procedure manual for the next inspection method, taking the user's feelings into consideration. This can increase work efficiency while reducing user stress.
[0455] Prompt Sentence Examples
[0456] "Please record the product inspection process in the factory with a camera and generate the next inspection procedure based on the analysis results."
[0457] This system can provide flexible instructions that reflect the user's emotions and specific work content, which is expected to improve work efficiency and user satisfaction.
[0458] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0459] Step 1:
[0460] User: A user working in a factory records their work using a camera attached to the robot.
[0461] Input: The video data you are working with.
[0462] Output: Captured working video data.
[0463] Specific operation: The camera captures video data in real time and stores it in the robot's built-in storage.
[0464] Step 2:
[0465] Server: The robot sends the video data to the server via an internet connection.
[0466] Input: Video data of work inside the robot.
[0467] Output: Video data stored on the server.
[0468] Specific operation: The robot sends video data to the server's upload endpoint using Wi-Fi or a wired connection.
[0469] Step 3:
[0470] Server: Preprocesses the received video data and passes it to the AI analysis engine.
[0471] Input: Received video data.
[0472] Output: Pre-processed video data.
[0473] Specific operation: The server performs preprocessing such as checking data consistency, removing noise, and splitting frames.
[0474] Step 4:
[0475] Server: An AI engine using OpenCV and TensorFlow analyzes video data and extracts important actions and objects.
[0476] Input: Preprocessed video data.
[0477] Output: Extracted important action and object data.
[0478] Specific actions: The AI engine runs object recognition and action detection algorithms to list important actions and objects.
[0479] Step 5:
[0480] Server: Uses an emotion analysis engine to recognize emotions from the user's facial expressions and voice.
[0481] Input: Preprocessed video and audio data.
[0482] Output: User emotion data.
[0483] How it works: An emotion analysis engine such as Microsoft Azure Cognitive Services analyzes facial expressions and vocal tones to digitize the user's emotional state.
[0484] Step 6:
[0485] Server: Using a generative AI model, it generates instructions in natural language based on the extracted key actions and emotion data.
[0486] Input: Important action and emotion data.
[0487] Output: Instructions generated in natural language.
[0488] How it works: A generative AI model such as GPT-4 analyzes key action lists and emotional data to generate instructions in natural language.
[0489] Step 7:
[0490] Server: Outputs the generated procedure manual in a specified format and provides it to the user.
[0491] Input: Instructions generated in natural language.
[0492] Output: Instructions in PDF format.
[0493] Specific operation: The server formats the instruction manual in PDF format and displays it on the robot's display or sends it to the user by email.
[0494] Step 8:
[0495] Server: Collects feedback from users after using the procedure and uses that data to improve the next analysis model.
[0496] Input: User feedback data.
[0497] Output: An improved analytical model.
[0498] How it works: The server stores the feedback data in a database and retrains the generative AI model (GPT-4) to improve the quality of future instructions.
[0499] Through the above processing steps, the present invention automatically generates a personalized procedure manual that takes the user's feelings into consideration, thereby improving work efficiency and user satisfaction.
[0500] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0501] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0502] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0503] [Second embodiment]
[0504] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0505] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0506] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0507] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0508] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0509] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0510] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0511] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0512] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0513] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0514] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0515] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0516] As an embodiment of the present invention, a system including the following processes is provided.
[0517] System Overview
[0518] The user records the work content with a camera, and the video data is received by the terminal and sent to the server. The server analyzes the received video data and extracts important actions and objects. Then, based on the extracted data, a procedure manual is generated in natural language and output in a specified format. A specific embodiment is shown below.
[0519] Program Overview
[0520] 1. Recording Stage
[0521] User: Record the work process with a camera. Specifically, the user records the entire process from start to finish using a camera-equipped smartphone or dedicated recording device.
[0522] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, the video data is automatically saved to the internal storage.
[0523] 2. Uploading videos
[0524] Device: Sends saved video data to the server via the Internet. The device sends saved video data to the specified server upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to inform the user of progress.
[0525] 3. Video Reception and Preprocessing
[0526] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. The server checks the consistency of the video data and performs a simple check to see if there are any missing or corrupted data.
[0527] 4. Video Analysis
[0528] Server: Sends video data to the AI analysis engine. The server divides the stored video data into frames and performs pre-processing to extract the important parts of the frames.
[0529] Server: The AI analysis engine analyzes the video data. It performs frame-by-frame object recognition, motion detection, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[0530] 5. Text Generation
[0531] Server: Generates a procedure manual in natural language based on the extracted important actions and objects. The server analyzes the action list extracted by AI and generates procedures such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[0532] Server: Formats the generated text data into a manual. Formats the text according to the manual template and divides it into sections.
[0533] 6. Manual Output
[0534] Server: Generates the finished manual in PDF or other specified formats. Converts the entire manual to PDF and adds tables, diagrams, screenshots, etc. (if needed).
[0535] Server: Sends the generated manual to the user's device. The completed manual is sent to the user as an email attachment or a dedicated download link is provided.
[0536] Specific examples
[0537] As a specific example, a situation in which an operating manual for a machine in a factory is to be created will be described.
[0538] 1. Recording Stage
[0539] User: A factory worker uses his smartphone to film how to operate a new machine.
[0540] Device: Your smartphone will save this footage to its internal storage.
[0541] 2. Uploading videos
[0542] Device: The smartphone uses Wi-Fi to upload video data to a designated server.
[0543] 3. Video Reception and Preprocessing
[0544] Server: The server receives the video data and performs preprocessing for AI analysis.
[0545] 4. Video Analysis
[0546] Server: The server sends the video data to an AI analysis engine, which recognizes and lists the actions in each frame.
[0547] 5. Text Generation
[0548] Server: Based on the extracted actions, generate a procedure manual in natural language and format it into a manual.
[0549] 6. Manual Output
[0550] Server: Generates the completed manual in PDF format and sends it to the worker via email. This allows users to easily create work manuals and quickly update them.
[0551] This significantly reduces time and effort, and enables the latest manuals to be provided promptly.
[0552] The processing flow will be explained below.
[0553] Step 1:
[0554] Users record their work with a camera. Specifically, they acquire a smartphone with a camera or a dedicated recording device and record the entire process from start to finish.
[0555] Step 2:
[0556] The device will temporarily store the recorded video data. After the recording device or smartphone has finished recording, the video data will automatically be saved to its internal storage.
[0557] Step 3:
[0558] The device sends the stored video data to the server via the Internet. The device sends the stored video data to the specified server's upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[0559] Step 4:
[0560] The server saves the received video data in the server's storage. The server organizes and saves the received data in a dedicated directory. It checks the consistency of the video data and performs a simple check to see if there are any missing or damaged data.
[0561] Step 5:
[0562] The server sends the video data to an AI analysis engine, which divides the stored video data into frames and performs preprocessing to extract the important parts of the frames.
[0563] Step 6:
[0564] The server's AI analysis engine analyzes the video data, performing frame-by-frame object recognition, motion detection, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[0565] Step 7:
[0566] The server generates a procedure manual in natural language based on the extracted important actions and objects. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[0567] Step 8:
[0568] The server formats the generated text data into a manual, formatting the text according to the manual template and dividing it into sections.
[0569] Step 9:
[0570] The server generates the finished manual in PDF or other specified format, converting the entire manual to PDF and adding tables, diagrams, screenshots, etc. (if needed).
[0571] Step 10:
[0572] The server sends the generated manual to the user's device, either by attaching the completed manual to an email or by providing a dedicated download link.
[0573] Example 1
[0574] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0575] There has been no system that can easily record video data of work content and automatically generate procedure manuals based on that data. In particular, there is a need for an efficient and accurate process that can analyze the video data, extract important actions and objects, and generate procedure manuals. This has led to the need for a system that can reduce work hours and labor, standardize work procedures, and enable rapid update support.
[0576] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0577] In this invention, the server includes: [means for acquiring video data;] [means for extracting important actions and objects using an AI analysis unit that analyzes the acquired video data; and] [means for generating a procedure manual in natural language based on the extracted data. This makes it possible to automatically generate a procedure manual based on the video data of the work content and output it in a specified format.
[0578] "Video data" refers to digital information in the form of moving images captured using a camera or other imaging device.
[0579] A "server" is a computing device used to receive, analyze, store, and transmit data over a network.
[0580] "AI Analysis Unit" means a collection of software and hardware that implements artificial intelligence technology used to analyze video data and identify and extract significant actions and objects.
[0581] "Critical actions" refer to actions or operations that require special attention when carrying out business procedures.
[0582] "Objects" refer to specific items or devices that exist within the video data and are required to explain business procedures.
[0583] A "procedure manual" is a document that explains business procedures in a specific and orderly manner as a series of steps.
[0584] "Means for acquiring" refers to methods and devices for collecting video data using cameras and recording devices.
[0585] "Means for generating in natural language" refers to the method and process for generating a procedure manual in a natural language that is easy for humans to understand, based on the extracted data.
[0586] A "prescribed format" is a rule that ensures that procedures are organized in a certain format and displayed in a way that is visually easy to understand.
[0587] As an embodiment of this invention, we provide a system that automatically captures and analyzes work procedures at companies or work sites and generates procedure manuals. Specifically, this system allows users to record work content with a camera, and the video data is received by a terminal and sent to a server. The details of the system are described below.
[0588] Recording Stage
[0589] User:
[0590] Work is recorded with a camera. Specifically, the user uses a camera-equipped smartphone or a dedicated recording device to record work from start to finish. For example, a factory worker uses a smartphone to record how to operate a new machine.
[0591] Device:
[0592] Recorded video data is saved to the internal storage. After recording is complete, the smartphone automatically saves the video data temporarily.
[0593] Video upload
[0594] Device:
[0595] The saved video data is sent to a server via the Internet. The device uploads the video data to the specified server via Wi-Fi or a mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[0596] Video reception and preprocessing
[0597] server:
[0598] The server stores the received video data in a dedicated directory, checks for missing or corrupted data, and checks the data for consistency. Organizes the video data into batches for analysis.
[0599] Video Analysis
[0600] server:
[0601] The server divides the stored video data into frames and sends them to the AI analysis unit. As preprocessing, the video resolution is adjusted appropriately and unnecessary frames are removed. The AI analysis unit (e.g., TensorFlow or OpenCV) is used to perform object recognition and motion detection in each frame, and convert text and voice as needed. For example, scenes in which the user presses a button or pulls a lever are extracted and listed.
[0602] Text Generation
[0603] server:
[0604] Generates instructions in natural language based on the extracted important actions and objects. Using an AI model (e.g., GPT-3), it creates steps such as "1. Turn on the power" and "2. Select the settings menu" from the listed actions. Provides appropriate explanations and cautions for each step. For example, it adds instructions such as "Check the power cord and be careful not to press the buttons too hard."
[0605] server:
[0606] The generated text data is formatted into a manual, with each step divided into sections and visually organized using bulleted and numbered lists.
[0607] Manual Output
[0608] server:
[0609] Generate the finished manual in PDF or other format of your choice. Use tools such as Adobe Acrobat or LaTeX to add tables, charts, and screenshots to create a visually appealing manual.
[0610] server:
[0611] The generated manual is sent to the user's device. The completed manual is sent as an email attachment to the user, or a dedicated download link is provided.
[0612] Specific examples
[0613] As a specific example, we will explain a situation in which an operating manual for a new machine is being created in a factory.
[0614] User:
[0615] Factory workers use their smartphones to record themselves operating a new machine.
[0616] Device:
[0617] The smartphone saves this footage to its internal storage.
[0618] Device:
[0619] The smartphone uses Wi-Fi to upload the video data to a designated server.
[0620] server:
[0621] The server receives the video data and performs preprocessing for AI analysis.
[0622] server:
[0623] The server sends the video data to an AI analysis unit, which recognizes and lists the movements in each frame.
[0624] server:
[0625] Based on the extracted actions, a procedure manual is generated in natural language and formatted into a manual.
[0626] server:
[0627] The completed manual is generated in PDF format and sent to the worker via email, allowing users to easily create work manuals and quickly update them.
[0628] Examples of prompt statements
[0629] "Please film how to operate the new machine and create a manual."
[0630] "Please tell me the procedure for uploading video data to the server."
[0631] "Please explain in detail how to generate the procedure manual."
[0632] This will provide a system that standardizes business procedures, improves efficiency, and enables rapid update support.
[0633] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0634] Step 1:
[0635] User: Records work with a camera. The user uses a smartphone with a camera or other recording device to record the entire process from start to finish. For example, a factory worker uses a smartphone to record how to operate a new machine. The user takes care to clearly capture specific operating steps (e.g., how to press a button, pull a lever). Input: The device used is a smartphone with a camera. Output: Recorded video data.
[0636] Step 2:
[0637] Device: After recording is complete, the device saves the video data to its internal storage. At the same time, the device sends the video data to the server via Wi-Fi or the mobile network. During uploading, the device displays a progress bar to inform the user of the transfer status. Input: Recorded video data. Output: Video data sent to the server.
[0638] Step 3:
[0639] Server: The server stores the received video data in a dedicated directory in an organized manner. It first checks the data for consistency and performs a simple check to see if there are any missing or corrupted data. If the data is OK, it is broken down into frames for analysis and the resolution is adjusted if necessary. Input: Video data sent to the server. Output: Video data that has been split into frames and pre-processed.
[0640] Step 4:
[0641] Server: Sends preprocessed video data to an AI analysis unit (e.g., TensorFlow or OpenCV), which performs object recognition and action detection for each frame, and converts text and voice as needed. For example, it extracts and lists scenes in which a button is pressed or a lever is pulled. Input: Preprocessed video data for each frame. Output: A list of extracted important actions and objects.
[0642] Step 5:
[0643] Server: Based on the extracted important actions and objects, a generative AI model (e.g., GPT-3) is used to generate instructions in natural language. Detailed explanations and precautions are added for each action (e.g., pressing a button), and steps such as "1. Turn on the power" and "2. Select the settings menu" are formed. Input: A list of extracted actions and objects. Output: Instructions generated in natural language.
[0644] Step 6:
[0645] Server: Formats the generated procedure manual into a manual format and outputs it in PDF or other specified format. Tables, diagrams, and in some cases screenshots are added to the procedure manual to make it visually easy to understand. Input: Procedure manual generated in natural language. Output: Procedure manual formatted in PDF or other specified format.
[0646] Step 7:
[0647] Server: Provides the completed procedure to the user. The server sends the procedure as an email attachment or provides a dedicated download link. It may also store it in a database. Input: Procedure formatted in PDF or other specified format. Output: Procedure provided to the user.
[0648] Through this series of processes, a system is realized that can automatically generate procedure manuals from video data of work content and provide them to users quickly and efficiently, making it possible to standardize work procedures and respond to updates quickly.
[0649] (Application example 1)
[0650] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0651] In modern factories and manufacturing sites, specialized robot maintenance work is complex, and creating procedure manuals requires a great deal of time and effort. Therefore, there is a need for an effective automatic procedure manual generation system to improve the efficiency and standardization of maintenance work. However, existing manual procedure manual creation methods have issues such as difficulty in recording and analyzing work, and difficulty in quickly reflecting the latest information.
[0652] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0653] In this invention, the server includes: [means for acquiring video data using a device with a camera; [means for extracting important actions and objects using an AI analysis engine that analyzes the acquired video data; and [means for generating a procedure manual in natural language based on the extracted data.] This makes it possible [to achieve work efficiency and standardization, and to quickly reflect the latest information in the procedure manual].
[0654] "Video data" is a series of image information captured by a device with a camera, and is digital data that includes details of actions and objects.
[0655] "AI analysis engine" is a general term for software and hardware that uses artificial intelligence technology to analyze video data and extract important actions and objects.
[0656] A "procedure" is a document that provides step-by-step instructions for performing a specific task or operation.
[0657] "Natural language" is a language used by humans on a daily basis and a set of textual forms generated by computer programs.
[0658] A "generative AI model" is an artificial intelligence model that learns large amounts of data in advance to generate instruction manuals and other text data.
[0659] A "prompt" is a series of phrases or documents that are input to a generative AI model to determine the content and format of the generated text.
[0660] A "server" is a computer that receives requests from client devices over a network and provides data processing and data storage functions.
[0661] "Preprocessing" refers to the process of removing noise from video data and converting its format before analysis by the AI analysis engine.
[0662] A "data processing device" is a computer system for analyzing video data and extracting important information.
[0663] A "template" is a predefined format or style guide for formatting procedures and reports.
[0664] "Digital document format" means a document format that can be stored and displayed electronically, including PDF.
[0665] MODE FOR CARRYING OUT THE INVENTION
[0666] System Overview
[0667] As an embodiment of the present invention, we provide a system that performs the following process. In this system, a user records work content using a camera-equipped smartphone and sends the video data to a server. The server analyzes the received video data and extracts important actions and objects. Then, based on the extracted data, it generates a procedure manual in natural language and outputs it in a specified format.
[0668] Explanation of program processing
[0669] Hardware and software used
[0670] Smartphone: A smart device used as a recording device.
[0671] Server: A computer used for data analysis that receives requests from client devices over the network and processes the data.
[0672] OpenCV: A library for recording and processing video.
[0673] Requests library: Used to send video data to the server.
[0674] AI analysis engine: An engine for analyzing video data and extracting important actions and objects.
[0675] PDF generation library: A library for converting instructions into PDF format (e.g., ReportLab).
[0676] Processing flow
[0677] 1. Recording: The user records the robot's maintenance work using a smartphone. The recorded data is saved in the smartphone's internal storage.
[0678] 2. Video Upload: Once the recording is complete, the device (smartphone) sends the saved video data to the server via Wi-Fi or mobile network. This data transmission is performed using the Requests library.
[0679] 3. Data processing: The server receives the transmitted video data and performs preprocessing, which includes noise reduction and format conversion. The AI analysis engine then divides the video data into frames and extracts important actions and objects.
[0680] 4. Procedure generation: A procedure manual is generated in natural language based on the data extracted by the AI analysis engine. A generative AI model is used here. A prompt sentence for generating the procedure manual is also created. By inputting this prompt sentence into the generative AI model, automatically generated text can be obtained.
[0681] 5. Output of procedure manual: The generated procedure manual is output in PDF format using a PDF generation library, and the output PDF is provided to the user.
[0682] Specific examples
[0683] For example, let's say a maintenance task of adding lubricant to the joints of a robot is recorded in a factory. This recorded data is sent to a server, which then uses an analysis engine to analyze the video and extract important actions (such as "remove the cover," "prepare the lubricant," and "apply lubricant to the specified location"). Based on the extracted actions, the AI model automatically generates a procedure manual and provides it to the engineer in PDF format.
[0684] Prompt Sentence Examples
[0685] Video data URL: {video_url}
[0686] Analysis output:
[0687] Object Recognition
[0688] Action Detection
[0689] Procedure generation
[0690] The above is an embodiment of the present invention. This system makes it possible to improve the efficiency and standardize maintenance work, and to quickly update the latest information in the procedure manual.
[0691] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0692] Program processing flow
[0693] Step 1:
[0694] Recording
[0695] The user records the robot's maintenance work using a smartphone. At this time, a camera-equipped device is used to capture the entire process from start to finish. The recorded data (video data) is saved in the smartphone's internal storage.
[0696] Input: Maintenance work video
[0697] Output: Video data stored in the internal storage
[0698] Step 2:
[0699] Video upload
[0700] Once recording is complete, the device (smartphone) sends the saved video data to the server via the Internet. The device uses the Requests library to send the video data to the specified server's upload endpoint. At this time, the file transfer status is displayed to inform the user of the progress.
[0701] Input: Video data stored in the internal storage
[0702] Output: Video data sent to the server
[0703] Step 3:
[0704] Data reception and preprocessing
[0705] The server receives the transmitted video data and stores it in storage. Upon receiving the data, it performs pre-processing such as noise removal and format conversion. This process converts the data into a format that can be smoothly processed by the analysis engine.
[0706] Input: Video data sent to the server
[0707] Output: Pre-processed video data
[0708] Step 4:
[0709] Video Analysis
[0710] The preprocessed video data is passed to the server's AI analysis engine, which divides the video data into frames and extracts important actions and objects from each frame. Here, it performs tasks such as object recognition and action detection for each frame, and creates a list of important actions.
[0711] Input: Preprocessed video data
[0712] Output: Data listing important actions and objects
[0713] Step 5:
[0714] Procedure generation
[0715] Based on the data listing important actions and objects, the server uses a generative AI model to generate instructions in natural language. A prompt is entered into the generative AI model, which then automatically generates text. An example of a prompt might be something like "Video data URL: {video_url}\nAnalysis output:\n- Object recognition\n- Action detection\n- Instruction manual generation."
[0716] Input: Data listing important actions and objects, prompts
[0717] Output: Natural language instructions
[0718] Step 6:
[0719] Output of procedure manual
[0720] The server converts the generated instructions into PDF format using a PDF generation library, and the generated PDF is sent to the user's device or a download link is provided.
[0721] Input: Natural language instructions
[0722] Output: Instructions in PDF format
[0723] By following the above steps, a procedure manual can be automatically generated from a series of maintenance tasks and provided efficiently.
[0724] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0725] As an embodiment of the present invention, we provide a system that includes the following processing. In addition to acquiring and analyzing video data, this system recognizes the user's emotions and reflects them in the generation of procedure manuals, thereby providing more personalized manuals.
[0726] System Overview
[0727] The user records their work with a camera, and the video data is received by the device and sent to the server. The server analyzes the received video data and extracts important actions and objects. It then generates a procedure manual in natural language based on the extracted data and outputs it in a specified format. The system also incorporates an emotion engine that recognizes the user's emotions and reflects the results in the generation of the procedure manual.
[0728] Program Overview
[0729] 1. Recording Stage
[0730] User: Record work content with a camera. Specifically, the user acquires a smartphone with a camera or a dedicated recording device and records the process of work from start to finish.
[0731] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, the video data is automatically saved to the internal storage.
[0732] 2. Uploading videos
[0733] Device: Sends saved video data to the server via the Internet. The device sends saved video data to the specified server upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to inform the user of progress.
[0734] 3. Video Reception and Preprocessing
[0735] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. The server checks the consistency of the video data and performs a simple check to see if there are any missing or corrupted data.
[0736] 4. Video Analysis
[0737] Server: Sends video data to the AI analysis engine. The server divides the stored video data into frames and performs pre-processing to extract the important parts of the frames.
[0738] Server: The AI analytics engine analyzes the video data. It performs frame-by-frame object recognition, motion detection, emotion analysis, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly fashion.
[0739] 5. Emotion analysis
[0740] Server: Analyzes the user's facial expressions and voice in the video data and uses an emotion engine to recognize emotions. Evaluates the user's stress level and satisfaction level and reflects that data in the analysis results.
[0741] Server: Adjust the content and tone of instructions based on the perceived emotion, for example adding more detailed explanations or additional information where the user is feeling stressed.
[0742] 6. Text Generation
[0743] Server: Generates instructions in natural language based on the extracted important actions and user emotion data. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[0744] Server: Formats the generated text data into a manual. Formats the text according to the manual template and divides it into sections.
[0745] 7. Manual Output
[0746] Server: Generates the finished manual in PDF or other specified formats. Converts the entire manual to PDF and adds tables, diagrams, screenshots, etc. (if needed).
[0747] Server: Sends the generated manual to the user's device. The completed manual is sent to the user as an email attachment or a dedicated download link is provided.
[0748] 8. Feedback and Learning
[0749] Server: Records the user's emotion recognition results and reflects them in the next procedure manual generation. Based on the recorded emotion data, the recommended procedure manual content is personalized.
[0750] Users: Provide feedback after using the manual, so the system continually learns and improves the quality of the procedure manual.
[0751] Specific examples
[0752] As a specific example, we will explain a situation in which an operating manual for a new machine is being created in a factory.
[0753] 1. Recording Stage
[0754] User: A factory worker uses his smartphone to film how to operate a new machine.
[0755] Device: Your smartphone will save this footage to its internal storage.
[0756] 2. Uploading videos
[0757] Device: The smartphone uses Wi-Fi to upload video data to a designated server.
[0758] 3. Video Reception and Preprocessing
[0759] Server: The server receives the video data and performs preprocessing for AI analysis.
[0760] 4. Video Analysis
[0761] Server: The server sends the video data to an AI analysis engine, which recognizes and lists the actions in each frame.
[0762] 5. Emotion analysis
[0763] Server: Recognizes emotions from the user's facial expressions and voice and adjusts the contents of the instruction manual.
[0764] 6. Text Generation
[0765] Server: Generates instructions in natural language based on the extracted action and emotion data and formats them into a manual.
[0766] 7. Manual Output
[0767] Server: Generates the completed manual in PDF format and sends it to the worker's email address.
[0768] 8. Feedback and Learning
[0769] Server: The system continues to learn based on user feedback and improves the quality of the instructions.
[0770] This will significantly reduce time and effort, and enable the rapid provision of up-to-date manuals that take user feelings into consideration.
[0771] The processing flow will be explained below.
[0772] Step 1:
[0773] Users record their work with a camera. Specifically, they record the entire process from start to finish using a camera-equipped smartphone or a dedicated recording device.
[0774] Step 2:
[0775] The device will temporarily store the recorded video data. After the recording device or smartphone has finished recording, the video data will automatically be saved to its internal storage.
[0776] Step 3:
[0777] The device sends the stored video data to the server via the Internet. The device sends the stored video data to the specified server's upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[0778] Step 4:
[0779] The server saves the received video data in the server's storage. The server organizes and saves the received data in a dedicated directory. It checks the consistency of the video data and performs a simple check to see if there are any missing or damaged data.
[0780] Step 5:
[0781] The server sends the video data to an AI analysis engine, which divides the stored video data into frames and performs preprocessing to extract the important parts of the frames.
[0782] Step 6:
[0783] The server's AI analysis engine analyzes the video data, performing frame-by-frame object recognition, motion detection, emotion analysis, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[0784] Step 7:
[0785] The server uses an emotion engine to analyze the user's facial expressions and voice in the video data and recognize emotions, such as whether the user is confused or happy.
[0786] Step 8:
[0787] The server adjusts the content and tone of the instructions based on the perceived emotion, for example adding detailed explanations or additional information if it detects that the user is stressed.
[0788] Step 9:
[0789] The server generates a procedure manual in natural language based on the extracted important actions and user emotion data. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[0790] Step 10:
[0791] The server formats the generated text data into a manual, formatting the text according to the manual template and dividing it into sections.
[0792] Step 11:
[0793] The server generates the finished manual in PDF or other specified format, converting the entire manual to PDF and adding tables, diagrams, screenshots, etc. (if needed).
[0794] Step 12:
[0795] The server sends the generated manual to the user's device, either by attaching the completed manual to an email or by providing a dedicated download link.
[0796] Step 13:
[0797] The server records the user's emotion recognition results and reflects them in the next procedure manual generation. The content of the procedure manual is personalized using the user's past emotion data.
[0798] Step 14:
[0799] The server continues to learn from feedback after the manual is used, and by combining user feedback with emotional data, the quality of the manual is improved.
[0800] Example 2
[0801] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0802] Conventional procedure manual generation systems often create procedures without considering the user's emotions, resulting in a lack of means to identify and address parts that users find emotionally difficult. Furthermore, feedback is not effectively incorporated into the procedure manual generation process, resulting in inconsistent quality. This can result in procedures that are not suited to actual work.
[0803] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0804] In this invention, the server includes: [means for providing an emotion engine that analyzes emotions from the user's facial expressions and voice; [means for generating a procedure manual in natural language based on the extracted data and emotion data; and [means for collecting feedback from the user and allowing the system to learn.] This makes it possible [to provide information that takes the user's emotions into consideration, thereby improving the quality and usability of the procedure manual].
[0805] "Video data" refers to video files in which users record their work.
[0806] "AI engine" refers to an artificial intelligence algorithm that analyzes video data and extracts important actions and objects.
[0807] "Important actions" refer to the work steps and operations required to generate a procedure manual.
[0808] An "object" refers to an object that is recognized in video data.
[0809] An "emotion engine" refers to a system that recognizes and analyzes emotions from the user's facial expressions and voice in video data.
[0810] A "procedure manual" refers to a document that describes the steps a user takes to perform a specific task.
[0811] "Predetermined format" refers to the predetermined manner in which the generated procedure manual is displayed or printed.
[0812] "Terminal" refers to a device for acquiring and transmitting video data.
[0813] "Feedback" refers to evaluations and opinions about the system provided by users.
[0814] "Preprocessing" refers to the process of organizing and checking consistency of video data before analyzing it.
[0815] "PDF format" refers to the Portable Document Format, an electronic document format.
[0816] "Download link" refers to the URL for obtaining a file via the Internet.
[0817] MODE FOR CARRYING OUT THE INVENTION
[0818] The following hardware and software are primarily used in the embodiment of this invention. A user records work content using a smartphone or dedicated recording device, and the terminal temporarily stores this video data. The terminal then transmits the video data to a server, which analyzes the data using an AI engine and an emotion engine. A procedure manual is then generated in natural language based on the extracted information and output in a specified format. The procedure manual is output in PDF format and provided to the user. The system also collects user feedback and continues to learn.
[0819] Hardware and software used
[0820] Hardware
[0821] A smartphone or dedicated recording device
[0822] server
[0823] software
[0824] AI analysis engine (analysis of video data)
[0825] Emotion engine (user emotion analysis)
[0826] PDF generation software (generating instruction manuals)
[0827] Internet connection software (data transmission and reception)
[0828] Data processing and calculation flow
[0829] 1. User: Records work using a smartphone or dedicated recording device. When a user wants to record how to operate a new machine, they can record all the operating procedures at once.
[0830] 2. Device: Temporarily stores the recorded video data. When the recording is finished, the device automatically saves the data to the internal storage and notifies the user of the progress.
[0831] 3. Device: The video data is sent to the server via the Internet. The data is uploaded to the specified server via Wi-Fi or mobile network. The upload progress is displayed on the screen.
[0832] 4. Server: Stores the received video data in storage and checks the data consistency. It performs a simple check to see if there are any missing or corrupted data.
[0833] 5. Server: The video data is divided into frames, and preprocessing such as object recognition and motion detection is performed using an AI analysis engine.
[0834] 6. Server: The AI analysis engine analyzes the video data, identifies and lists important actions and procedures, such as "pressing a button" or "pulling a lever."
[0835] 7. Server: The emotion engine recognizes emotions from the user's facial expressions and voice, evaluates the data, determines the user's stress level and satisfaction, and reflects this in the operating instructions.
[0836] 8. Server: Generates a natural language instruction manual based on the extracted important actions and emotion data. The generated instruction manual includes specific steps such as "1. Turn on the power" and "2. Select the settings menu."
[0837] 9. Server: The generated text data is organized into a specified manual format, categorizing it into sections and adding tables and charts as needed.
[0838] 10. Server: The completed procedure manual is output in PDF format and sent to the user's device. The procedure manual is attached to an email or a download link is provided.
[0839] 11. Server: Collects user feedback and the system continues to learn. Based on the feedback, it will be reflected in the next generation of the procedure manual.
[0840] Specific examples
[0841] When creating an operating manual for a new machine in a factory, a user uses a smartphone to film how to operate the new machine in the factory. The video data is saved by the device and uploaded to a server via the Internet. The server receives the video data, analyzes the data using an AI analysis engine and an emotion engine, and generates a manual in natural language. This manual is created in PDF format and sent to the user. Feedback provided by the user after using the manual is used to help generate manuals in the future.
[0842] Prompt Sentence Examples
[0843] "I recorded the machine's operating procedures on my smartphone. Please generate an operating manual using this video data. Please list important actions and customize the manual with consideration for the user's emotions."
[0844] By inputting this prompt into the generative AI model, the system can efficiently generate instructions.
[0845] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0846] Step 1:
[0847] Recording Stage
[0848] User: Record work with a camera. Users record the process from start to finish using a smartphone or dedicated recording device. For example, record the process of learning how to operate a new machine. During this process, it is recommended to adjust the camera angle and lighting.
[0849] Input: User operation during work
[0850] Output: Recorded video data
[0851] Step 2:
[0852] Data storage
[0853] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, it automatically saves the video data to its internal storage. It notifies the user that the save is complete. The saved data is usually in MP4 or MOV format.
[0854] Input: Recorded video data
[0855] Output: Video data stored in the internal storage
[0856] Step 3:
[0857] Video upload
[0858] Device: Video data is sent to the server via the internet. Data is uploaded to the specified server via Wi-Fi or mobile network. The transfer progress is displayed on the screen and the user is notified when the transfer is complete. Data is encrypted during transfer as a security measure.
[0859] Input: Video data stored in the internal storage
[0860] Output: Video data sent to the server
[0861] Step 4:
[0862] Video reception and preprocessing
[0863] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. After saving, the data is checked for consistency and a simple check is performed to see if there are any missing or corrupted data. This preprocessing also includes checking the frame rate and renaming the data.
[0864] Input: Video data sent to the server
[0865] Output: Pre-processed video data
[0866] Step 5:
[0867] Video Analysis
[0868] Server: Sends video data to the AI analysis engine. The server divides the data into frames and performs preprocessing to extract important parts. It identifies important objects and actions in each frame.
[0869] Input: Preprocessed video data
[0870] Output: Object and movement data for each frame analyzed by the AI analysis engine
[0871] Server: The AI analysis engine analyzes the video data, identifies and lists important actions and procedures, for example, recognizing operations such as "pressing a button" or "pulling a lever."
[0872] Input: Object and movement data per frame
[0873] Output: A list of extracted important actions and steps
[0874] Step 6:
[0875] Emotion analysis
[0876] Server: Analyzes the user's facial expressions and voice in the video data and uses an emotion engine to recognize emotions. The server evaluates the user's stress level and satisfaction level and reflects this data in the analysis results. Analyzing subtle changes in facial expressions and tone of voice in particular enables more accurate emotion recognition.
[0877] Input: User's facial expressions and voice in video data
[0878] Output: User emotion data
[0879] Server: Adjust the content and tone of instructions based on the perceived emotion, for example adding more detailed explanations or additional information where the user is feeling stressed.
[0880] Input: User sentiment data, list of important actions and steps
[0881] Output: An outline of the procedure manual with the sentiment data reflected
[0882] Step 7:
[0883] Text Generation
[0884] Server: Generates instructions in natural language based on the extracted important actions and emotion data. AI analyzes the action list and generates specific steps such as "1. Turn on the power" and "2. Select the settings menu." Appropriate explanations and cautions are added to each step.
[0885] Input: Instructions outline with sentiment data reflected
[0886] Output: Natural language generated instructions
[0887] Server: Prepares the generated text data into a predetermined manual format. Formats the text according to a template, categorizes it into sections, and adds tables and charts as needed.
[0888] Input: Natural language generated instructions
[0889] Output: Formatted instructions
[0890] Step 8:
[0891] Manual Output
[0892] Server: Prints the completed manual in PDF format and sends it to the user's device. The instructions are attached to an email or a download link is provided. Adds tables and diagrams as needed (if necessary).
[0893] Input: Formatted instructions
[0894] Output: Instructions in PDF format provided to users
[0895] Step 9:
[0896] Feedback and Learning
[0897] Server: Collects feedback from users and the system continues to learn. Based on the feedback, it reflects it in the next generation of the procedure manual. Specifically, it analyzes evaluation points and comments and collects data to improve the procedure manual.
[0898] Input: User feedback
[0899] Output: Learning data that will be useful for the next procedure generation
[0900] (Application example 2)
[0901] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0902] Generating work instructions is important for improving the work efficiency of factory robots. However, conventional methods generate uniform instructions, making it difficult to provide personalized instructions that take into account the emotions and situations of individual users. Furthermore, analyzing work content without considering the user's emotions can lead to the risk of overlooking areas that cause stress to the user or areas for improvement.
[0903] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data, means for extracting important actions and objects using an AI engine that analyzes the acquired video data, means for using an emotion analysis engine that recognizes the user's emotions and reflects the results in generating a procedure manual, means for generating a procedure manual in natural language based on the extracted data and emotion data, and means for outputting the generated procedure manual in a predetermined format and providing it to the user. This makes it possible to provide a personalized procedure manual that takes into account the user's work situation and emotions.
[0904] "Video data" refers to a series of image frames captured by a camera, and is the digital data that is the subject of analysis.
[0905] "AI Engine" is a software platform for analyzing and extracting data using artificial intelligence technology.
[0906] "Critical actions" refer to actions or behaviors that are particularly essential in work procedures or processes.
[0907] "Object" refers to an object, such as a person or object, that is recognized in video data.
[0908] An "emotion analysis engine" is a software platform that recognizes emotions from a user's facial expressions and voice in video data and extracts that emotional state as data.
[0909] A "procedure manual" refers to a document that provides detailed instructions for a specific task or operation.
[0910] A "natural language" is a language that humans use on a daily basis, and is a written expression used to describe specific processes or instructions.
[0911] "Format" refers to a template or form that defines the structure and format of a document or data.
[0912] A "server" is a computer system that stores data, processes data, and provides services over a network.
[0913] "Analysis" is the process of breaking down data into detail to reveal its meaning and structure.
[0914] This invention is a system for automatically generating personalized operation manuals that incorporate the emotions of users by utilizing factory robots. The system of the present invention captures video of the user working, analyzes it using an AI engine, and generates operation manuals based on the extracted data and emotion analysis. Specifically, the system performs the following processes.
[0915] System Overview
[0916] 1. Recording Stage
[0917] User: Users working in the factory record their work using cameras attached to the robots, and through this process, video data is acquired in real time.
[0918] Hardware: Video data is temporarily stored in the robot's internal storage.
[0919] 2. Uploading videos
[0920] Server: The robot sends video data to the server via an internet connection, either wireless LAN or wired data communication.
[0921] 3. Video Reception and Preprocessing
[0922] Server: The received video data is stored in the server's storage. Next, as preprocessing, the data consistency is checked and processing such as noise removal and frame division is performed.
[0923] 4. Video Analysis
[0924] Server: An AI engine using OpenCV and TensorFlow analyzes the video data and extracts important actions and objects, using algorithms for object recognition and motion detection.
[0925] 5. Emotion analysis
[0926] Server: Uses an emotion analysis engine such as Microsoft Azure Cognitive Services to recognize emotions from the user's facial expressions and voice, allowing it to assess whether the user is feeling stressed.
[0927] 6. Text Generation
[0928] Server: Using a generative AI model such as GPT-4, a natural language instruction manual is generated based on the extracted key actions and emotion data. The instruction manual is then formatted in a predetermined format.
[0929] 7. Manual Output
[0930] Server: Generates the completed procedure manual in PDF format and provides it to the user by displaying it on the robot's display or by sending it to the user's email address.
[0931] 8. Feedback and Learning
[0932] Server: Collects feedback from users after using the procedure and uses that data to improve the analysis model for the next run, allowing the system to continuously learn and improve the quality of the procedure.
[0933] Specific examples
[0934] For example, when inspecting products in a factory, a robot can record the work, analyze the video, and compile a procedure manual for the next inspection method, taking the user's feelings into consideration. This can increase work efficiency while reducing user stress.
[0935] Prompt Sentence Examples
[0936] "Please record the product inspection process in the factory with a camera and generate the next inspection procedure based on the analysis results."
[0937] This system can provide flexible instructions that reflect the user's emotions and specific work content, which is expected to improve work efficiency and user satisfaction.
[0938] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0939] Step 1:
[0940] User: A user working in a factory records their work using a camera attached to the robot.
[0941] Input: The video data you are working with.
[0942] Output: Captured working video data.
[0943] Specific operation: The camera captures video data in real time and stores it in the robot's built-in storage.
[0944] Step 2:
[0945] Server: The robot sends the video data to the server via an internet connection.
[0946] Input: Video data of work inside the robot.
[0947] Output: Video data stored on the server.
[0948] Specific operation: The robot sends video data to the server's upload endpoint using Wi-Fi or a wired connection.
[0949] Step 3:
[0950] Server: Preprocesses the received video data and passes it to the AI analysis engine.
[0951] Input: Received video data.
[0952] Output: Pre-processed video data.
[0953] Specific operation: The server performs preprocessing such as checking data consistency, removing noise, and splitting frames.
[0954] Step 4:
[0955] Server: An AI engine using OpenCV and TensorFlow analyzes video data and extracts important actions and objects.
[0956] Input: Preprocessed video data.
[0957] Output: Extracted important action and object data.
[0958] Specific actions: The AI engine runs object recognition and action detection algorithms to list important actions and objects.
[0959] Step 5:
[0960] Server: Uses an emotion analysis engine to recognize emotions from the user's facial expressions and voice.
[0961] Input: Preprocessed video and audio data.
[0962] Output: User emotion data.
[0963] How it works: An emotion analysis engine such as Microsoft Azure Cognitive Services analyzes facial expressions and vocal tones to digitize the user's emotional state.
[0964] Step 6:
[0965] Server: Using a generative AI model, it generates instructions in natural language based on the extracted key actions and emotion data.
[0966] Input: Important action and emotion data.
[0967] Output: Instructions generated in natural language.
[0968] How it works: A generative AI model such as GPT-4 analyzes key action lists and emotional data to generate instructions in natural language.
[0969] Step 7:
[0970] Server: Outputs the generated procedure manual in a specified format and provides it to the user.
[0971] Input: Instructions generated in natural language.
[0972] Output: Instructions in PDF format.
[0973] Specific operation: The server formats the instruction manual in PDF format and displays it on the robot's display or sends it to the user by email.
[0974] Step 8:
[0975] Server: Collects feedback from users after using the procedure and uses that data to improve the next analysis model.
[0976] Input: User feedback data.
[0977] Output: An improved analytical model.
[0978] How it works: The server stores the feedback data in a database and retrains the generative AI model (GPT-4) to improve the quality of future instructions.
[0979] Through the above processing steps, the present invention automatically generates a personalized procedure manual that takes the user's feelings into consideration, thereby improving work efficiency and user satisfaction.
[0980] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0981] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0982] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0983] [Third embodiment]
[0984] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0985] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0986] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0987] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0988] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0989] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0990] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0991] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0992] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0993] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0994] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0995] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0996] As an embodiment of the present invention, a system including the following processes is provided.
[0997] System Overview
[0998] The user records the work content with a camera, and the video data is received by the terminal and sent to the server. The server analyzes the received video data and extracts important actions and objects. Then, based on the extracted data, a procedure manual is generated in natural language and output in a specified format. A specific embodiment is shown below.
[0999] Program Overview
[1000] 1. Recording Stage
[1001] User: Record the work process with a camera. Specifically, the user records the entire process from start to finish using a camera-equipped smartphone or dedicated recording device.
[1002] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, the video data is automatically saved to the internal storage.
[1003] 2. Uploading videos
[1004] Device: Sends saved video data to the server via the Internet. The device sends saved video data to the specified server upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to inform the user of progress.
[1005] 3. Video Reception and Preprocessing
[1006] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. The server checks the consistency of the video data and performs a simple check to see if there are any missing or corrupted data.
[1007] 4. Video Analysis
[1008] Server: Sends video data to the AI analysis engine. The server divides the stored video data into frames and performs pre-processing to extract the important parts of the frames.
[1009] Server: The AI analysis engine analyzes the video data. It performs frame-by-frame object recognition, motion detection, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[1010] 5. Text Generation
[1011] Server: Generates a procedure manual in natural language based on the extracted important actions and objects. The server analyzes the action list extracted by AI and generates procedures such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[1012] Server: Formats the generated text data into a manual. Formats the text according to the manual template and divides it into sections.
[1013] 6. Manual Output
[1014] Server: Generates the finished manual in PDF or other specified formats. Converts the entire manual to PDF and adds tables, diagrams, screenshots, etc. (if needed).
[1015] Server: Sends the generated manual to the user's device. The completed manual is sent to the user as an email attachment or a dedicated download link is provided.
[1016] Specific examples
[1017] As a specific example, a situation in which an operating manual for a machine in a factory is to be created will be described.
[1018] 1. Recording Stage
[1019] User: A factory worker uses his smartphone to film how to operate a new machine.
[1020] Device: Your smartphone will save this footage to its internal storage.
[1021] 2. Uploading videos
[1022] Device: The smartphone uses Wi-Fi to upload video data to a designated server.
[1023] 3. Video Reception and Preprocessing
[1024] Server: The server receives the video data and performs preprocessing for AI analysis.
[1025] 4. Video Analysis
[1026] Server: The server sends the video data to an AI analysis engine, which recognizes and lists the actions in each frame.
[1027] 5. Text Generation
[1028] Server: Based on the extracted actions, generate a procedure manual in natural language and format it into a manual.
[1029] 6. Manual Output
[1030] Server: Generates the completed manual in PDF format and sends it to the worker via email. This allows users to easily create work manuals and quickly update them.
[1031] This significantly reduces time and effort, and enables the latest manuals to be provided promptly.
[1032] The processing flow will be explained below.
[1033] Step 1:
[1034] Users record their work with a camera. Specifically, they acquire a smartphone with a camera or a dedicated recording device and record the entire process from start to finish.
[1035] Step 2:
[1036] The device will temporarily store the recorded video data. After the recording device or smartphone has finished recording, the video data will automatically be saved to its internal storage.
[1037] Step 3:
[1038] The device sends the stored video data to the server via the Internet. The device sends the stored video data to the specified server's upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[1039] Step 4:
[1040] The server saves the received video data in the server's storage. The server organizes and saves the received data in a dedicated directory. It checks the consistency of the video data and performs a simple check to see if there are any missing or damaged data.
[1041] Step 5:
[1042] The server sends the video data to an AI analysis engine, which divides the stored video data into frames and performs preprocessing to extract the important parts of the frames.
[1043] Step 6:
[1044] The server's AI analysis engine analyzes the video data, performing frame-by-frame object recognition, motion detection, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever, etc.) and steps, and creates a sequenced list of them.
[1045] Step 7:
[1046] The server generates a procedure manual in natural language based on the extracted important actions and objects. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[1047] Step 8:
[1048] The server formats the generated text data into a manual, formatting the text according to the manual template and dividing it into sections.
[1049] Step 9:
[1050] The server generates the finished manual in PDF or other specified format, converting the entire manual to PDF and adding tables, diagrams, screenshots, etc. (if needed).
[1051] Step 10:
[1052] The server sends the generated manual to the user's device, either by attaching the completed manual to an email or by providing a dedicated download link.
[1053] Example 1
[1054] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1055] There has been no system that can easily record video data of work content and automatically generate procedure manuals based on that data. In particular, there is a need for an efficient and accurate process that can analyze the video data, extract important actions and objects, and generate procedure manuals. This has led to the need for a system that can reduce work hours and labor, standardize work procedures, and enable rapid update support.
[1056] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1057] In this invention, the server includes: [means for acquiring video data;] [means for extracting important actions and objects using an AI analysis unit that analyzes the acquired video data; and] [means for generating a procedure manual in natural language based on the extracted data. This makes it possible to automatically generate a procedure manual based on the video data of the work content and output it in a specified format.
[1058] "Video data" refers to digital information in the form of moving images captured using a camera or other imaging device.
[1059] A "server" is a computing device used to receive, analyze, store, and transmit data over a network.
[1060] "AI Analysis Unit" means a collection of software and hardware that implements artificial intelligence technology used to analyze video data and identify and extract significant actions and objects.
[1061] "Critical actions" refer to actions or operations that require special attention when carrying out business procedures.
[1062] "Objects" refer to specific items or devices that exist within the video data and are required to explain business procedures.
[1063] A "procedure manual" is a document that explains business procedures in a specific and orderly manner as a series of steps.
[1064] "Means for acquiring" refers to methods and devices for collecting video data using cameras and recording devices.
[1065] "Means for generating in natural language" refers to the method and process for generating a procedure manual in a natural language that is easy for humans to understand, based on the extracted data.
[1066] A "prescribed format" is a rule that ensures that procedures are organized in a certain format and displayed in a way that is visually easy to understand.
[1067] As an embodiment of this invention, we provide a system that automatically captures and analyzes work procedures at companies or work sites and generates procedure manuals. Specifically, this system allows users to record work content with a camera, and the video data is received by a terminal and sent to a server. The details of the system are described below.
[1068] Recording Stage
[1069] User:
[1070] Work is recorded with a camera. Specifically, the user uses a camera-equipped smartphone or a dedicated recording device to record work from start to finish. For example, a factory worker uses a smartphone to record how to operate a new machine.
[1071] Device:
[1072] Recorded video data is saved to the internal storage. After recording is complete, the smartphone automatically saves the video data temporarily.
[1073] Video upload
[1074] Device:
[1075] The saved video data is sent to a server via the Internet. The device uploads the video data to the specified server via Wi-Fi or a mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[1076] Video reception and preprocessing
[1077] server:
[1078] The server stores the received video data in a dedicated directory, checks for missing or corrupted data, and checks the data for consistency. Organizes the video data into batches for analysis.
[1079] Video Analysis
[1080] server:
[1081] The server divides the stored video data into frames and sends them to the AI analysis unit. As preprocessing, the video resolution is adjusted appropriately and unnecessary frames are removed. The AI analysis unit (e.g., TensorFlow or OpenCV) is used to perform object recognition and motion detection in each frame, and convert text and voice as needed. For example, scenes in which the user presses a button or pulls a lever are extracted and listed.
[1082] Text Generation
[1083] server:
[1084] Generates instructions in natural language based on the extracted important actions and objects. Using an AI model (e.g., GPT-3), it creates steps such as "1. Turn on the power" and "2. Select the settings menu" from the listed actions. Provides appropriate explanations and cautions for each step. For example, it adds instructions such as "Check the power cord and be careful not to press the buttons too hard."
[1085] server:
[1086] The generated text data is formatted into a manual, with each step divided into sections and visually organized using bulleted and numbered lists.
[1087] Manual Output
[1088] server:
[1089] Generate the finished manual in PDF or other format of your choice. Use tools such as Adobe Acrobat or LaTeX to add tables, charts, and screenshots to create a visually appealing manual.
[1090] server:
[1091] The generated manual is sent to the user's device. The completed manual is sent as an email attachment to the user, or a dedicated download link is provided.
[1092] Specific examples
[1093] As a specific example, we will explain a situation in which an operating manual for a new machine is being created in a factory.
[1094] User:
[1095] Factory workers use their smartphones to record themselves operating a new machine.
[1096] Device:
[1097] The smartphone saves this footage to its internal storage.
[1098] Device:
[1099] The smartphone uses Wi-Fi to upload the video data to a designated server.
[1100] server:
[1101] The server receives the video data and performs preprocessing for AI analysis.
[1102] server:
[1103] The server sends the video data to an AI analysis unit, which recognizes and lists the movements in each frame.
[1104] server:
[1105] Based on the extracted actions, a procedure manual is generated in natural language and formatted into a manual.
[1106] server:
[1107] The completed manual is generated in PDF format and sent to the worker via email, allowing users to easily create work manuals and quickly update them.
[1108] Examples of prompt statements
[1109] "Please film how to operate the new machine and create a manual."
[1110] "Please tell me the procedure for uploading video data to the server."
[1111] "Please explain in detail how to generate the procedure manual."
[1112] This will provide a system that standardizes business procedures, improves efficiency, and enables rapid update support.
[1113] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1114] Step 1:
[1115] User: Records work with a camera. The user uses a smartphone with a camera or other recording device to record the entire process from start to finish. For example, a factory worker uses a smartphone to record how to operate a new machine. The user takes care to clearly capture specific operating steps (e.g., how to press a button, pull a lever). Input: The device used is a smartphone with a camera. Output: Recorded video data.
[1116] Step 2:
[1117] Device: After recording is complete, the device saves the video data to its internal storage. At the same time, the device sends the video data to the server via Wi-Fi or the mobile network. During uploading, the device displays a progress bar to inform the user of the transfer status. Input: Recorded video data. Output: Video data sent to the server.
[1118] Step 3:
[1119] Server: The server stores the received video data in a dedicated directory in an organized manner. It first checks the data for consistency and performs a simple check to see if there are any missing or corrupted data. If the data is OK, it is broken down into frames for analysis and the resolution is adjusted if necessary. Input: Video data sent to the server. Output: Video data that has been split into frames and pre-processed.
[1120] Step 4:
[1121] Server: Sends preprocessed video data to an AI analysis unit (e.g., TensorFlow or OpenCV), which performs object recognition and action detection for each frame, and converts text and voice as needed. For example, it extracts and lists scenes in which a button is pressed or a lever is pulled. Input: Preprocessed video data for each frame. Output: A list of extracted important actions and objects.
[1122] Step 5:
[1123] Server: Based on the extracted important actions and objects, a generative AI model (e.g., GPT-3) is used to generate instructions in natural language. Detailed explanations and precautions are added for each action (e.g., pressing a button), and steps such as "1. Turn on the power" and "2. Select the settings menu" are formed. Input: A list of extracted actions and objects. Output: Instructions generated in natural language.
[1124] Step 6:
[1125] Server: Formats the generated procedure manual into a manual format and outputs it in PDF or other specified format. Tables, diagrams, and in some cases screenshots are added to the procedure manual to make it visually easy to understand. Input: Procedure manual generated in natural language. Output: Procedure manual formatted in PDF or other specified format.
[1126] Step 7:
[1127] Server: Provides the completed procedure to the user. The server sends the procedure as an email attachment or provides a dedicated download link. It may also store it in a database. Input: Procedure formatted in PDF or other specified format. Output: Procedure provided to the user.
[1128] Through this series of processes, a system is realized that can automatically generate procedure manuals from video data of work content and provide them to users quickly and efficiently, making it possible to standardize work procedures and respond to updates quickly.
[1129] (Application example 1)
[1130] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1131] In modern factories and manufacturing sites, specialized robot maintenance work is complex, and creating procedure manuals requires a great deal of time and effort. Therefore, there is a need for an effective automatic procedure manual generation system to improve the efficiency and standardization of maintenance work. However, existing manual procedure manual creation methods have issues such as difficulty in recording and analyzing work, and difficulty in quickly reflecting the latest information.
[1132] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1133] In this invention, the server includes: [means for acquiring video data using a device with a camera; [means for extracting important actions and objects using an AI analysis engine that analyzes the acquired video data; and [means for generating a procedure manual in natural language based on the extracted data.] This makes it possible [to achieve work efficiency and standardization, and to quickly reflect the latest information in the procedure manual].
[1134] "Video data" is a series of image information captured by a device with a camera, and is digital data that includes details of actions and objects.
[1135] "AI analysis engine" is a general term for software and hardware that uses artificial intelligence technology to analyze video data and extract important actions and objects.
[1136] A "procedure" is a document that provides step-by-step instructions for performing a specific task or operation.
[1137] "Natural language" is a language used by humans on a daily basis and a set of textual forms generated by computer programs.
[1138] A "generative AI model" is an artificial intelligence model that learns large amounts of data in advance to generate instruction manuals and other text data.
[1139] A "prompt" is a series of phrases or documents that are input to a generative AI model to determine the content and format of the generated text.
[1140] A "server" is a computer that receives requests from client devices over a network and provides data processing and data storage functions.
[1141] "Preprocessing" refers to the process of removing noise from video data and converting its format before analysis by the AI analysis engine.
[1142] A "data processing device" is a computer system for analyzing video data and extracting important information.
[1143] A "template" is a predefined format or style guide for formatting procedures and reports.
[1144] "Digital document format" means a document format that can be stored and displayed electronically, including PDF.
[1145] MODE FOR CARRYING OUT THE INVENTION
[1146] System Overview
[1147] As an embodiment of the present invention, we provide a system that performs the following process. In this system, a user records work content using a camera-equipped smartphone and sends the video data to a server. The server analyzes the received video data and extracts important actions and objects. Then, based on the extracted data, it generates a procedure manual in natural language and outputs it in a specified format.
[1148] Explanation of program processing
[1149] Hardware and software used
[1150] Smartphone: A smart device used as a recording device.
[1151] Server: A computer used for data analysis that receives requests from client devices over the network and processes the data.
[1152] OpenCV: A library for recording and processing video.
[1153] Requests library: Used to send video data to the server.
[1154] AI analysis engine: An engine for analyzing video data and extracting important actions and objects.
[1155] PDF generation library: A library for converting instructions into PDF format (e.g., ReportLab).
[1156] Processing flow
[1157] 1. Recording: The user records the robot's maintenance work using a smartphone. The recorded data is saved in the smartphone's internal storage.
[1158] 2. Video Upload: Once the recording is complete, the device (smartphone) sends the saved video data to the server via Wi-Fi or mobile network. This data transmission is performed using the Requests library.
[1159] 3. Data processing: The server receives the transmitted video data and performs preprocessing, which includes noise reduction and format conversion. The AI analysis engine then divides the video data into frames and extracts important actions and objects.
[1160] 4. Procedure generation: A procedure manual is generated in natural language based on the data extracted by the AI analysis engine. A generative AI model is used here. A prompt sentence for generating the procedure manual is also created. By inputting this prompt sentence into the generative AI model, automatically generated text can be obtained.
[1161] 5. Output of procedure manual: The generated procedure manual is output in PDF format using a PDF generation library, and the output PDF is provided to the user.
[1162] Specific examples
[1163] For example, let's say a maintenance task of adding lubricant to the joints of a robot is recorded in a factory. This recorded data is sent to a server, which then uses an analysis engine to analyze the video and extract important actions (such as "remove the cover," "prepare the lubricant," and "apply lubricant to the specified location"). Based on the extracted actions, the AI model automatically generates a procedure manual and provides it to the engineer in PDF format.
[1164] Prompt Sentence Examples
[1165] Video data URL: {video_url}
[1166] Analysis output:
[1167] Object Recognition
[1168] Action Detection
[1169] Procedure generation
[1170] The above is an embodiment of the present invention. This system makes it possible to improve the efficiency and standardize maintenance work, and to quickly update the latest information in the procedure manual.
[1171] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1172] Program processing flow
[1173] Step 1:
[1174] Recording
[1175] The user records the robot's maintenance work using a smartphone. At this time, a camera-equipped device is used to capture the entire process from start to finish. The recorded data (video data) is saved in the smartphone's internal storage.
[1176] Input: Maintenance work video
[1177] Output: Video data stored in the internal storage
[1178] Step 2:
[1179] Video upload
[1180] Once recording is complete, the device (smartphone) sends the saved video data to the server via the Internet. The device uses the Requests library to send the video data to the specified server's upload endpoint. At this time, the file transfer status is displayed to inform the user of the progress.
[1181] Input: Video data stored in the internal storage
[1182] Output: Video data sent to the server
[1183] Step 3:
[1184] Data reception and preprocessing
[1185] The server receives the transmitted video data and stores it in storage. Upon receiving the data, it performs pre-processing such as noise removal and format conversion. This process converts the data into a format that can be smoothly processed by the analysis engine.
[1186] Input: Video data sent to the server
[1187] Output: Pre-processed video data
[1188] Step 4:
[1189] Video Analysis
[1190] The preprocessed video data is passed to the server's AI analysis engine, which divides the video data into frames and extracts important actions and objects from each frame. Here, it performs tasks such as object recognition and action detection for each frame, and creates a list of important actions.
[1191] Input: Preprocessed video data
[1192] Output: Data listing important actions and objects
[1193] Step 5:
[1194] Procedure generation
[1195] Based on the data listing important actions and objects, the server uses a generative AI model to generate instructions in natural language. A prompt is entered into the generative AI model, which then automatically generates text. An example of a prompt might be something like "Video data URL: {video_url}\nAnalysis output:\n- Object recognition\n- Action detection\n- Instruction manual generation."
[1196] Input: Data listing important actions and objects, prompts
[1197] Output: Natural language instructions
[1198] Step 6:
[1199] Output of procedure manual
[1200] The server converts the generated instructions into PDF format using a PDF generation library, and the generated PDF is sent to the user's device or a download link is provided.
[1201] Input: Natural language instructions
[1202] Output: Instructions in PDF format
[1203] By following the above steps, a procedure manual can be automatically generated from a series of maintenance tasks and provided efficiently.
[1204] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1205] As an embodiment of the present invention, we provide a system that includes the following processing. In addition to acquiring and analyzing video data, this system recognizes the user's emotions and reflects them in the generation of procedure manuals, thereby providing more personalized manuals.
[1206] System Overview
[1207] The user records their work with a camera, and the video data is received by the device and sent to the server. The server analyzes the received video data and extracts important actions and objects. It then generates a procedure manual in natural language based on the extracted data and outputs it in a specified format. The system also incorporates an emotion engine that recognizes the user's emotions and reflects the results in the generation of the procedure manual.
[1208] Program Overview
[1209] 1. Recording Stage
[1210] User: Record work content with a camera. Specifically, the user acquires a smartphone with a camera or a dedicated recording device and records the process of work from start to finish.
[1211] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, the video data is automatically saved to the internal storage.
[1212] 2. Uploading videos
[1213] Device: Sends saved video data to the server via the Internet. The device sends saved video data to the specified server upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to inform the user of progress.
[1214] 3. Video Reception and Preprocessing
[1215] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. The server checks the consistency of the video data and performs a simple check to see if there are any missing or corrupted data.
[1216] 4. Video Analysis
[1217] Server: Sends video data to the AI analysis engine. The server divides the stored video data into frames and performs pre-processing to extract the important parts of the frames.
[1218] Server: The AI analytics engine analyzes the video data. It performs frame-by-frame object recognition, motion detection, emotion analysis, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly fashion.
[1219] 5. Emotion analysis
[1220] Server: Analyzes the user's facial expressions and voice in the video data and uses an emotion engine to recognize emotions. Evaluates the user's stress level and satisfaction level and reflects that data in the analysis results.
[1221] Server: Adjust the content and tone of instructions based on the perceived emotion, for example adding more detailed explanations or additional information where the user is feeling stressed.
[1222] 6. Text Generation
[1223] Server: Generates instructions in natural language based on the extracted important actions and user emotion data. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[1224] Server: Formats the generated text data into a manual. Formats the text according to the manual template and divides it into sections.
[1225] 7. Manual Output
[1226] Server: Generates the finished manual in PDF or other specified formats. Converts the entire manual to PDF and adds tables, diagrams, screenshots, etc. (if needed).
[1227] Server: Sends the generated manual to the user's device. The completed manual is sent to the user as an email attachment or a dedicated download link is provided.
[1228] 8. Feedback and Learning
[1229] Server: Records the user's emotion recognition results and reflects them in the next procedure manual generation. Based on the recorded emotion data, the recommended procedure manual content is personalized.
[1230] Users: Provide feedback after using the manual, so the system continually learns and improves the quality of the procedure manual.
[1231] Specific examples
[1232] As a specific example, we will explain a situation in which an operating manual for a new machine is being created in a factory.
[1233] 1. Recording Stage
[1234] User: A factory worker uses his smartphone to film how to operate a new machine.
[1235] Device: Your smartphone will save this footage to its internal storage.
[1236] 2. Uploading videos
[1237] Device: The smartphone uses Wi-Fi to upload video data to a designated server.
[1238] 3. Video Reception and Preprocessing
[1239] Server: The server receives the video data and performs preprocessing for AI analysis.
[1240] 4. Video Analysis
[1241] Server: The server sends the video data to an AI analysis engine, which recognizes and lists the actions in each frame.
[1242] 5. Emotion analysis
[1243] Server: Recognizes emotions from the user's facial expressions and voice and adjusts the contents of the instruction manual.
[1244] 6. Text Generation
[1245] Server: Generates instructions in natural language based on the extracted action and emotion data and formats them into a manual.
[1246] 7. Manual Output
[1247] Server: Generates the completed manual in PDF format and sends it to the worker's email address.
[1248] 8. Feedback and Learning
[1249] Server: The system continues to learn based on user feedback and improves the quality of the instructions.
[1250] This will significantly reduce time and effort, and enable the rapid provision of up-to-date manuals that take user feelings into consideration.
[1251] The processing flow will be explained below.
[1252] Step 1:
[1253] Users record their work with a camera. Specifically, they record the entire process from start to finish using a camera-equipped smartphone or a dedicated recording device.
[1254] Step 2:
[1255] The device will temporarily store the recorded video data. After the recording device or smartphone has finished recording, the video data will automatically be saved to its internal storage.
[1256] Step 3:
[1257] The device sends the stored video data to the server via the Internet. The device sends the stored video data to the specified server's upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[1258] Step 4:
[1259] The server saves the received video data in the server's storage. The server organizes and saves the received data in a dedicated directory. It checks the consistency of the video data and performs a simple check to see if there are any missing or damaged data.
[1260] Step 5:
[1261] The server sends the video data to an AI analysis engine, which divides the stored video data into frames and performs preprocessing to extract the important parts of the frames.
[1262] Step 6:
[1263] The server's AI analysis engine analyzes the video data, performing frame-by-frame object recognition, motion detection, emotion analysis, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[1264] Step 7:
[1265] The server uses an emotion engine to analyze the user's facial expressions and voice in the video data and recognize emotions, such as whether the user is confused or happy.
[1266] Step 8:
[1267] The server adjusts the content and tone of the instructions based on the perceived emotion, for example adding detailed explanations or additional information if it detects that the user is stressed.
[1268] Step 9:
[1269] The server generates a procedure manual in natural language based on the extracted important actions and user emotion data. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[1270] Step 10:
[1271] The server formats the generated text data into a manual, formatting the text according to the manual template and dividing it into sections.
[1272] Step 11:
[1273] The server generates the finished manual in PDF or other specified format, converting the entire manual to PDF and adding tables, diagrams, screenshots, etc. (if needed).
[1274] Step 12:
[1275] The server sends the generated manual to the user's device, either by attaching the completed manual to an email or by providing a dedicated download link.
[1276] Step 13:
[1277] The server records the user's emotion recognition results and reflects them in the next procedure manual generation. The content of the procedure manual is personalized using the user's past emotion data.
[1278] Step 14:
[1279] The server continues to learn from feedback after the manual is used, and by combining user feedback with emotional data, the quality of the manual is improved.
[1280] Example 2
[1281] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1282] Conventional procedure manual generation systems often create procedures without considering the user's emotions, resulting in a lack of means to identify and address parts that users find emotionally difficult. Furthermore, feedback is not effectively incorporated into the procedure manual generation process, resulting in inconsistent quality. This can result in procedures that are not suited to actual work.
[1283] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1284] In this invention, the server includes: [means for providing an emotion engine that analyzes emotions from the user's facial expressions and voice; [means for generating a procedure manual in natural language based on the extracted data and emotion data; and [means for collecting feedback from the user and allowing the system to learn.] This makes it possible [to provide information that takes the user's emotions into consideration, thereby improving the quality and usability of the procedure manual].
[1285] "Video data" refers to video files in which users record their work.
[1286] "AI engine" refers to an artificial intelligence algorithm that analyzes video data and extracts important actions and objects.
[1287] "Important actions" refer to the work steps and operations required to generate a procedure manual.
[1288] An "object" refers to an object that is recognized in video data.
[1289] An "emotion engine" refers to a system that recognizes and analyzes emotions from the user's facial expressions and voice in video data.
[1290] A "procedure manual" refers to a document that describes the steps a user takes to perform a specific task.
[1291] "Predetermined format" refers to the predetermined manner in which the generated procedure manual is displayed or printed.
[1292] "Terminal" refers to a device for acquiring and transmitting video data.
[1293] "Feedback" refers to evaluations and opinions about the system provided by users.
[1294] "Preprocessing" refers to the process of organizing and checking consistency of video data before analyzing it.
[1295] "PDF format" refers to the Portable Document Format, an electronic document format.
[1296] "Download link" refers to the URL for obtaining a file via the Internet.
[1297] MODE FOR CARRYING OUT THE INVENTION
[1298] The following hardware and software are primarily used in the embodiment of this invention. A user records work content using a smartphone or dedicated recording device, and the terminal temporarily stores this video data. The terminal then transmits the video data to a server, which analyzes the data using an AI engine and an emotion engine. A procedure manual is then generated in natural language based on the extracted information and output in a specified format. The procedure manual is output in PDF format and provided to the user. The system also collects user feedback and continues to learn.
[1299] Hardware and software used
[1300] Hardware
[1301] A smartphone or dedicated recording device
[1302] server
[1303] software
[1304] AI analysis engine (analysis of video data)
[1305] Emotion engine (user emotion analysis)
[1306] PDF generation software (generating instruction manuals)
[1307] Internet connection software (data transmission and reception)
[1308] Data processing and calculation flow
[1309] 1. User: Records work using a smartphone or dedicated recording device. When a user wants to record how to operate a new machine, they can record all the operating procedures at once.
[1310] 2. Device: Temporarily stores the recorded video data. When the recording is finished, the device automatically saves the data to the internal storage and notifies the user of the progress.
[1311] 3. Device: The video data is sent to the server via the Internet. The data is uploaded to the specified server via Wi-Fi or mobile network. The upload progress is displayed on the screen.
[1312] 4. Server: Stores the received video data in storage and checks the data consistency. It performs a simple check to see if there are any missing or corrupted data.
[1313] 5. Server: The video data is divided into frames, and preprocessing such as object recognition and motion detection is performed using an AI analysis engine.
[1314] 6. Server: The AI analysis engine analyzes the video data, identifies and lists important actions and procedures, such as "pressing a button" or "pulling a lever."
[1315] 7. Server: The emotion engine recognizes emotions from the user's facial expressions and voice, evaluates the data, determines the user's stress level and satisfaction, and reflects this in the operating instructions.
[1316] 8. Server: Generates a natural language instruction manual based on the extracted important actions and emotion data. The generated instruction manual includes specific steps such as "1. Turn on the power" and "2. Select the settings menu."
[1317] 9. Server: The generated text data is organized into a specified manual format, categorizing it into sections and adding tables and charts as needed.
[1318] 10. Server: The completed procedure manual is output in PDF format and sent to the user's device. The procedure manual is attached to an email or a download link is provided.
[1319] 11. Server: Collects user feedback and the system continues to learn. Based on the feedback, it will be reflected in the next generation of the procedure manual.
[1320] Specific examples
[1321] When creating an operating manual for a new machine in a factory, a user uses a smartphone to film how to operate the new machine in the factory. The video data is saved by the device and uploaded to a server via the Internet. The server receives the video data, analyzes the data using an AI analysis engine and an emotion engine, and generates a manual in natural language. This manual is created in PDF format and sent to the user. Feedback provided by the user after using the manual is used to help generate manuals in the future.
[1322] Prompt Sentence Examples
[1323] "I recorded the machine's operating procedures on my smartphone. Please generate an operating manual using this video data. Please list important actions and customize the manual with consideration for the user's emotions."
[1324] By inputting this prompt into the generative AI model, the system can efficiently generate instructions.
[1325] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1326] Step 1:
[1327] Recording Stage
[1328] User: Record work with a camera. Users record the process from start to finish using a smartphone or dedicated recording device. For example, record the process of learning how to operate a new machine. During this process, it is recommended to adjust the camera angle and lighting.
[1329] Input: User operation during work
[1330] Output: Recorded video data
[1331] Step 2:
[1332] Data storage
[1333] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, it automatically saves the video data to its internal storage. It notifies the user that the save is complete. The saved data is usually in MP4 or MOV format.
[1334] Input: Recorded video data
[1335] Output: Video data stored in the internal storage
[1336] Step 3:
[1337] Video upload
[1338] Device: Video data is sent to the server via the internet. Data is uploaded to the specified server via Wi-Fi or mobile network. The transfer progress is displayed on the screen and the user is notified when the transfer is complete. Data is encrypted during transfer as a security measure.
[1339] Input: Video data stored in the internal storage
[1340] Output: Video data sent to the server
[1341] Step 4:
[1342] Video reception and preprocessing
[1343] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. After saving, the data is checked for consistency and a simple check is performed to see if there are any missing or corrupted data. This preprocessing also includes checking the frame rate and renaming the data.
[1344] Input: Video data sent to the server
[1345] Output: Pre-processed video data
[1346] Step 5:
[1347] Video Analysis
[1348] Server: Sends video data to the AI analysis engine. The server divides the data into frames and performs preprocessing to extract important parts. It identifies important objects and actions in each frame.
[1349] Input: Preprocessed video data
[1350] Output: Object and movement data for each frame analyzed by the AI analysis engine
[1351] Server: The AI analysis engine analyzes the video data, identifies and lists important actions and procedures, for example, recognizing operations such as "pressing a button" or "pulling a lever."
[1352] Input: Object and movement data per frame
[1353] Output: A list of extracted important actions and steps
[1354] Step 6:
[1355] Emotion analysis
[1356] Server: Uses an emotion engine that analyzes the user's facial expressions and voice in the video data and recognizes their emotions. The server evaluates the user's stress level and satisfaction level and reflects this data in the analysis results. Analyzing subtle changes in facial expressions and tone of voice in particular enables more accurate emotion recognition.
[1357] Input: User's facial expressions and voice in video data
[1358] Output: User emotion data
[1359] Server: Adjust the content and tone of instructions based on the perceived emotion, for example adding more detailed explanations or additional information where the user is feeling stressed.
[1360] Input: User sentiment data, list of important actions and steps
[1361] Output: An outline of the procedure manual with the sentiment data reflected
[1362] Step 7:
[1363] Text Generation
[1364] Server: Generates instructions in natural language based on the extracted important actions and emotion data. AI analyzes the action list and generates specific steps such as "1. Turn on the power" and "2. Select the settings menu." Appropriate explanations and cautions are added to each step.
[1365] Input: Instructions outline with sentiment data reflected
[1366] Output: Natural language generated instructions
[1367] Server: Prepares the generated text data into a predetermined manual format. Formats the text according to a template, categorizes it into sections, and adds tables and charts as needed.
[1368] Input: Natural language generated instructions
[1369] Output: Formatted instructions
[1370] Step 8:
[1371] Manual Output
[1372] Server: Prints the completed manual in PDF format and sends it to the user's device. The instructions are attached to an email or a download link is provided. Adds tables and diagrams as needed (if necessary).
[1373] Input: Formatted instructions
[1374] Output: Instructions in PDF format provided to users
[1375] Step 9:
[1376] Feedback and Learning
[1377] Server: Collects feedback from users and the system continues to learn. Based on the feedback, it reflects it in the next generation of the procedure manual. Specifically, it analyzes evaluation points and comments and collects data to improve the procedure manual.
[1378] Input: User feedback
[1379] Output: Learning data that will be useful for the next procedure generation
[1380] (Application example 2)
[1381] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1382] Generating work instructions is important for improving the work efficiency of factory robots. However, conventional methods generate uniform instructions, making it difficult to provide personalized instructions that take into account the emotions and situations of individual users. Furthermore, analyzing work content without considering the user's emotions can lead to the risk of overlooking areas that cause stress to the user or areas for improvement.
[1383] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data, means for extracting important actions and objects using an AI engine that analyzes the acquired video data, means for using an emotion analysis engine that recognizes the user's emotions and reflects the results in generating a procedure manual, means for generating a procedure manual in natural language based on the extracted data and emotion data, and means for outputting the generated procedure manual in a predetermined format and providing it to the user. This makes it possible to provide a personalized procedure manual that takes into account the user's work situation and emotions.
[1384] "Video data" refers to a series of image frames captured by a camera, and is the digital data that is the subject of analysis.
[1385] "AI Engine" is a software platform for analyzing and extracting data using artificial intelligence technology.
[1386] "Critical actions" refer to actions or behaviors that are particularly essential in work procedures or processes.
[1387] "Object" refers to an object, such as a person or object, that is recognized in video data.
[1388] An "emotion analysis engine" is a software platform that recognizes emotions from a user's facial expressions and voice in video data and extracts that emotional state as data.
[1389] A "procedure manual" refers to a document that provides detailed instructions for a specific task or operation.
[1390] A "natural language" is a language that humans use on a daily basis, and is a written expression used to describe specific processes or instructions.
[1391] "Format" refers to a template or form that defines the structure and format of a document or data.
[1392] A "server" is a computer system that stores data, processes data, and provides services over a network.
[1393] "Analysis" is the process of breaking down data into detail to reveal its meaning and structure.
[1394] This invention is a system for automatically generating personalized operation manuals that incorporate the emotions of users by utilizing factory robots. The system of the present invention captures video of the user working, analyzes it using an AI engine, and generates operation manuals based on the extracted data and emotion analysis. Specifically, the system performs the following processes.
[1395] System Overview
[1396] 1. Recording Stage
[1397] User: Users working in the factory record their work using cameras attached to the robots, and through this process, video data is acquired in real time.
[1398] Hardware: Video data is temporarily stored in the robot's internal storage.
[1399] 2. Uploading videos
[1400] Server: The robot sends video data to the server via an internet connection, either wireless LAN or wired data communication.
[1401] 3. Video Reception and Preprocessing
[1402] Server: The received video data is stored in the server's storage. Next, as preprocessing, the data consistency is checked and processing such as noise removal and frame division is performed.
[1403] 4. Video Analysis
[1404] Server: An AI engine using OpenCV and TensorFlow analyzes the video data and extracts important actions and objects, using algorithms for object recognition and motion detection.
[1405] 5. Emotion analysis
[1406] Server: Uses an emotion analysis engine such as Microsoft Azure Cognitive Services to recognize emotions from the user's facial expressions and voice, allowing it to assess whether the user is feeling stressed.
[1407] 6. Text Generation
[1408] Server: Using a generative AI model such as GPT-4, a natural language instruction manual is generated based on the extracted key actions and emotion data. The instruction manual is then formatted in a predetermined format.
[1409] 7. Manual Output
[1410] Server: Generates the completed procedure manual in PDF format and provides it to the user by displaying it on the robot's display or by sending it to the user's email address.
[1411] 8. Feedback and Learning
[1412] Server: Collects feedback from users after using the procedure and uses that data to improve the analysis model for the next run, allowing the system to continuously learn and improve the quality of the procedure.
[1413] Specific examples
[1414] For example, when inspecting products in a factory, a robot can record the work, analyze the video, and compile a procedure manual for the next inspection method, taking the user's feelings into consideration. This can increase work efficiency while reducing user stress.
[1415] Prompt Sentence Examples
[1416] "Please record the product inspection process in the factory with a camera and generate the next inspection procedure based on the analysis results."
[1417] This system can provide flexible instructions that reflect the user's emotions and specific work content, which is expected to improve work efficiency and user satisfaction.
[1418] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1419] Step 1:
[1420] User: A user working in a factory records their work using a camera attached to the robot.
[1421] Input: The video data you are working with.
[1422] Output: Captured working video data.
[1423] Specific operation: The camera captures video data in real time and stores it in the robot's built-in storage.
[1424] Step 2:
[1425] Server: The robot sends the video data to the server via an internet connection.
[1426] Input: Video data of work inside the robot.
[1427] Output: Video data stored on the server.
[1428] Specific operation: The robot sends video data to the server's upload endpoint using Wi-Fi or a wired connection.
[1429] Step 3:
[1430] Server: Preprocesses the received video data and passes it to the AI analysis engine.
[1431] Input: Received video data.
[1432] Output: Pre-processed video data.
[1433] Specific operation: The server performs preprocessing such as checking data consistency, removing noise, and splitting frames.
[1434] Step 4:
[1435] Server: An AI engine using OpenCV and TensorFlow analyzes video data and extracts important actions and objects.
[1436] Input: Preprocessed video data.
[1437] Output: Extracted important action and object data.
[1438] Specific actions: The AI engine runs object recognition and action detection algorithms to list important actions and objects.
[1439] Step 5:
[1440] Server: Uses an emotion analysis engine to recognize emotions from the user's facial expressions and voice.
[1441] Input: Preprocessed video and audio data.
[1442] Output: User emotion data.
[1443] How it works: An emotion analysis engine, such as Microsoft Azure Cognitive Services, analyzes facial expressions and vocal tones to digitize the user's emotional state.
[1444] Step 6:
[1445] Server: Using a generative AI model, it generates instructions in natural language based on the extracted key actions and emotion data.
[1446] Input: Important action and emotion data.
[1447] Output: Instructions generated in natural language.
[1448] How it works: A generative AI model such as GPT-4 analyzes key action lists and emotional data to generate instructions in natural language.
[1449] Step 7:
[1450] Server: Outputs the generated procedure manual in a specified format and provides it to the user.
[1451] Input: Instructions generated in natural language.
[1452] Output: Instructions in PDF format.
[1453] Specific operation: The server formats the instruction manual in PDF format and displays it on the robot's display or sends it to the user by email.
[1454] Step 8:
[1455] Server: Collects feedback from users after using the procedure and uses that data to improve the next analysis model.
[1456] Input: User feedback data.
[1457] Output: An improved analytical model.
[1458] How it works: The server stores the feedback data in a database and retrains the generative AI model (GPT-4) to improve the quality of future instructions.
[1459] Through the above processing steps, the present invention automatically generates a personalized procedure manual that takes the user's feelings into consideration, thereby improving work efficiency and user satisfaction.
[1460] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1461] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1462] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1463] [Fourth embodiment]
[1464] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1465] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1466] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1467] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1468] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1469] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1470] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1471] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1472] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1473] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1474] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1475] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1476] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1477] As an embodiment of the present invention, a system including the following processes is provided.
[1478] System Overview
[1479] The user records the work content with a camera, and the video data is received by the terminal and sent to the server. The server analyzes the received video data and extracts important actions and objects. Then, based on the extracted data, a procedure manual is generated in natural language and output in a specified format. A specific embodiment is shown below.
[1480] Program Overview
[1481] 1. Recording Stage
[1482] User: Record the work process with a camera. Specifically, the user records the entire process from start to finish using a camera-equipped smartphone or dedicated recording device.
[1483] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, the video data is automatically saved to the internal storage.
[1484] 2. Uploading videos
[1485] Device: Sends saved video data to the server via the Internet. The device sends saved video data to the specified server upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to inform the user of progress.
[1486] 3. Video Reception and Preprocessing
[1487] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. The server checks the consistency of the video data and performs a simple check to see if there are any missing or corrupted data.
[1488] 4. Video Analysis
[1489] Server: Sends video data to the AI analysis engine. The server divides the stored video data into frames and performs pre-processing to extract the important parts of the frames.
[1490] Server: The AI analysis engine analyzes the video data. It performs frame-by-frame object recognition, motion detection, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[1491] 5. Text Generation
[1492] Server: Generates a procedure manual in natural language based on the extracted important actions and objects. The server analyzes the action list extracted by AI and generates procedures such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[1493] Server: Formats the generated text data into a manual. Formats the text according to the manual template and divides it into sections.
[1494] 6. Manual Output
[1495] Server: Generates the finished manual in PDF or other specified formats. Converts the entire manual to PDF and adds tables, diagrams, screenshots, etc. (if needed).
[1496] Server: Sends the generated manual to the user's device. The completed manual is sent to the user as an email attachment or a dedicated download link is provided.
[1497] Specific examples
[1498] As a specific example, a situation in which an operating manual for a machine in a factory is to be created will be described.
[1499] 1. Recording Stage
[1500] User: A factory worker uses his smartphone to film how to operate a new machine.
[1501] Device: Your smartphone will save this footage to its internal storage.
[1502] 2. Uploading videos
[1503] Device: The smartphone uses Wi-Fi to upload video data to a designated server.
[1504] 3. Video Reception and Preprocessing
[1505] Server: The server receives the video data and performs preprocessing for AI analysis.
[1506] 4. Video Analysis
[1507] Server: The server sends the video data to an AI analysis engine, which recognizes and lists the actions in each frame.
[1508] 5. Text Generation
[1509] Server: Based on the extracted actions, generate a procedure manual in natural language and format it into a manual.
[1510] 6. Manual Output
[1511] Server: Generates the completed manual in PDF format and sends it to the worker via email. This allows users to easily create work manuals and quickly update them.
[1512] This significantly reduces time and effort, and enables the latest manuals to be provided promptly.
[1513] The processing flow will be explained below.
[1514] Step 1:
[1515] Users record their work with a camera. Specifically, they acquire a smartphone with a camera or a dedicated recording device and record the entire process from start to finish.
[1516] Step 2:
[1517] The device will temporarily store the recorded video data. After the recording device or smartphone has finished recording, the video data will automatically be saved to its internal storage.
[1518] Step 3:
[1519] The device sends the stored video data to the server via the Internet. The device sends the stored video data to the specified server's upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[1520] Step 4:
[1521] The server saves the received video data in the server's storage. The server organizes and saves the received data in a dedicated directory. It checks the consistency of the video data and performs a simple check to see if there are any missing or damaged data.
[1522] Step 5:
[1523] The server sends the video data to an AI analysis engine, which divides the stored video data into frames and performs preprocessing to extract the important parts of the frames.
[1524] Step 6:
[1525] The server's AI analysis engine analyzes the video data, performing frame-by-frame object recognition, motion detection, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever, etc.) and steps, and creates a sequenced list of them.
[1526] Step 7:
[1527] The server generates a procedure manual in natural language based on the extracted important actions and objects. The server analyzes the action list extracted by AI and generates procedures such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[1528] Step 8:
[1529] The server formats the generated text data into a manual, formatting the text according to the manual template and dividing it into sections.
[1530] Step 9:
[1531] The server generates the finished manual in PDF or other specified format, converting the entire manual to PDF and adding tables, diagrams, screenshots, etc. (if needed).
[1532] Step 10:
[1533] The server sends the generated manual to the user's device, either by attaching the completed manual to an email or by providing a dedicated download link.
[1534] Example 1
[1535] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1536] There has been no system that can easily record video data of work content and automatically generate procedure manuals based on that data. In particular, there is a need for an efficient and accurate process that can analyze the video data, extract important actions and objects, and generate procedure manuals. This has led to the need for a system that can reduce work hours and labor, standardize work procedures, and enable rapid update support.
[1537] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1538] In this invention, the server includes: [means for acquiring video data;] [means for extracting important actions and objects using an AI analysis unit that analyzes the acquired video data; and] [means for generating a procedure manual in natural language based on the extracted data. This makes it possible to automatically generate a procedure manual based on the video data of the work content and output it in a specified format.
[1539] "Video data" refers to digital information in the form of moving images captured using a camera or other imaging device.
[1540] A "server" is a computing device used to receive, analyze, store, and transmit data over a network.
[1541] "AI Analysis Unit" means a collection of software and hardware that implements artificial intelligence technology used to analyze video data and identify and extract significant actions and objects.
[1542] "Critical actions" refer to actions or operations that require special attention when carrying out business procedures.
[1543] "Objects" refer to specific items or devices that exist within the video data and are required to explain business procedures.
[1544] A "procedure manual" is a document that explains business procedures in a specific and orderly manner as a series of steps.
[1545] "Means for acquiring" refers to methods and devices for collecting video data using cameras and recording devices.
[1546] "Means for generating in natural language" refers to the method and process for generating a procedure manual in a natural language that is easy for humans to understand, based on the extracted data.
[1547] A "prescribed format" is a rule that ensures that procedures are organized in a certain format and displayed in a way that is visually easy to understand.
[1548] As an embodiment of this invention, we provide a system that automatically captures and analyzes work procedures at companies or work sites and generates procedure manuals. Specifically, this system allows users to record work content with a camera, and the video data is received by a terminal and sent to a server. The details of the system are described below.
[1549] Recording Stage
[1550] User:
[1551] Work is recorded with a camera. Specifically, the user uses a camera-equipped smartphone or a dedicated recording device to record work from start to finish. For example, a factory worker uses a smartphone to record how to operate a new machine.
[1552] Device:
[1553] Recorded video data is saved to the internal storage. After recording is complete, the smartphone automatically saves the video data temporarily.
[1554] Video upload
[1555] Device:
[1556] The saved video data is sent to a server via the Internet. The device uploads the video data to the specified server via Wi-Fi or a mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[1557] Video reception and preprocessing
[1558] server:
[1559] The server stores the received video data in a dedicated directory, checks for missing or corrupted data, and checks the data for consistency. Organizes the video data into batches for analysis.
[1560] Video Analysis
[1561] server:
[1562] The server divides the stored video data into frames and sends them to the AI analysis unit. As preprocessing, the video resolution is adjusted appropriately and unnecessary frames are removed. The AI analysis unit (e.g., TensorFlow or OpenCV) is used to perform object recognition and motion detection in each frame, and convert text and voice as needed. For example, scenes in which the user presses a button or pulls a lever are extracted and listed.
[1563] Text Generation
[1564] server:
[1565] Generates instructions in natural language based on the extracted important actions and objects. Using an AI model (e.g., GPT-3), it creates steps such as "1. Turn on the power" and "2. Select the settings menu" from the listed actions. Provides appropriate explanations and cautions for each step. For example, it adds instructions such as "Check the power cord and be careful not to press the buttons too hard."
[1566] server:
[1567] The generated text data is formatted into a manual, with each step divided into sections and visually organized using bulleted and numbered lists.
[1568] Manual Output
[1569] server:
[1570] Generate the finished manual in PDF or other format of your choice. Use tools such as Adobe Acrobat or LaTeX to add tables, charts, and screenshots to create a visually appealing manual.
[1571] server:
[1572] The generated manual is sent to the user's device. The completed manual is sent as an email attachment to the user, or a dedicated download link is provided.
[1573] Specific examples
[1574] As a specific example, we will explain a situation in which an operating manual for a new machine is being created in a factory.
[1575] User:
[1576] Factory workers use their smartphones to record themselves operating a new machine.
[1577] Device:
[1578] The smartphone saves this footage to its internal storage.
[1579] Device:
[1580] The smartphone uses Wi-Fi to upload the video data to a designated server.
[1581] server:
[1582] The server receives the video data and performs preprocessing for AI analysis.
[1583] server:
[1584] The server sends the video data to an AI analysis unit, which recognizes and lists the movements in each frame.
[1585] server:
[1586] Based on the extracted actions, a procedure manual is generated in natural language and formatted into a manual.
[1587] server:
[1588] The completed manual is generated in PDF format and sent to the worker via email, allowing users to easily create work manuals and quickly update them.
[1589] Examples of prompt statements
[1590] "Please film how to operate the new machine and create a manual."
[1591] "Please tell me the procedure for uploading video data to the server."
[1592] "Please explain in detail how to generate the procedure manual."
[1593] This will provide a system that standardizes business procedures, improves efficiency, and enables rapid update support.
[1594] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1595] Step 1:
[1596] User: Records work with a camera. The user uses a smartphone with a camera or other recording device to record the entire process from start to finish. For example, a factory worker uses a smartphone to record how to operate a new machine. The user takes care to clearly capture specific operating steps (e.g., how to press a button, pull a lever). Input: The device used is a smartphone with a camera. Output: Recorded video data.
[1597] Step 2:
[1598] Device: After recording is complete, the device saves the video data to its internal storage. At the same time, the device sends the video data to the server via Wi-Fi or the mobile network. During uploading, the device displays a progress bar to inform the user of the transfer status. Input: Recorded video data. Output: Video data sent to the server.
[1599] Step 3:
[1600] Server: The server stores the received video data in a dedicated directory in an organized manner. It first checks the data for consistency and performs a simple check to see if there are any missing or corrupted data. If the data is OK, it is broken down into frames for analysis and the resolution is adjusted if necessary. Input: Video data sent to the server. Output: Video data that has been split into frames and pre-processed.
[1601] Step 4:
[1602] Server: Sends preprocessed video data to an AI analysis unit (e.g., TensorFlow or OpenCV), which performs object recognition and action detection for each frame, and converts text and voice as needed. For example, it extracts and lists scenes in which a button is pressed or a lever is pulled. Input: Preprocessed video data for each frame. Output: A list of extracted important actions and objects.
[1603] Step 5:
[1604] Server: Based on the extracted important actions and objects, a generative AI model (e.g., GPT-3) is used to generate instructions in natural language. Detailed explanations and precautions are added for each action (e.g., pressing a button), and steps such as "1. Turn on the power" and "2. Select the settings menu" are formed. Input: A list of extracted actions and objects. Output: Instructions generated in natural language.
[1605] Step 6:
[1606] Server: Formats the generated procedure manual into a manual format and outputs it in PDF or other specified format. Tables, diagrams, and in some cases screenshots are added to the procedure manual to make it visually easy to understand. Input: Procedure manual generated in natural language. Output: Procedure manual formatted in PDF or other specified format.
[1607] Step 7:
[1608] Server: Provides the completed procedure to the user. The server sends the procedure as an email attachment or provides a dedicated download link. It may also store it in a database. Input: Procedure formatted in PDF or other specified format. Output: Procedure provided to the user.
[1609] Through this series of processes, a system is realized that can automatically generate procedure manuals from video data of work content and provide them to users quickly and efficiently, making it possible to standardize work procedures and respond to updates quickly.
[1610] (Application example 1)
[1611] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1612] In modern factories and manufacturing sites, specialized robot maintenance work is complex, and creating procedure manuals requires a great deal of time and effort. Therefore, there is a need for an effective automatic procedure manual generation system to improve the efficiency and standardization of maintenance work. However, existing manual procedure manual creation methods have issues such as difficulty in recording and analyzing work, and difficulty in quickly reflecting the latest information.
[1613] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1614] In this invention, the server includes: [means for acquiring video data using a device with a camera; [means for extracting important actions and objects using an AI analysis engine that analyzes the acquired video data; and [means for generating a procedure manual in natural language based on the extracted data.] This makes it possible [to achieve work efficiency and standardization, and to quickly reflect the latest information in the procedure manual].
[1615] "Video data" is a series of image information captured by a device with a camera, and is digital data that includes details of actions and objects.
[1616] "AI analysis engine" is a general term for software and hardware that uses artificial intelligence technology to analyze video data and extract important actions and objects.
[1617] A "procedure" is a document that provides step-by-step instructions for performing a specific task or operation.
[1618] "Natural language" is a language used by humans on a daily basis and a set of textual forms generated by computer programs.
[1619] A "generative AI model" is an artificial intelligence model that learns large amounts of data in advance to generate instruction manuals and other text data.
[1620] A "prompt" is a series of phrases or documents that are input to a generative AI model to determine the content and format of the generated text.
[1621] A "server" is a computer that receives requests from client devices over a network and provides data processing and data storage functions.
[1622] "Preprocessing" refers to the process of removing noise from video data and converting its format before analysis by the AI analysis engine.
[1623] A "data processing device" is a computer system for analyzing video data and extracting important information.
[1624] A "template" is a predefined format or style guide for formatting procedures and reports.
[1625] "Digital document format" means a document format that can be stored and displayed electronically, including PDF.
[1626] MODE FOR CARRYING OUT THE INVENTION
[1627] System Overview
[1628] As an embodiment of the present invention, we provide a system that performs the following process. In this system, a user records work content using a camera-equipped smartphone and sends the video data to a server. The server analyzes the received video data and extracts important actions and objects. Then, based on the extracted data, it generates a procedure manual in natural language and outputs it in a specified format.
[1629] Explanation of program processing
[1630] Hardware and software used
[1631] Smartphone: A smart device used as a recording device.
[1632] Server: A computer used for data analysis that receives requests from client devices over the network and processes the data.
[1633] OpenCV: A library for recording and processing video.
[1634] Requests library: Used to send video data to the server.
[1635] AI analysis engine: An engine for analyzing video data and extracting important actions and objects.
[1636] PDF generation library: A library for converting instructions into PDF format (e.g., ReportLab).
[1637] Processing flow
[1638] 1. Recording: The user records the robot's maintenance work using a smartphone. The recorded data is saved in the smartphone's internal storage.
[1639] 2. Video Upload: Once the recording is complete, the device (smartphone) sends the saved video data to the server via Wi-Fi or mobile network. This data transmission is performed using the Requests library.
[1640] 3. Data processing: The server receives the transmitted video data and performs preprocessing, which includes noise reduction and format conversion. The AI analysis engine then divides the video data into frames and extracts important actions and objects.
[1641] 4. Procedure generation: A procedure manual is generated in natural language based on the data extracted by the AI analysis engine. A generative AI model is used here. A prompt sentence for generating the procedure manual is also created. By inputting this prompt sentence into the generative AI model, automatically generated text can be obtained.
[1642] 5. Output of procedure manual: The generated procedure manual is output in PDF format using a PDF generation library, and the output PDF is provided to the user.
[1643] Specific examples
[1644] For example, let's say a maintenance task of adding lubricant to the joints of a robot is recorded in a factory. This recorded data is sent to a server, which then uses an analysis engine to analyze the video and extract important actions (such as "remove the cover," "prepare the lubricant," and "apply lubricant to the specified location"). Based on the extracted actions, the AI model automatically generates a procedure manual and provides it to the engineer in PDF format.
[1645] Prompt Sentence Examples
[1646] Video data URL: {video_url}
[1647] Analysis output:
[1648] Object Recognition
[1649] Action Detection
[1650] Procedure generation
[1651] The above is an embodiment of the present invention. This system makes it possible to improve the efficiency and standardize maintenance work, and to quickly update the latest information in the procedure manual.
[1652] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1653] Program processing flow
[1654] Step 1:
[1655] Recording
[1656] The user records the robot's maintenance work using a smartphone. At this time, a camera-equipped device is used to capture the entire process from start to finish. The recorded data (video data) is saved in the smartphone's internal storage.
[1657] Input: Maintenance work video
[1658] Output: Video data stored in the internal storage
[1659] Step 2:
[1660] Video upload
[1661] Once recording is complete, the device (smartphone) sends the saved video data to the server via the Internet. The device uses the Requests library to send the video data to the specified server's upload endpoint. At this time, the file transfer status is displayed to inform the user of the progress.
[1662] Input: Video data stored in the internal storage
[1663] Output: Video data sent to the server
[1664] Step 3:
[1665] Data reception and preprocessing
[1666] The server receives the transmitted video data and stores it in storage. Upon receiving the data, it performs pre-processing such as noise removal and format conversion. This process converts the data into a format that can be smoothly processed by the analysis engine.
[1667] Input: Video data sent to the server
[1668] Output: Pre-processed video data
[1669] Step 4:
[1670] Video Analysis
[1671] The preprocessed video data is passed to the server's AI analysis engine, which divides the video data into frames and extracts important actions and objects from each frame. Here, it performs tasks such as object recognition and action detection for each frame, and creates a list of important actions.
[1672] Input: Preprocessed video data
[1673] Output: Data listing important actions and objects
[1674] Step 5:
[1675] Procedure generation
[1676] Based on the data listing important actions and objects, the server uses a generative AI model to generate instructions in natural language. A prompt is entered into the generative AI model, which then automatically generates text. An example of a prompt might be something like "Video data URL: {video_url}\nAnalysis output:\n- Object recognition\n- Action detection\n- Instruction manual generation."
[1677] Input: Data listing important actions and objects, prompts
[1678] Output: Natural language instructions
[1679] Step 6:
[1680] Output of procedure manual
[1681] The server converts the generated instructions into PDF format using a PDF generation library, and the generated PDF is sent to the user's device or a download link is provided.
[1682] Input: Natural language instructions
[1683] Output: Instructions in PDF format
[1684] By following the above steps, a procedure manual can be automatically generated from a series of maintenance tasks and provided efficiently.
[1685] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1686] As an embodiment of the present invention, we provide a system that includes the following processing. In addition to acquiring and analyzing video data, this system recognizes the user's emotions and reflects them in the generation of procedure manuals, thereby providing more personalized manuals.
[1687] System Overview
[1688] The user records their work with a camera, and the video data is received by the device and sent to the server. The server analyzes the received video data and extracts important actions and objects. It then generates a procedure manual in natural language based on the extracted data and outputs it in a specified format. The system also incorporates an emotion engine that recognizes the user's emotions and reflects the results in the generation of the procedure manual.
[1689] Program Overview
[1690] 1. Recording Stage
[1691] User: Record work content with a camera. Specifically, the user acquires a smartphone with a camera or a dedicated recording device and records the process of work from start to finish.
[1692] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, the video data is automatically saved to the internal storage.
[1693] 2. Uploading videos
[1694] Device: Sends saved video data to the server via the Internet. The device sends saved video data to the specified server upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to inform the user of progress.
[1695] 3. Video Reception and Preprocessing
[1696] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. The server checks the consistency of the video data and performs a simple check to see if there are any missing or corrupted data.
[1697] 4. Video Analysis
[1698] Server: Sends video data to the AI analysis engine. The server divides the stored video data into frames and performs pre-processing to extract the important parts of the frames.
[1699] Server: The AI analytics engine analyzes the video data. It performs frame-by-frame object recognition, motion detection, emotion analysis, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly fashion.
[1700] 5. Emotion analysis
[1701] Server: Analyzes the user's facial expressions and voice in the video data and uses an emotion engine to recognize emotions. Evaluates the user's stress level and satisfaction level and reflects that data in the analysis results.
[1702] Server: Adjust the content and tone of instructions based on the perceived emotion, for example adding more detailed explanations or additional information where the user is feeling stressed.
[1703] 6. Text Generation
[1704] Server: Generates instructions in natural language based on the extracted important actions and user emotion data. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[1705] Server: Formats the generated text data into a manual. Formats the text according to the manual template and divides it into sections.
[1706] 7. Manual Output
[1707] Server: Generates the finished manual in PDF or other specified formats. Converts the entire manual to PDF and adds tables, diagrams, screenshots, etc. (if needed).
[1708] Server: Sends the generated manual to the user's device. The completed manual is sent to the user as an email attachment or a dedicated download link is provided.
[1709] 8. Feedback and Learning
[1710] Server: Records the user's emotion recognition results and reflects them in the next procedure manual generation. Based on the recorded emotion data, the recommended procedure manual content is personalized.
[1711] Users: Provide feedback after using the manual, so the system continually learns and improves the quality of the procedure manual.
[1712] Specific examples
[1713] As a specific example, we will explain a situation in which an operating manual for a new machine is being created in a factory.
[1714] 1. Recording Stage
[1715] User: A factory worker uses his smartphone to film how to operate a new machine.
[1716] Device: Your smartphone will save this footage to its internal storage.
[1717] 2. Uploading videos
[1718] Device: The smartphone uses Wi-Fi to upload the video data to a designated server.
[1719] 3. Video Reception and Preprocessing
[1720] Server: The server receives the video data and performs preprocessing for AI analysis.
[1721] 4. Video Analysis
[1722] Server: The server sends the video data to an AI analysis engine, which recognizes and lists the actions in each frame.
[1723] 5. Emotion analysis
[1724] Server: Recognizes emotions from the user's facial expressions and voice and adjusts the contents of the instruction manual.
[1725] 6. Text Generation
[1726] Server: Generates instructions in natural language based on the extracted action and emotion data and formats them into a manual.
[1727] 7. Manual Output
[1728] Server: Generates the completed manual in PDF format and sends it to the worker's email address.
[1729] 8. Feedback and Learning
[1730] Server: The system continues to learn based on user feedback and improves the quality of the instructions.
[1731] This will significantly reduce time and effort, and enable the rapid provision of up-to-date manuals that take user feelings into consideration.
[1732] The processing flow will be explained below.
[1733] Step 1:
[1734] Users record their work with a camera. Specifically, they record the entire process from start to finish using a camera-equipped smartphone or a dedicated recording device.
[1735] Step 2:
[1736] The device will temporarily store the recorded video data. After the recording device or smartphone has finished recording, the video data will automatically be saved to its internal storage.
[1737] Step 3:
[1738] The device sends the stored video data to the server via the Internet. The device sends the stored video data to the specified server's upload endpoint via Wi-Fi or mobile network. During upload, the file transfer status is displayed to show the progress to the user.
[1739] Step 4:
[1740] The server saves the received video data in the server's storage. The server organizes and saves the received data in a dedicated directory. It checks the consistency of the video data and performs a simple check to see if there are any missing or damaged data.
[1741] Step 5:
[1742] The server sends the video data to an AI analysis engine, which divides the stored video data into frames and performs preprocessing to extract the important parts of the frames.
[1743] Step 6:
[1744] The server's AI analysis engine analyzes the video data, performing frame-by-frame object recognition, motion detection, emotion analysis, and text-to-speech conversion (if necessary). It identifies important actions (e.g., pressing a button, pulling a lever) and steps, and lists them in an orderly manner.
[1745] Step 7:
[1746] The server uses an emotion engine to analyze the user's facial expressions and voice in the video data and recognize emotions, such as whether the user is confused or happy.
[1747] Step 8:
[1748] The server adjusts the content and tone of the instructions based on the perceived emotion, for example adding detailed explanations or additional information if it detects that the user is stressed.
[1749] Step 9:
[1750] The server generates a procedure manual in natural language based on the extracted important actions and user emotion data. The server analyzes the action list extracted by AI and generates steps such as "1. Turn on the power," "2. Select the settings menu," and "3. Press the start button." Appropriate explanations and precautions are added to each step.
[1751] Step 10:
[1752] The server formats the generated text data into a manual, formatting the text according to the manual template and dividing it into sections.
[1753] Step 11:
[1754] The server generates the finished manual in PDF or other specified format, converting the entire manual to PDF and adding tables, diagrams, screenshots, etc. (if needed).
[1755] Step 12:
[1756] The server sends the generated manual to the user's device, either by attaching the completed manual to an email or by providing a dedicated download link.
[1757] Step 13:
[1758] The server records the user's emotion recognition results and reflects them in the next procedure manual generation. The content of the procedure manual is personalized using the user's past emotion data.
[1759] Step 14:
[1760] The server continues to learn from feedback after the manual is used, and by combining user feedback with emotional data, the quality of the manual is improved.
[1761] Example 2
[1762] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1763] Conventional procedure manual generation systems often create procedures without considering the user's emotions, resulting in a lack of means to identify and address parts that users find emotionally difficult. Furthermore, feedback is not effectively incorporated into the procedure manual generation process, resulting in inconsistent quality. This can result in procedures that are not suited to actual work.
[1764] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1765] In this invention, the server includes: [means for providing an emotion engine that analyzes emotions from the user's facial expressions and voice; [means for generating a procedure manual in natural language based on the extracted data and emotion data; and [means for collecting feedback from the user and allowing the system to learn.] This makes it possible [to provide information that takes the user's emotions into consideration, thereby improving the quality and usability of the procedure manual].
[1766] "Video data" refers to video files in which users record their work.
[1767] "AI engine" refers to an artificial intelligence algorithm that analyzes video data and extracts important actions and objects.
[1768] "Important actions" refer to the work steps and operations required to generate a procedure manual.
[1769] An "object" refers to an object that is recognized in video data.
[1770] An "emotion engine" refers to a system that recognizes and analyzes emotions from the user's facial expressions and voice in video data.
[1771] A "procedure manual" refers to a document that describes the steps a user takes to perform a specific task.
[1772] "Predetermined format" refers to the predetermined manner in which the generated procedure manual is displayed or printed.
[1773] "Terminal" refers to a device for acquiring and transmitting video data.
[1774] "Feedback" refers to evaluations and opinions about the system provided by users.
[1775] "Preprocessing" refers to the process of organizing and checking consistency of video data before analyzing it.
[1776] "PDF format" refers to the Portable Document Format, an electronic document format.
[1777] "Download link" refers to the URL for obtaining a file via the Internet.
[1778] MODE FOR CARRYING OUT THE INVENTION
[1779] The following hardware and software are primarily used in the embodiment of this invention. A user records work content using a smartphone or dedicated recording device, and the terminal temporarily stores this video data. The terminal then transmits the video data to a server, which analyzes the data using an AI engine and an emotion engine. A procedure manual is then generated in natural language based on the extracted information and output in a specified format. The procedure manual is output in PDF format and provided to the user. The system also collects user feedback and continues to learn.
[1780] Hardware and software used
[1781] Hardware
[1782] A smartphone or dedicated recording device
[1783] server
[1784] software
[1785] AI analysis engine (analysis of video data)
[1786] Emotion engine (user emotion analysis)
[1787] PDF generation software (generating instruction manuals)
[1788] Internet connection software (data transmission and reception)
[1789] Data processing and calculation flow
[1790] 1. User: Records work using a smartphone or dedicated recording device. When a user wants to record how to operate a new machine, they can record all the operating procedures at once.
[1791] 2. Device: Temporarily stores the recorded video data. When the recording is finished, the device automatically saves the data to the internal storage and notifies the user of the progress.
[1792] 3. Device: The video data is sent to the server via the Internet. The data is uploaded to the specified server via Wi-Fi or mobile network. The upload progress is displayed on the screen.
[1793] 4. Server: Stores the received video data in storage and checks the data consistency. It performs a simple check to see if there are any missing or corrupted data.
[1794] 5. Server: The video data is divided into frames, and preprocessing such as object recognition and motion detection is performed using an AI analysis engine.
[1795] 6. Server: The AI analysis engine analyzes the video data, identifies and lists important actions and procedures, such as "pressing a button" or "pulling a lever."
[1796] 7. Server: The emotion engine recognizes emotions from the user's facial expressions and voice, evaluates the data, determines the user's stress level and satisfaction, and reflects this in the operating instructions.
[1797] 8. Server: Generates a natural language instruction manual based on the extracted important actions and emotion data. The generated instruction manual includes specific steps such as "1. Turn on the power" and "2. Select the settings menu."
[1798] 9. Server: The generated text data is organized into a specified manual format, categorizing it into sections and adding tables and charts as needed.
[1799] 10. Server: The completed procedure manual is output in PDF format and sent to the user's device. The procedure manual is attached to an email or a download link is provided.
[1800] 11. Server: Collects user feedback and the system continues to learn. Based on the feedback, it will be reflected in the next generation of the procedure manual.
[1801] Specific examples
[1802] When creating an operating manual for a new machine in a factory, a user uses a smartphone to film how to operate the new machine in the factory. The video data is saved by the device and uploaded to a server via the Internet. The server receives the video data, analyzes the data using an AI analysis engine and an emotion engine, and generates a manual in natural language. This manual is created in PDF format and sent to the user. Feedback provided by the user after using the manual is used to help generate manuals in the future.
[1803] Prompt Sentence Examples
[1804] "I recorded the machine's operating procedures on my smartphone. Please generate an operating manual using this video data. Please list important actions and customize the manual with consideration for the user's emotions."
[1805] By inputting this prompt into the generative AI model, the system can efficiently generate instructions.
[1806] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1807] Step 1:
[1808] Recording Stage
[1809] User: Record work with a camera. Users record the process from start to finish using a smartphone or dedicated recording device. For example, record the process of learning how to operate a new machine. During this process, it is recommended to adjust the camera angle and lighting.
[1810] Input: User operation during work
[1811] Output: Recorded video data
[1812] Step 2:
[1813] Data storage
[1814] Device: Temporarily stores recorded video data. After the recording device or smartphone finishes recording, it automatically saves the video data to its internal storage. It notifies the user that the save is complete. The saved data is usually in MP4 or MOV format.
[1815] Input: Recorded video data
[1816] Output: Video data stored in the internal storage
[1817] Step 3:
[1818] Video upload
[1819] Device: Video data is sent to the server via the internet. Data is uploaded to the specified server via Wi-Fi or mobile network. The transfer progress is displayed on the screen and the user is notified when the transfer is complete. Data is encrypted during transfer as a security measure.
[1820] Input: Video data stored in the internal storage
[1821] Output: Video data sent to the server
[1822] Step 4:
[1823] Video reception and preprocessing
[1824] Server: The received video data is stored in the server's storage. The server organizes and stores the received data in a dedicated directory. After saving, the data is checked for consistency and a simple check is performed to see if there are any missing or corrupted data. This preprocessing also includes checking the frame rate and renaming the data.
[1825] Input: Video data sent to the server
[1826] Output: Pre-processed video data
[1827] Step 5:
[1828] Video Analysis
[1829] Server: Sends video data to the AI analysis engine. The server divides the data into frames and performs preprocessing to extract important parts. It identifies important objects and actions in each frame.
[1830] Input: Preprocessed video data
[1831] Output: Object and movement data for each frame analyzed by the AI analysis engine
[1832] Server: The AI analysis engine analyzes the video data, identifies and lists important actions and procedures, for example, recognizing operations such as "pressing a button" or "pulling a lever."
[1833] Input: Object and movement data per frame
[1834] Output: A list of extracted important actions and steps
[1835] Step 6:
[1836] Emotion analysis
[1837] Server: Analyzes the user's facial expressions and voice in the video data and uses an emotion engine to recognize emotions. The server evaluates the user's stress level and satisfaction level and reflects this data in the analysis results. Analyzing subtle changes in facial expressions and tone of voice in particular enables more accurate emotion recognition.
[1838] Input: User's facial expressions and voice in video data
[1839] Output: User emotion data
[1840] Server: Adjust the content and tone of instructions based on the perceived emotion, for example adding more detailed explanations or additional information where the user is feeling stressed.
[1841] Input: User sentiment data, list of important actions and steps
[1842] Output: An outline of the procedure manual with the sentiment data reflected
[1843] Step 7:
[1844] Text Generation
[1845] Server: Generates instructions in natural language based on the extracted important actions and emotion data. AI analyzes the action list and generates specific steps such as "1. Turn on the power" and "2. Select the settings menu." Appropriate explanations and cautions are added to each step.
[1846] Input: Instructions outline with sentiment data reflected
[1847] Output: Natural language generated instructions
[1848] Server: Prepares the generated text data into a predetermined manual format. Formats the text according to a template, categorizes it into sections, and adds tables and charts as needed.
[1849] Input: Natural language generated instructions
[1850] Output: Formatted instructions
[1851] Step 8:
[1852] Manual Output
[1853] Server: Prints the completed manual in PDF format and sends it to the user's device. The instructions are attached to an email or a download link is provided. Adds tables and diagrams as needed (if necessary).
[1854] Input: Formatted instructions
[1855] Output: Instructions in PDF format provided to users
[1856] Step 9:
[1857] Feedback and Learning
[1858] Server: Collects feedback from users and the system continues to learn. Based on the feedback, it reflects it in the next generation of the procedure manual. Specifically, it analyzes evaluation points and comments and collects data to improve the procedure manual.
[1859] Input: User feedback
[1860] Output: Learning data that will be useful for the next procedure generation
[1861] (Application example 2)
[1862] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1863] Generating work instructions is important for improving the work efficiency of factory robots. However, conventional methods generate uniform instructions, making it difficult to provide personalized instructions that take into account the emotions and situations of individual users. Furthermore, analyzing work content without considering the user's emotions can lead to the risk of overlooking areas that cause stress to the user or areas for improvement.
[1864] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data, means for extracting important actions and objects using an AI engine that analyzes the acquired video data, means for using an emotion analysis engine that recognizes the user's emotions and reflects the results in generating a procedure manual, means for generating a procedure manual in natural language based on the extracted data and emotion data, and means for outputting the generated procedure manual in a predetermined format and providing it to the user. This makes it possible to provide a personalized procedure manual that takes into account the user's work situation and emotions.
[1865] "Video data" refers to a series of image frames captured by a camera, and is the digital data that is the subject of analysis.
[1866] "AI Engine" is a software platform for analyzing and extracting data using artificial intelligence technology.
[1867] "Critical actions" refer to actions or behaviors that are particularly essential in work procedures or processes.
[1868] "Object" refers to an object, such as a person or object, that is recognized in video data.
[1869] An "emotion analysis engine" is a software platform that recognizes emotions from a user's facial expressions and voice in video data and extracts that emotional state as data.
[1870] A "procedure manual" refers to a document that provides detailed instructions for a specific task or operation.
[1871] A "natural language" is a language that humans use on a daily basis, and is a written expression used to describe specific processes or instructions.
[1872] "Format" refers to a template or form that defines the structure and format of a document or data.
[1873] A "server" is a computer system that stores data, processes data, and provides services over a network.
[1874] "Analysis" is the process of breaking down data into detail to reveal its meaning and structure.
[1875] This invention is a system for automatically generating personalized operation manuals that incorporate the emotions of users by utilizing factory robots. The system of the present invention captures video of the user working, analyzes it using an AI engine, and generates operation manuals based on the extracted data and emotion analysis. Specifically, the system performs the following processes.
[1876] System Overview
[1877] 1. Recording Stage
[1878] User: Users working in the factory record their work using cameras attached to the robots, and through this process, video data is acquired in real time.
[1879] Hardware: Video data is temporarily stored in the robot's internal storage.
[1880] 2. Uploading videos
[1881] Server: The robot sends video data to the server via an internet connection, either wireless LAN or wired data communication.
[1882] 3. Video Reception and Preprocessing
[1883] Server: The received video data is stored in the server's storage. Next, as preprocessing, the data consistency is checked and processing such as noise removal and frame division is performed.
[1884] 4. Video Analysis
[1885] Server: An AI engine using OpenCV and TensorFlow analyzes the video data and extracts important actions and objects, using algorithms for object recognition and motion detection.
[1886] 5. Emotion analysis
[1887] Server: Uses an emotion analysis engine such as Microsoft Azure Cognitive Services to recognize emotions from the user's facial expressions and voice, allowing it to assess whether the user is feeling stressed.
[1888] 6. Text Generation
[1889] Server: Using a generative AI model such as GPT-4, a natural language instruction manual is generated based on the extracted key actions and emotion data. The instruction manual is then formatted in a predetermined format.
[1890] 7. Manual Output
[1891] Server: Generates the completed procedure manual in PDF format and provides it to the user by displaying it on the robot's display or by sending it to the user's email address.
[1892] 8. Feedback and Learning
[1893] Server: Collects feedback from users after using the procedure and uses that data to improve the analysis model for the next run, allowing the system to continuously learn and improve the quality of the procedure.
[1894] Specific examples
[1895] For example, when inspecting products in a factory, a robot can record the work, analyze the video, and compile a procedure manual for the next inspection method, taking the user's feelings into consideration. This can increase work efficiency while reducing user stress.
[1896] Prompt Sentence Examples
[1897] "Please record the product inspection process in the factory with a camera and generate the next inspection procedure based on the analysis results."
[1898] This system can provide flexible instructions that reflect the user's emotions and specific work content, which is expected to improve work efficiency and user satisfaction.
[1899] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1900] Step 1:
[1901] User: A user working in a factory records their work using a camera attached to the robot.
[1902] Input: The video data you are working with.
[1903] Output: Captured working video data.
[1904] Specific operation: The camera captures video data in real time and stores it in the robot's built-in storage.
[1905] Step 2:
[1906] Server: The robot sends the video data to the server via an internet connection.
[1907] Input: Video data of work inside the robot.
[1908] Output: Video data stored on the server.
[1909] Specific operation: The robot sends video data to the server's upload endpoint using Wi-Fi or a wired connection.
[1910] Step 3:
[1911] Server: Preprocesses the received video data and passes it to the AI analysis engine.
[1912] Input: Received video data.
[1913] Output: Pre-processed video data.
[1914] Specific operation: The server performs preprocessing such as checking data consistency, removing noise, and splitting frames.
[1915] Step 4:
[1916] Server: An AI engine using OpenCV and TensorFlow analyzes video data and extracts important actions and objects.
[1917] Input: Preprocessed video data.
[1918] Output: Extracted important action and object data.
[1919] Specific actions: The AI engine runs object recognition and action detection algorithms to list important actions and objects.
[1920] Step 5:
[1921] Server: Uses an emotion analysis engine to recognize emotions from the user's facial expressions and voice.
[1922] Input: Preprocessed video and audio data.
[1923] Output: User emotion data.
[1924] How it works: An emotion analysis engine such as Microsoft Azure Cognitive Services analyzes facial expressions and vocal tones to digitize the user's emotional state.
[1925] Step 6:
[1926] Server: Using a generative AI model, it generates instructions in natural language based on the extracted key actions and emotion data.
[1927] Input: Important action and emotion data.
[1928] Output: Instructions generated in natural language.
[1929] How it works: A generative AI model such as GPT-4 analyzes key action lists and emotional data to generate instructions in natural language.
[1930] Step 7:
[1931] Server: Outputs the generated procedure manual in a specified format and provides it to the user.
[1932] Input: Instructions generated in natural language.
[1933] Output: Instructions in PDF format.
[1934] Specific operation: The server formats the instruction manual in PDF format and displays it on the robot's display or sends it to the user by email.
[1935] Step 8:
[1936] Server: Collects feedback from users after using the procedure and uses that data to improve the next analysis model.
[1937] Input: User feedback data.
[1938] Output: An improved analytical model.
[1939] How it works: The server stores the feedback data in a database and retrains the generative AI model (GPT-4) to improve the quality of future instructions.
[1940] Through the above processing steps, the present invention automatically generates a personalized procedure manual that takes the user's feelings into consideration, thereby improving work efficiency and user satisfaction.
[1941] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1942] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1943] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1944] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1945] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1946] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1947] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1948] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1949] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1950] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1951] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1952] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1953] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1954] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1955] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1956] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1957] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1958] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1959] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1960] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1961] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1962] The following is further disclosed regarding the above embodiment.
[1963] (Claim 1)
[1964] [Means for acquiring video data;
[1965] [Means of extracting important actions and objects using an AI engine that analyzes the acquired video data;
[1966] [Means for generating a procedure manual in natural language based on the extracted data;
[1967] [Means for outputting the generated procedure manual in a predetermined format;
[1968] A system including:
[1969] (Claim 2)
[1970] [Means for transmitting video data to a server via the Internet;
[1971] [Means for preprocessing received video data and passing it to the AI analysis engine,
[1972] The system of claim 1 further comprising:
[1973] (Claim 3)
[1974] [Means for formatting data according to a predetermined template when generating a procedure manual;
[1975] [Means to output the generated procedure manual in PDF format and provide it to users,
[1976] The system of claim 1 further comprising:
[1977] "Example 1"
[1978] (Claim 1)
[1979] [Means for acquiring video data;
[1980] [Means for extracting important actions and objects using an AI analysis unit that analyzes the acquired video data;
[1981] [Means for generating a procedure manual in natural language based on the extracted data;
[1982] [Means for outputting the generated procedure manual in a predetermined format;
[1983] A system including:
[1984] (Claim 2)
[1985] [Means for transmitting video data to a server via a communication network;
[1986] [Means for preprocessing the received video data and passing it to the AI analysis unit;
[1987] The system of claim 1 further comprising:
[1988] (Claim 3)
[1989] [Means for formatting data according to a predetermined template when generating a procedure manual;
[1990] [Means for outputting the generated procedure manual in electronic document format and providing it to users;
[1991] The system of claim 1 further comprising:
[1992] "Application Example 1"
[1993] (Claim 1)
[1994] [Means for acquiring video data using a device with a camera;
[1995] [Means to extract important actions and objects using an AI analysis engine that analyzes the acquired video data;
[1996] [Means for generating a procedure manual in natural language based on the extracted data;
[1997] [Means for outputting the generated procedure manual in a predetermined format;
[1998] [Means for creating prompt sentences for generating instructions using a generative AI model;
[1999] A system including:
[2000] (Claim 2)
[2001] [Means for transmitting video data to a data processing device via the Internet;
[2002] [Means for preprocessing received video data and passing it to the AI analysis engine,
[2003] [Means for the data processing device to extract relevant maintenance actions from the video data;
[2004] The system of claim 1 further comprising:
[2005] (Claim 3)
[2006] [Means for formatting data according to a predetermined template when generating a procedure manual;
[2007] [Means for outputting the generated procedure manual in a digital document format and providing it to the user;
[2008] The system of claim 1 further comprising:
[2009] "Example 2: Combining Emotion Engines"
[2010] (Claim 1)
[2011] [Means for acquiring video data;
[2012] [Means of extracting important actions and objects using an AI engine that analyzes the acquired video data;
[2013] [Means for providing an emotion engine that analyzes emotions from the user's facial expressions and voice;
[2014] [Means for generating a procedure manual in natural language based on the extracted data and emotion data;
[2015] [Means for outputting the generated procedure manual in a predetermined format;
[2016] [Means for transmitting the generated procedure manual to the user's terminal;
[2017] [Means for gathering user feedback and for the system to learn from it;
[2018] A system including:
[2019] (Claim 2)
[2020] [Means for transmitting video data to a server via the Internet;
[2021] [Means for preprocessing received video data and passing it to the AI analysis engine,
[2022] [Means for listing action data extracted from video data in an orderly manner;
[2023] [The system according to claim 1, further comprising means for reflecting the emotion analysis results in the operating manual.]
[2024] (Claim 3)
[2025] [Means for formatting data according to a predetermined template when generating a procedure manual;
[2026] [Means to output the generated procedure manual in PDF format and provide it to users,
[2027] [Including means to enable output in specified formats other than PDF format,
[2028] The system of claim 1, further comprising means for providing a download link for the generated instruction manual to a user.
[2029] "Application example 2 when combining emotion engines"
[2030] (Claim 1)
[2031] [Means for acquiring video data;
[2032] [Means of extracting important actions and objects using an AI engine that analyzes the acquired video data;
[2033] [Means for using an emotion analysis engine to recognize the user's emotions and reflect the results in the generation of the procedure manual;
[2034] [Means for generating a procedure manual in natural language based on the extracted data and emotion data;
[2035] [Means for outputting the generated procedure manual in a predetermined format and providing it to the user;
[2036] A system including:
[2037] (Claim 2)
[2038] [Means for transmitting video data to a server via the Internet;
[2039] [The system according to claim 1, wherein the received video data is preprocessed and passed to an AI analysis engine.
[2040] (Claim 3)
[2041] [Means for formatting data according to a predetermined template when generating a procedure manual;
[2042] [The system according to claim 1, wherein the generated procedure manual is output in PDF format and provided to the user. [Explanation of symbols]
[2043] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring video data; A means of extracting important actions and objects using an AI engine that analyzes the acquired video data; A means for generating a procedure manual in natural language based on the extracted data; means for outputting the generated procedure manual in a predetermined format; A system including:
2. means for transmitting the video data to a server via the internet; A means to preprocess the received video data and pass it to the AI analysis engine, The system of claim 1 further comprising:
3. A means for formatting data according to a predetermined template when generating a procedure manual; A means for outputting the generated procedure manual in PDF format and providing it to the user; The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A