system

A system that records and converts user screen operations and voice explanations into text to automatically generate operation manuals addresses the inefficiencies and inaccuracies of manual creation, providing efficient and accurate documentation.

JP2026064666APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Creating operation manuals for computer screens is labor-intensive, time-consuming, and prone to errors due to the complexity of integrating screen captures and voice explanations, leading to inconsistent and inaccurate documentation.

Method used

A system that records user screen operations and voice explanations in real-time, converts audio to text, and automatically generates a manual by integrating the text and video data to create a consistent operation manual.

Benefits of technology

Enables efficient and accurate generation of operation manuals, reducing the effort required and ensuring high-quality documentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064666000001_ABST
    Figure 2026064666000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of recording video of the screen being operated by the user, A means of recording the user's voice through a microphone, A means of sending recorded video files and recorded audio files to a server, A means of converting audio files to text on a server, A means of generating an operation manual based on converted text and video files, A means of providing the generated operation manual to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When creating a manual for explaining the operation procedure of a computer screen, it is a problem that it takes a lot of labor and time. In particular, in order to integrate a screen capture and its corresponding voice explanation and finish it as a consistent procedure manual, a plurality of specialized operations are required. This process is very complicated, and mistakes are likely to occur when performed manually, and the quality of the final document is not uniform. Therefore, there is a need for a system that can automatically generate an operation manual efficiently and accurately.

Means for Solving the Problems

[0005] The present invention provides a system that includes means for recording video of a screen operated by a user, means for recording the user's voice through a microphone, means for transmitting the recorded video file and the recorded audio file to a server, means for converting the audio file to text on the server, means for generating an operation manual based on the converted text and video file, and means for providing the generated operation manual to the user. Specifically, by recording the user's operations and explanations in real time, and then automatically analyzing the video and audio to generate a manual, an efficient and accurate operation manual can be created quickly.

[0006] A "user" is a person who operates a system.

[0007] A "screen" is the area of ​​visual information displayed on a computer or other electronic device.

[0008] "Video" refers to a collection of moving images, which are recorded as screen captures or videos.

[0009] "Means of recording" refers to hardware and software used to capture and save on-screen operations as video.

[0010] A "microphone" is a device that captures sound and converts it into an electrical signal.

[0011] "Sound" refers to vibrational energy that includes human speech and other sounds.

[0012] "Means of recording" refers to hardware and software for recording audio acquired through a microphone as digital data.

[0013] A "server" is a central processing unit that processes data and provides services to other computers.

[0014] "The means for transmitting" refers to network components and protocols for moving data from one device to another.

[0015] "The audio file" refers to the recorded audio data stored in digital format.

[0016] "The means for converting to text" refers to speech recognition technology and software for analyzing audio data and converting it into character data.

[0017] "The operation manual" refers to a document or guide created to explain specific operation procedures.

[0018] "The means for generating" refers to software and algorithms for automatically creating new documents or files based on the collected data.

[0019] "The means for providing" refers to an interface and distribution technology for sharing the generated operation manual in a form accessible to users.

Brief Description of the Drawings

[0020] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7]It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0022] First, the language used in the following description will be explained.

[0023] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0024] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0028] [First Embodiment]

[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0041] This invention is a system that instantly records the computer screen operations performed by a user and their explanations, and later automatically generates an operation manual. The embodiments thereof are described in detail below.

[0042] System Overview

[0043] 1. Initial Setup

[0044] The user starts the system and grants permission for microphone and screen capture.

[0045] The device checks the microphone and screen capture settings and completes the setup.

[0046] 2. Recording screen operations and audio.

[0047] Recording begins when the user clicks the "Start Recording" button.

[0048] The device records all actions performed on the screen as video files.

[0049] At the same time, the device records the user's voice description through the microphone.

[0050] As a concrete example, consider a scenario where a user clicks "New File" and says aloud, "I have created a new file." This action and the accompanying audio are recorded simultaneously.

[0051] 3. Sending data

[0052] Once all recording and recording are complete, the user clicks the "Stop Recording" button.

[0053] The terminal sends the generated video and audio files to the server. This transmission takes place over the internet.

[0054] 4. Data Processing

[0055] The server analyzes the received video and audio files.

[0056] The server uses speech recognition technology to convert audio files into text data. For example, the audio file "A new file has been created" is converted to text "A new file has been created".

[0057] 5. Generation of the operation manual

[0058] The server combines the text with the corresponding timestamp information and integrates it with a specific portion (screenshot) of the video file.

[0059] Specifically, text and screenshots are compiled into a consistent manual as a series of steps.

[0060] 6. Distribution of manuals

[0061] The completed operation manual will be generated in PDF or HTML format.

[0062] The server will provide this manual to users by displaying a download link via email or on the dashboard.

[0063] Users can download and view the manual.

[0064] Specific example

[0065] 1. A scene where the user clicks the "Create New Project" button and the process is explained verbally.

[0066] User: "I'll create a new project here."

[0067] The device records audio simultaneously with operation, along with a timestamp.

[0068] 2. Data transmission and processing after recording stops

[0069] The device sends the recorded data to the server.

[0070] The server uses speech recognition to extract the text "We will create a new project here" and matches it with the video of the operation.

[0071] 3. Example of manual generation

[0072] The server organizes the steps on the timeline and generates a page for the "1. Create a new project" section, combining corresponding screenshots and text.

[0073] In this way, by using the system of the present invention, an efficient and accurate operation manual is automatically generated based on the user's operations and explanations. The implementation of this invention significantly reduces the effort required to create manuals and enables the provision of highly accurate documentation.

[0074] The following describes the processing flow.

[0075] Step 1:

[0076] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[0077] Step 2:

[0078] The user clicks the "Start Recording" button. This causes the device to begin recording the entire current screen, saving all actions as a video file. Simultaneously, the device records the user's voice via the microphone. Specifically, the screen capture and audio recording processes are executed at the same time.

[0079] Step 3:

[0080] As the user performs screen operations, the system provides real-time audio explanations of those operations. For example, when the user clicks "New File" and selects it from the "File" menu, the system will provide the audio explanation, "Now you will create a new file." The device records these operations and audio, along with a timestamp.

[0081] Step 4:

[0082] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and saves the generated video and audio files to a temporary directory.

[0083] Step 5:

[0084] The terminal sends video and audio files stored in a temporary directory to the server. The files are sent using an HTTP POST request. Specifically, the video and audio files are uploaded to the server as form data.

[0085] Step 6:

[0086] The server receives the uploaded video and audio files. After receiving them, it passes the audio files to a speech recognition engine to begin the process of converting them to text. Here, speech recognition technology (e.g., a speech recognition API) is used to convert the audio data into text data.

[0087] Step 7:

[0088] The server analyzes the video file and generates frame-by-frame screenshots. Based on the video timeline, it extracts specific action events (e.g., button clicks or form inputs) and takes screenshots of the moments when these events occur.

[0089] Step 8:

[0090] The server integrates the converted text and screenshots to generate an operation manual. Based on the text and timestamp information, the screenshots are arranged as a sequence of steps to create a consistent operation manual. The generated operation manual is in PDF or HTML format.

[0091] Step 9:

[0092] The server generates a download link to provide the user with the completed operation manual. This link is sent to the system dashboard or to the user's email address.

[0093] Step 10:

[0094] Users can download and view the operation manual by clicking the provided download link. This operation manual is available for reference as needed and supports the use of the system.

[0095] (Example 1)

[0096] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0097] Traditional methods for creating user manuals require users to manually record each operation step and then transcribe the information, which is time-consuming and labor-intensive. Furthermore, the operation and its explanation may not always be synchronized, potentially leading to the creation of inaccurate manuals. Therefore, there is a need for a system that allows users to efficiently and accurately generate user manuals without having to spend time on incidental tasks.

[0098] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0099] In this invention, the server includes means for recording video of the screen operated by the user, means for recording the user's voice through a microphone, means for transmitting the recorded video file and the recorded audio file from the device to an information processing device, means for converting the audio file into text data on the information processing device, means for generating an operation manual based on the converted text data and video file, and means for providing the generated operation manual to the user. This makes it possible for the user to generate an operation manual simply and efficiently.

[0100] A "user" refers to anyone who intends to use this system to generate an operation manual.

[0101] "Screen images" refers to all visual information displayed on the screen of a device operated by the user.

[0102] A "microphone" is an audio input device used to record the user's voice.

[0103] A "recorded video file" is a digital file created by recording the video of the screen being operated by the user.

[0104] "Recorded audio files" refer to digital files generated by recording the user's voice.

[0105] A "device" refers to electronic devices such as computers and smartphones that are operated by the user.

[0106] An "information processing device" refers to a server or cloud computer used to process video and audio files transmitted from a device.

[0107] "Speech recognition technology" is a technology that converts audio files into text data; it analyzes audio data and turns it into text.

[0108] "Text data" refers to text information converted by speech recognition technology.

[0109] An "operation manual" refers to a document that summarizes the user's operating procedures and their explanations, and is generated based on video files and text data.

[0110] "Electronic document format" refers to digital document formats such as PDF and HTML.

[0111] "Web page format" refers to a web-based document that can be viewed on an internet browser.

[0112] This invention provides a system that instantly records user-performed computer screen operations and their descriptions, and automatically generates an operation manual later. This system allows users to efficiently and accurately generate operation manuals.

[0113] Hardware and software to be used

[0114] Terminal: A device used by the user to perform operations (e.g., personal computer, smartphone)

[0115] Microphone: A voice input device for recording the user's voice.

[0116] Screen capture tool: Software that records screen activity (e.g., OBS Studio)

[0117] Server: An information processing device that performs data processing and manual generation (e.g., cloud computing service).

[0118] Speech recognition technology: Technology that converts speech data into text data (e.g., Google® Speech-to-Text API)

[0119] Image processing software: Technology for automatically generating screenshots from video (e.g., OpenCV)

[0120] HTML template engine: A technology for generating operation manuals in HTML format (e.g., Jinja2).

[0121] PDF generation software: Technology for generating manuals in PDF format (e.g., wkhtmltopdf).

[0122] Process Overview

[0123] 1. The user starts the system and grants permission for microphone and screen capture.

[0124] 2. The device checks the screen capture and microphone settings and completes the setup.

[0125] 3. When the user clicks the "Start Recording" button, the device simultaneously records the video and audio of the screen. For example, the user clicks "New File" and says aloud, "A new file has been created."

[0126] 4. When the user clicks the "Stop Recording" button, the device sends the recorded video and audio files to the server.

[0127] 5. The server analyzes the received video and audio files and uses speech recognition technology to convert the audio files into text data. For example, the audio data "A new file has been created" is converted into the text data "A new file has been created".

[0128] 6. The server integrates text data and timestamp information with video files to generate a consistent operation manual. The operation manual is output in HTML or PDF format using an HTML template engine or PDF generation software.

[0129] 7. The server will provide users with a download link for the generated operation manual via email or on the dashboard. Users can download and view the operation manual from this link.

[0130] Specific example

[0131] A scene where the user clicks the "Create New Project" button and the process is explained via voice.

[0132] User: "I'll create a new project here."

[0133] The device records audio simultaneously with this operation, along with a timestamp.

[0134] Data transmission and processing after recording stops

[0135] The device sends the recorded data to the server.

[0136] The server uses speech recognition to extract the text "We will create a new project here" and matches it with the video of the operation.

[0137] Manual generation example

[0138] The server organizes the steps on the timeline and generates a page for the step "1. Create a new project," combining corresponding screenshots and text.

[0139] Examples of prompt statements

[0140] 1. "Please record the steps to create a new project using both screen operations and audio."

[0141] 2. "Please continue by explaining the current screen operation aloud."

[0142] 3. "Once recording is complete, please click the 'End' button to send the data to the server."

[0143] The above details describe the embodiments of the present invention. This system enables users to generate highly accurate operation manuals in a short amount of time.

[0144] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0145] Step 1:

[0146] The user starts the system and grants permission for microphone and screen capture.

[0147] Input: Command to start the system, permission request popup.

[0148] Output: Permission granted

[0149] Specific action: When the user starts the system, a pop-up appears asking for permission to use the microphone and screen capture. The user clicks "Allow".

[0150] Step 2:

[0151] The device checks the microphone and screen capture settings and completes the setup.

[0152] Input: User-granted permissions

[0153] Output: Microphone and screen capture ready to use status

[0154] Specific operation: The device automatically checks the settings for the microphone and screen capture tool (e.g., OBS Studio) and displays a "Ready" message.

[0155] Step 3:

[0156] The user clicks the "Start Recording" button.

[0157] Input: User click of the "Start Recording" button.

[0158] Output: Trigger for recording start

[0159] Specific action: The user clicks the "Start Recording" button in the application window.

[0160] Step 4:

[0161] The device records all actions performed on the screen.

[0162] Input: Trigger for recording start

[0163] Output: Video data being recorded

[0164] Specific action: The OBS Studio software begins recording the user's screen activity. For example, it records the user clicking the "Create New Project" button.

[0165] Step 5:

[0166] The device records the user's voice description via the microphone.

[0167] Input: Trigger for recording start, user's voice

[0168] Output: Audio data being recorded

[0169] Specific actions: As the user performs an action, the system provides a voice explanation of that action. For example, it might say, "Here, we will create a new project."

[0170] Step 6:

[0171] The user clicks the "Stop Recording" button.

[0172] Input: User click of the "Stop Recording" button.

[0173] Output: Trigger for recording termination

[0174] Specific action: The user clicks the "Stop Recording" button after completing the operation.

[0175] Step 7:

[0176] The device sends the recorded video file and the recorded audio file to the server.

[0177] Input: Trigger for recording termination, video file, audio file

[0178] Output: Status of successful file transfer to server

[0179] Specific operation: The device uploads video and audio files to the server via the internet connection. For example, it sends files using an API and displays a "transmission complete" notification.

[0180] Step 8:

[0181] The server analyzes the received video and audio files.

[0182] Input: Video file, audio file

[0183] Output: Analyzed audio data, video data

[0184] Specific operation: The server temporarily stores the received data and separates the video file from the audio file.

[0185] Step 9:

[0186] The server converts the audio file into text data.

[0187] Input: Audio file

[0188] Output: Character data (text file)

[0189] Specific operation: Use the Google Speech-to-Text API to convert audio data into text data. For example, the audio "I have created a new file" will be converted into text.

[0190] Step 10:

[0191] The server integrates text and timestamp information with the video file.

[0192] Input: Text data, timestamp, video file

[0193] Output: Operation procedure data, screenshots

[0194] Specific operation: Use OpenCV to obtain appropriate screenshots from video, and generate a consistent instruction manual based on text data and timestamp information.

[0195] Step 11:

[0196] The server generates an operation manual.

[0197] Input: Operation procedure data, screenshots

[0198] Output: Operation manual (PDF or HTML format)

[0199] Specific operation: Use the Jinja2 template engine to generate an operation manual in HTML format, and then convert it to PDF format using wkhtmltopdf.

[0200] Step 12:

[0201] The server provides users with an operation manual.

[0202] Input: Generated operation manual

[0203] Output: Download link, notification email

[0204] Specific operation: The server saves the generated manual and provides a download link to the user. The user can download the manual by displaying the link in an email or on the dashboard.

[0205] (Application Example 1)

[0206] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0207] Creating manuals for operating robots and machinery within a factory is time-consuming and labor-intensive, and accurately conveying advanced operations and complex procedures is difficult. Furthermore, training new workers is lengthy and prone to errors, highlighting the need for improved production efficiency and safety.

[0208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0209] In this invention, the server includes means for recording video of the screen operated by the user, means for recording the user's voice through a microphone, means for using a visual device to record video of the real environment, means for transmitting the recorded video file and recorded audio file to the server, means for converting the audio file to text on the server, means for generating an operation manual based on the converted text and video file, and means for providing the generated operation manual to the user. This enables accurate and rapid recording of complex procedures and operations within the factory, efficient automatic generation of operation manuals, reduction of worker training time, and prevention of errors.

[0210] "Means for recording video of the user's screen" refers to a function that records all operations performed on the screen of a computer or device being operated by the user in video format.

[0211] "Means of recording user voice via microphone" refers to a function that uses a microphone to record the voice of the user when they give explanations or instructions while operating the device.

[0212] "Means of using visual devices to record images of the real environment" refers to a function that allows users to record images of their surroundings or the object they are operating in real time using devices such as smart glasses or cameras.

[0213] "Means for sending recorded video files and recorded audio files to a server" refers to a function for uploading recorded video data and audio data to a server via the internet or other means.

[0214] "Method for converting audio files to text on the server" refers to a function that uses speech recognition technology on the server side to automatically convert transmitted audio data into text format.

[0215] "A means of generating operation manuals based on converted text and video files" refers to a function that automatically creates consistent operation procedures and manuals by combining audio-to-text data and recorded video data.

[0216] "Means of providing the generated operation manual to the user" refers to a function that outputs the completed operation manual in PDF or HTML format, allowing the user to download or view it.

[0217] This invention is a system that records robot operation and maintenance procedures in a factory in real time and automatically generates an operation manual based on that data. The embodiments of this system are described in detail below.

[0218] System Overview

[0219] 1. Initial Setup

[0220] The user starts the system and grants permission for microphone and screen capture. Additionally, the user wears a visual device (e.g., smart glasses or a camera) and confirms the settings.

[0221] The device checks the settings for microphone, screen capture, and visual devices, and then completes the setup.

[0222] 2. Recording screen operations and audio.

[0223] Recording begins when the user clicks the "Start Recording" button.

[0224] The device records the user's actions on the device screen as a video file, and simultaneously records video of the real environment through a visual device.

[0225] At the same time, the device records the user's voice description through the microphone.

[0226] 3. Sending data

[0227] Once all recording and recording are complete, the user clicks the "Stop Recording" button.

[0228] The terminal sends the generated video and audio files to the server. This transmission takes place over the internet.

[0229] 4. Data Processing

[0230] The server analyzes the received video and audio files and uses speech recognition technology to convert the audio files into text data. For example, the audio data "I have created a new project" is converted to text "I have created a new project".

[0231] 5. Generation of the operation manual

[0232] The server combines the text with the corresponding timestamp information and integrates it with a specific portion (screenshot) of the video file.

[0233] Specifically, text and screenshots are compiled into a consistent manual as a series of steps.

[0234] 6. Distribution of manuals

[0235] The completed operation manual will be generated in PDF or HTML format.

[0236] The server will provide this manual to users by displaying a download link via email or on the dashboard.

[0237] Users can download and view the manual.

[0238] Specific example

[0239] Consider a scenario where factory workers wear smart glasses to perform robot maintenance procedures. For example, they might record the procedure for replacing a hydraulic cylinder.

[0240] The user performs the task while explaining verbally, "Here, we will remove the hydraulic cylinder," and both the operation and the audio are recorded simultaneously.

[0241] After recording stops, the terminal sends the recorded data to the server, converts the audio data into text, and generates a manual.

[0242] Example of a prompt

[0243] "Use video and audio capture to record robot operation procedures within the factory. After recording is complete, convert the audio data into text and generate an operation manual along with the video frames. For example, if you perform a task while explaining, 'Here we remove the hydraulic cylinder,' ensure that this procedure and explanation are integrated into a single procedure in the manual."

[0244] In this way, by using the system of the present invention, complex procedures and operations within a factory can be accurately and quickly recorded, and efficient operation manuals can be automatically generated, thereby reducing worker training time and preventing operational errors.

[0245] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0246] Step 1:

[0247] Initial settings:

[0248] The user starts the system and grants permissions for the microphone, screen capture, and visual devices (3D glasses or camera). The terminal verifies that these devices are working correctly and completes the setup. The input is the permission and setup required for each device, and the output is the result of verifying that the devices are ready for use.

[0249] Step 2:

[0250] Start recording screen operations and audio:

[0251] The recording process begins when the user clicks the "Start Recording" button. The device records all operations performed on the computer screen as a video file using screen capture software (e.g., PyAutoGUI). Simultaneously, audio recording software (e.g., PyAudio) records the user's voice explanations via the microphone. The visual device records the user's viewpoint in real time. Input data consists of user operations, audio, and visual environment, and the output is a recorded file.

[0252] Step 3:

[0253] Sending data:

[0254] Once recording is complete, the user clicks the "Stop Recording" button. The device then sends the recorded video files (screen captures and video from the visual device) and audio files to the server. The transmission takes place over the internet. The input data consists of the recorded video and audio files, and the output is the data uploaded to the server.

[0255] Step 4:

[0256] Data processing:

[0257] The server analyzes the received video and audio files. Audio files are converted to text using speech recognition technology (e.g., Google Cloud Speech-to-Text API). The input is an audio file, and the output is the corresponding text data. Video files are also organized based on timestamps.

[0258] Step 5:

[0259] Generating the operation manual:

[0260] The server combines the converted text data with corresponding timestamp information and integrates it with specific portions (screenshots) of the video file. This generates a consistent operation manual. The input is text data and video files, and the output is an integrated operation manual file.

[0261] Step 6:

[0262] Distribution of manuals:

[0263] The server generates a completed operation manual in PDF or HTML format and makes it available for users to download. The download link is provided via email or on the dashboard. The input is the generated operation manual, and the output is the download link that users can access.

[0264] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0265] This invention provides a system that instantly records user computer screen operations and their explanations, and later automatically generates an operation manual. Furthermore, by combining it with an emotion engine, it enables the recognition and utilization of user emotions. The embodiments thereof will be described in detail below.

[0266] System Overview

[0267] 1. Initial Setup

[0268] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[0269] 2. Recording screen operations and audio.

[0270] The user clicks the "Start Recording" button. This causes the device to begin recording the entire current screen and save all actions as a video file. Simultaneously, the device records the user's voice explanation via the microphone.

[0271] The device activates an emotion engine, which analyzes the user's emotions in real time from their voice and facial expressions, and collects emotional data.

[0272] 3. Sending data

[0273] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and save the generated video and audio files, as well as emotion data, to a temporary directory.

[0274] The device sends these files to the server. The transmission is done using an HTTP POST request, and video files, audio files, and emotion data are uploaded to the server as form data.

[0275] 4. Data Processing

[0276] The server receives uploaded video files, audio files, and emotion data. After receiving the data, it passes the audio files to a speech recognition engine to begin the process of converting them to text. Specifically, speech recognition technology is used to convert the audio data into text data.

[0277] The server also analyzes emotional data and records it along with the corresponding timestamp.

[0278] 5. Generation of the operation manual

[0279] The server integrates text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual.

[0280] In the generated operation manual, the user's emotions at each operation step are also displayed. For example, along with an explanation such as "A new file has been created", the emotions felt by the user at that moment (such as joy or confusion) are also shown.

[0281] 6. Distribution of the Manual

[0282] The completed operation manual is generated in PDF or HTML format. In the generated manual, in addition to the operation procedures, the emotion data is visually shown, so that it is clear what emotions the user had during which operations.

[0283] The server generates a download link to provide this manual to the user. This link is sent to the system dashboard or the user's email address.

[0284] The user can click on the provided download link to download and view the operation manual. This operation manual can be referred to as needed to support the use of the system.

[0285] Specific Example

[0286] 1. The user clicks the "Create New Project" button

[0287] User: "Here, I will create a new project."

[0288] The terminal records the voice simultaneously with the operation and records it together with the timestamp.

[0289] At the same time, the emotion engine analyzes the user's emotions and collects emotion data such as "joy" and "expectation".

[0290] 2. Data Transmission and Processing after Stopping the Recording

[0291] The terminal sends the recorded data to the server. The server uses speech recognition to extract the text "Here we will create a new project" and matches it with the video of the operation.

[0292] Emotional data is also analyzed and recorded along with the corresponding timestamp.

[0293] 3. Example of manual generation

[0294] The server organizes the steps on the timeline and generates a page for the "Create a new project" section, combining the corresponding screenshots, text, and sentiment data at that time.

[0295] The manual displays an emotion such as "Joy" below an explanation like, "Here we will create a new project."

[0296] In this way, the system of the present invention can generate a detailed and intuitive operation manual, including emotional data, based on the user's operations and explanations. This enriches the user's experience and allows for the provision of feedback utilizing emotional information.

[0297] The following describes the processing flow.

[0298] Step 1:

[0299] The user starts the system and accesses the initial setup screen. The terminal displays a dialog box requesting permission for microphone and screen capture. Once the user grants these permissions, the terminal confirms the settings and completes the setup.

[0300] Step 2:

[0301] The user clicks the "Start Recording" button. Thereby, the terminal starts recording the entire screen and saves the operations as a video file. At the same time, the terminal records the user's voice through the microphone. Also, the terminal activates the emotion engine and analyzes the emotion in real time from the user's facial expressions and voice.

[0302] Step 3:

[0303] While the user performs screen operations, the user explains the operations in voice in real time. For example, when clicking on "New File" and selecting it from the "File" menu, the user explains in voice "Create a new file here". The terminal records this operation and voice together with a timestamp. At the same time, it analyzes the emotion data and records it together with the timestamp.

[0304] Step 4:

[0305] The user clicks the "Stop Recording" button. Thereby, the terminal stops recording and voice recording, and saves the generated video file, voice file, and emotion data in a temporary directory.

[0306] Step 5:

[0307] The terminal sends the video file, voice file, and emotion data saved in the temporary directory to the server. The file transmission is performed using an HTTP POST request, and the video file, voice file, and emotion data are uploaded to the server as form data.

[0308] Step 6:

[0309] The server receives the uploaded video file, voice file, and emotion data and starts processing. First, it passes the voice file to the voice recognition engine to convert it into text. In this process, for example, voice data such as "Create a new file here" is converted into character data.

[0310] Step 7:

[0311] The server analyzes the video file and generates frame-by-frame screenshots on the timeline. Based on the video's timestamps, it extracts specific action events (e.g., button clicks or form inputs) and takes screenshots of the moments when these events occurred.

[0312] Step 8:

[0313] The server integrates text, screenshots, and sentiment data to generate an operation manual. The manual sequentially displays text and corresponding screenshots as operating procedures, and also shows the user's sentiment data at each step. For example, under the explanation "Create a new file," the corresponding screenshot and the sentiment data "Joy" are displayed.

[0314] Step 9:

[0315] The server saves the generated operation manual in PDF or HTML format and generates a download link for the user. This link is sent to the system dashboard or to the user's email address.

[0316] Step 10:

[0317] Users can download and view the operation manual by clicking the provided download link. This operation manual visually presents user sentiment data along with the operating procedures, making it easier to understand intuitively.

[0318] (Example 2)

[0319] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0320] Conventional operation manual generation systems could record user instructions and screen operations, but they had the problem of not being able to obtain feedback that included the user's emotions. As a result, feedback based on the user's emotions was not provided, leading to a lack of deeper understanding of how to operate the system.

[0321] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting audio files to text, means for generating an operation manual based on the converted text and video files, means for analyzing the user's emotions in real time and collecting emotion data, and means for integrating the collected emotion data into the operation manual. This makes it possible to generate a more detailed and intuitive operation manual that reflects not only the user's operation instructions but also their emotional information at that time.

[0322] A "user" is the entity that operates the system and performs various inputs.

[0323] "Means of recording screen footage" refers to functions or devices that record the user's screen in video format.

[0324] "Means of recording audio via a microphone" refers to functions or devices for recording the voice spoken by a user.

[0325] "Means of sending to a server" refers to the technology and processes used to upload recorded video files and audio files to a server via a network.

[0326] "Methods for converting audio files to text" refers to the process of converting audio data into text data using speech recognition technology.

[0327] "Means for generating operation manuals" refers to technologies and processes for creating documents that visually show user operation procedures based on acquired audio text and video files.

[0328] "Means of providing operation manuals to users" refers to the processes and functions for providing the generated operation manuals in a user-accessible format (e.g., PDF or HTML).

[0329] "Means of analyzing emotions in real time and collecting emotional data" refers to technologies and devices that analyze a user's facial expressions and voice to identify and record the emotions the user is feeling in real time.

[0330] "Means of integrating emotional data into the operation manual" refers to the techniques and processes used to link collected emotional data to each operation step in the operation manual and ensure consistency.

[0331] The system according to the present invention can instantly record the computer screen operations performed by the user and their explanations, and further recognize, analyze, and utilize the user's emotions by combining them with an emotion engine. Specific embodiments of the present invention will be described in detail below.

[0332] Overall overview

[0333] This system is implemented through a process that records the user's screen operations and their voice explanations via a microphone, sends this data to a server for analysis, and ultimately generates an operation manual. Furthermore, it analyzes the user's emotions using an emotion engine and incorporates this analysis into the operation manual.

[0334] Initial setup

[0335] When the user starts the system, the terminal displays an initial setup screen and a dialog box prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the preparation for recording and audio.

[0336] Screen operation and audio recording

[0337] When the user clicks the "Start Recording" button, the device begins recording the entire current screen and saves all operations as a video file. Simultaneously, the device records the user's voice description via the microphone. The device also activates an emotion engine and uses OpenCV to analyze the user's voice and facial expressions in real time, collecting emotion data.

[0338] Sending data

[0339] When the user clicks the "Stop Recording" button, the device stops recording video and audio, and saves the generated video and audio files, as well as emotion data, to a temporary directory. The device then uses the Python requests library to send these files to the server via an HTTP POST request.

[0340] Data processing

[0341] The server receives uploaded video and audio files, as well as sentiment data. Audio files are converted to text using the Google Cloud Speech-to-Text API. Additionally, the server analyzes the sentiment data and records it along with the corresponding timestamp.

[0342] Generating an operation manual

[0343] The server integrates the converted text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. The manual also visually displays the user's emotions at each operation step. Specifically, it generates a PDF version of the manual using LaTeX and also provides it in HTML format.

[0344] Distribution of manuals

[0345] The completed operation manual is generated in PDF or HTML format, and the server uses Flask to generate download links. Users can download and view the operation manual by clicking the download link provided via the dashboard or their email address.

[0346] Specific examples and prompt statements

[0347] Specific example

[0348] The scenario envisions a user clicking the "Create New Project" button and being instructed, "Here, we will create a new project." The device simultaneously records the user's actions and voice, and uses an emotion engine to analyze the user's emotions, such as "joy" and "expectation." Subsequently, the server generates an operation manual based on the collected data and provides it to the user.

[0349] Example of a prompt

[0350] "Record the steps for creating a new project and generate a manual that includes emotional data along with the operating instructions."

[0351] "Based on the recorded operating procedures, use the emotion engine to analyze the user's emotions and create a detailed operating manual."

[0352] The above describes specific embodiments for carrying out the present invention. By utilizing the present invention, it becomes possible to enrich the user's operational experience and provide valuable feedback that utilizes emotional information.

[0353] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0354] Step 1:

[0355] Initial setup

[0356] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant permission for microphone and screen capture. When the user clicks "Allow," the terminal confirms and saves these settings.

[0357] Input: User actions (system startup, permission granting)

[0358] Data processing: Permission verification and saving of settings

[0359] Output: Microphone and screen capture ready.

[0360] Step 2:

[0361] Screen recording

[0362] The user clicks the "Start Recording" button. The device begins recording the entire current screen and saves all user actions as an MP4 video file. FFmpeg is used as the recording tool.

[0363] Input: User action (click of the "Start Recording" button)

[0364] Data processing: Screen recording

[0365] Output: MP4 video file (recording.mp4)

[0366] Step 3:

[0367] Audio recording

[0368] The device simultaneously starts microphone input and saves the user's voice as a WAV audio file. Pyaudio is used as the audio library.

[0369] Input: User's voice

[0370] Data processing: Audio recording

[0371] Output: WAV format audio file (audio.wav)

[0372] Step 4:

[0373] Emotion analysis

[0374] The device uses OpenCV to analyze camera input, analyzes the user's emotions in real time from their voice and facial expressions, and saves the emotion data in JSON format.

[0375] Input: User's facial expressions, voice

[0376] Data processing: Real-time analysis of emotions

[0377] Output: Emotion data file in JSON format (emotions.json)

[0378] Step 5:

[0379] Sending data

[0380] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The generated video file, audio file, and emotion data are saved to a temporary directory. The device uses the Python requests library to send these files to the server via an HTTP POST request.

[0381] Input: User action (click of the "Stop Recording" button), generated file

[0382] Data processing: Saving and sending files

[0383] Output: File sent to the server

[0384] Step 6:

[0385] Converting audio data to text

[0386] The server sends the received audio file to the Google Cloud Speech-to-Text API, where it converts the audio data into text.

[0387] Input: Audio file (audio.wav)

[0388] Data processing: Text conversion using speech recognition.

[0389] Output: Text data

[0390] Step 7:

[0391] Analysis of emotional data

[0392] The server analyzes the received emotion data file and stores each emotion data item in the database along with its corresponding timestamp.

[0393] Input: Emotion data file (emotions.json)

[0394] Data processing: Analysis of emotional data and its correspondence with timestamps.

[0395] Output: Emotional data stored in the database

[0396] Step 8:

[0397] Generating an operation manual

[0398] The server integrates the received text data and timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. It generates the manual in PDF format using LaTeX, and also in HTML format.

[0399] Input: Text data, timestamp information, video files

[0400] Data processing: Capture and integrate screenshots, generate consistent operation manuals.

[0401] Output: Operation manual in PDF and HTML formats

[0402] Step 9:

[0403] Provision of operation manual

[0404] The server generates a download link for the operation manual using Flask and provides it to the user via their dashboard and email. Users can download and view the operation manual by clicking the download link.

[0405] Input: Generated operation manual

[0406] Data processing: Creating download links

[0407] Output: Providing download links to users

[0408] (Application Example 2)

[0409] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0410] Conventional user manual generation systems lacked the ability to consider user emotions, resulting in a failure to integrate emotional information into operating procedures and explanations. Especially in online learning platforms, there is a demand for learning content that accurately reflects the instructor's emotional state and explanations. Therefore, an intuitive user manual generation system with emotion analysis capabilities is needed to deepen users' understanding of the operations.

[0411] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting emotion data, means for converting audio files into text, and means for generating an operation manual based on the converted text, video files, and emotion data. This makes it possible to provide an operation manual that integrates user operations and explanations with the emotion information at that time.

[0412] "User" refers to the entity that operates this system and uses it to record screen operations and audio.

[0413] "Screen images" refer to the visual information displayed on the monitor of a computer operated by a user.

[0414] "Means of recording" refers to a device or software that captures and records video of the screen being operated by the user.

[0415] A "microphone" refers to a device that converts sound into electrical signals and records the user's voice.

[0416] "Means of recording audio" refers to a device or software that uses a microphone to record the user's voice.

[0417] "Means for analyzing user facial expressions" refers to devices or software that capture the user's facial movements and expressions, analyze them, and obtain emotional data.

[0418] "Emotional data" refers to emotional information analyzed from the user's facial expressions and voice.

[0419] A "server" refers to a computer system used to store, process, and generate operation manuals for recorded data.

[0420] A "video file" refers to a digital file containing recorded data of the user's screen.

[0421] An "audio file" refers to a digital file containing user voice data recorded via a microphone.

[0422] "Speech recognition technology" refers to algorithms and techniques for converting speech data into text data.

[0423] An "operation manual" refers to a set of instructions or manuals that include user actions, explanations, and emotional data.

[0424] "Means of provision" refers to devices or software for distributing or displaying the generated operation manual to the user.

[0425] "PDF format" is an abbreviation for Portable Document Format, and refers to a digital file format that allows users to view and print operation manuals.

[0426] "HTML format" is an abbreviation for HyperText Markup Language, and refers to a file format used to display instruction manuals on the web.

[0427] System Overview

[0428] Initial setup

[0429] The user starts the system and accesses the initial setup screen. At this point, the terminal displays a dialog box prompting the user to grant permission for microphone and screen capture. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[0430] Screen operation and audio recording

[0431] When the user clicks the "Start Recording" button, the device begins recording the entire current screen and saves all operations as a video file. Simultaneously, the device records the user's voice explanation via the microphone. In addition, the device captures the user's facial expressions and analyzes and collects emotion data in real time.

[0432] Sending data

[0433] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The generated video file, audio file, and emotion data are then saved to a temporary directory and sent to the server. The transmission is performed using an HTTP POST request, and the video file, audio file, and emotion data are uploaded to the server as form data.

[0434] Data processing

[0435] The server receives uploaded video files, audio files, and sentiment data. After receiving the data, the server passes the audio files to a speech recognition engine to begin the process of converting them to text. Specifically, speech recognition technology is used to convert the audio data into text data. The server also analyzes the sentiment data and records it along with the corresponding timestamp.

[0436] Generating an operation manual

[0437] The server integrates text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. The generated operation manual also displays the user's emotions at each operation step.

[0438] Distribution of manuals

[0439] The completed user manual is generated in PDF or HTML format. In addition to the operating procedures, the generated manual visually displays sentiment data, clearly showing what emotions the user experienced during each operation. The server generates a download link to provide this manual to the user. This link is sent to the system dashboard or to the user's email address. The user can download and view the user manual by clicking the provided download link.

[0440] Program Processing Overview

[0441] Hardware and software to be used

[0442] Hardware:

[0443] Webcam: Used for video and sentiment analysis.

[0444] Microphone: Used for voice recording

[0445] software:

[0446] OpenCV: Used for video capture and saving

[0447] pyaudio: Used for audio recording

[0448] wave: For saving audio files

[0449] Request: For sending files via HTTP POST request.

[0450] EmotionRecognizer: For emotion analysis

[0451] speech_recognition: For speech recognition

[0452] Processing details

[0453] The terminal simultaneously performs video and audio recording and emotional data analysis. The recorded video files, audio files, and emotional data are sent to the server. The server converts the audio data to text and integrates the emotional data with timestamps to generate an operation manual. Finally, the generated manual is distributed to the user in PDF or HTML format.

[0454] Specific example

[0455] If the system is started: "start_recording('Lesson Start', 'Audio Explanation')"

[0456] When the operation is finished: "analyze_emotions('lesson_audio.wav', 'lesson_video.avi')"

[0457] Generating a lesson operation manual: "generate_manual('generated text', 'emotion data', 'lesson_manual.txt')"

[0458] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0459] Step 1:

[0460] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant permission for microphone and screen capture. The input is the user's actions, and the output is the state of readiness for completing the microphone and screen capture permission settings. Specifically, the terminal verifies the user's permission and configures the microphone and screen capture settings.

[0461] Step 2:

[0462] When the user clicks the "Start Recording" button, the device begins recording the entire current screen. The input is the click of the "Start Recording" button, and the output is the start of recording the video file. Specifically, the device uses OpenCV to capture and save the screen video.

[0463] Step 3:

[0464] The device simultaneously records the user's voice description via the microphone. The input is the user's voice, and the output is an audio file. Specifically, the device uses pyaudio to record the audio and saves it using the wave library.

[0465] Step 4:

[0466] The device captures the user's facial expressions and analyzes and collects emotional data in real time. The input is a video of the user's face, and the output is emotional data. Specifically, the device uses its EmotionRecognizer to analyze facial expressions and saves them as emotional data.

[0467] Step 5:

[0468] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The input is the click of the "Stop Recording" button, and the output is the completion of collecting video and audio files, as well as emotion data. Specifically, it stops recording video and audio and saves the data to a temporary directory.

[0469] Step 6:

[0470] The device sends recorded video files, recorded audio files, and emotion data to the server. The inputs are video files, audio files, and emotion data, and the output is the uploading of this data to the server. Specifically, the device uses the requests library to send an HTTP POST request and upload the data to the server.

[0471] Step 7:

[0472] The server receives uploaded video files, audio files, and emotion data. The input is the transmitted data, and the output is the completion of data reception on the server side. Specifically, the server saves the files to storage.

[0473] Step 8:

[0474] The server passes the audio file to the speech recognition engine and begins the process of converting it to text. The input is an audio file, and the output is text data. Specifically, the server uses the speech_recognition library to convert the audio to text.

[0475] Step 9:

[0476] The server analyzes sentiment data and records it along with a timestamp. The input is sentiment data, and the output is the sentiment analysis result with a timestamp. Specifically, the server analyzes the sentiment data and integrates it with the corresponding time information.

[0477] Step 10:

[0478] The server generates operation manuals based on text, video files, and sentiment data. Inputs are text data, screenshots, and sentiment data, while output is the operation manual. Specifically, the server organizes this data into a consistent set of operating procedures and creates the manual in PDF or HTML format.

[0479] Step 11:

[0480] The server provides the user with a completed operation manual. The input is the generated operation manual, and the output is a download link. Specifically, the server creates a download link and sends it to the system dashboard or the user's email address.

[0481] Example of a prompt

[0482] If the system is started: "start_recording('Lesson Start', 'Audio Explanation')"

[0483] When the operation is finished: "analyze_emotions('lesson_audio.wav', 'lesson_video.avi')"

[0484] Generating a lesson operation manual: "generate_manual('generated text', 'emotion data', 'lesson_manual.txt')"

[0485] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0486] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0487] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0488] [Second Embodiment]

[0489] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0490] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0491] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0492] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0493] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0494] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0495] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0496] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0497] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0498] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0499] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0500] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0501] This invention is a system that instantly records the computer screen operations performed by a user and their explanations, and later automatically generates an operation manual. The embodiments thereof are described in detail below.

[0502] System Overview

[0503] 1. Initial Setup

[0504] The user starts the system and grants permission for microphone and screen capture.

[0505] The device checks the microphone and screen capture settings and completes the setup.

[0506] 2. Recording screen operations and audio.

[0507] Recording begins when the user clicks the "Start Recording" button.

[0508] The device records all actions performed on the screen as video files.

[0509] At the same time, the device records the user's voice description through the microphone.

[0510] As a concrete example, consider a scenario where a user clicks "New File" and says aloud, "I have created a new file." This action and the accompanying audio are recorded simultaneously.

[0511] 3. Sending data

[0512] Once all recording and recording are complete, the user clicks the "Stop Recording" button.

[0513] The terminal sends the generated video and audio files to the server. This transmission takes place over the internet.

[0514] 4. Data Processing

[0515] The server analyzes the received video and audio files.

[0516] The server uses speech recognition technology to convert audio files into text data. For example, the audio file "A new file has been created" is converted to text "A new file has been created".

[0517] 5. Generation of the operation manual

[0518] The server combines the text with the corresponding timestamp information and integrates it with a specific portion (screenshot) of the video file.

[0519] Specifically, text and screenshots are compiled into a consistent manual as a series of steps.

[0520] 6. Distribution of manuals

[0521] The completed operation manual will be generated in PDF or HTML format.

[0522] The server will provide this manual to users by displaying a download link via email or on the dashboard.

[0523] Users can download and view the manual.

[0524] Specific example

[0525] 1. A scene where the user clicks the "Create New Project" button and the process is explained verbally.

[0526] User: "I'll create a new project here."

[0527] The device records audio simultaneously with operation, along with a timestamp.

[0528] 2. Data transmission and processing after recording stops

[0529] The device sends the recorded data to the server.

[0530] The server uses speech recognition to extract the text "We will create a new project here" and matches it with the video of the operation.

[0531] 3. Example of manual generation

[0532] The server organizes the steps on the timeline and generates a page for the "1. Create a new project" section, combining corresponding screenshots and text.

[0533] In this way, by using the system of the present invention, an efficient and accurate operation manual is automatically generated based on the user's operations and explanations. The implementation of this invention significantly reduces the effort required to create manuals and enables the provision of highly accurate documentation.

[0534] The following describes the processing flow.

[0535] Step 1:

[0536] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[0537] Step 2:

[0538] The user clicks the "Start Recording" button. This causes the device to begin recording the entire current screen, saving all actions as a video file. Simultaneously, the device records the user's voice via the microphone. Specifically, the screen capture and audio recording processes are executed at the same time.

[0539] Step 3:

[0540] As the user performs screen operations, the system provides real-time audio explanations of those operations. For example, when the user clicks "New File" and selects it from the "File" menu, the system will provide the audio explanation, "Now you will create a new file." The device records these operations and audio, along with a timestamp.

[0541] Step 4:

[0542] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and saves the generated video and audio files to a temporary directory.

[0543] Step 5:

[0544] The terminal sends video and audio files stored in a temporary directory to the server. The files are sent using an HTTP POST request. Specifically, the video and audio files are uploaded to the server as form data.

[0545] Step 6:

[0546] The server receives the uploaded video and audio files. After receiving them, it passes the audio files to a speech recognition engine to begin the process of converting them to text. Here, speech recognition technology (e.g., a speech recognition API) is used to convert the audio data into text data.

[0547] Step 7:

[0548] The server analyzes the video file and generates frame-by-frame screenshots. Based on the video timeline, it extracts specific action events (e.g., button clicks or form inputs) and takes screenshots of the moments when these events occur.

[0549] Step 8:

[0550] The server integrates the converted text and screenshots to generate an operation manual. Based on the text and timestamp information, the screenshots are arranged as a sequence of steps to create a consistent operation manual. The generated operation manual is in PDF or HTML format.

[0551] Step 9:

[0552] The server generates a download link to provide the user with the completed operation manual. This link is sent to the system dashboard or to the user's email address.

[0553] Step 10:

[0554] Users can download and view the operation manual by clicking the provided download link. This operation manual is available for reference as needed and supports the use of the system.

[0555] (Example 1)

[0556] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0557] Traditional methods for creating user manuals require users to manually record each operation step and then transcribe the information, which is time-consuming and labor-intensive. Furthermore, the operation and its explanation may not always be synchronized, potentially leading to the creation of inaccurate manuals. Therefore, there is a need for a system that allows users to efficiently and accurately generate user manuals without having to spend time on incidental tasks.

[0558] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0559] In this invention, the server includes means for recording video of the screen operated by the user, means for recording the user's voice through a microphone, means for transmitting the recorded video file and the recorded audio file from the device to an information processing device, means for converting the audio file into text data on the information processing device, means for generating an operation manual based on the converted text data and video file, and means for providing the generated operation manual to the user. This makes it possible for the user to generate an operation manual simply and efficiently.

[0560] A "user" refers to anyone who intends to use this system to generate an operation manual.

[0561] "Screen images" refers to all visual information displayed on the screen of a device operated by the user.

[0562] A "microphone" is an audio input device used to record the user's voice.

[0563] A "recorded video file" is a digital file created by recording the video of the screen being operated by the user.

[0564] "Recorded audio files" refer to digital files generated by recording the user's voice.

[0565] A "device" refers to electronic devices such as computers and smartphones that are operated by the user.

[0566] An "information processing device" refers to a server or cloud computer used to process video and audio files transmitted from a device.

[0567] "Speech recognition technology" is a technology that converts audio files into text data; it analyzes audio data and turns it into text.

[0568] "Text data" refers to text information converted by speech recognition technology.

[0569] An "operation manual" refers to a document that summarizes the user's operating procedures and their explanations, and is generated based on video files and text data.

[0570] "Electronic document format" refers to digital document formats such as PDF and HTML.

[0571] "Web page format" refers to a web-based document that can be viewed on an internet browser.

[0572] This invention provides a system that instantly records user-performed computer screen operations and their descriptions, and automatically generates an operation manual later. This system allows users to efficiently and accurately generate operation manuals.

[0573] Hardware and software to be used

[0574] Terminal: A device used by the user to perform operations (e.g., personal computer, smartphone)

[0575] Microphone: A voice input device for recording the user's voice.

[0576] Screen capture tool: Software that records screen activity (e.g., OBS Studio)

[0577] Server: An information processing device that performs data processing and manual generation (e.g., cloud computing service).

[0578] Speech recognition technology: A technology that converts speech data into text data (e.g., Google Speech-to-Text API).

[0579] Image processing software: Technology for automatically generating screenshots from video (e.g., OpenCV)

[0580] HTML template engine: A technology for generating operation manuals in HTML format (e.g., Jinja2).

[0581] PDF generation software: Technology for generating manuals in PDF format (e.g., wkhtmltopdf).

[0582] Process Overview

[0583] 1. The user starts the system and grants permission for microphone and screen capture.

[0584] 2. The device checks the screen capture and microphone settings and completes the setup.

[0585] 3. When the user clicks the "Start Recording" button, the device simultaneously records the video and audio of the screen. For example, the user clicks "New File" and says aloud, "A new file has been created."

[0586] 4. When the user clicks the "Stop Recording" button, the device sends the recorded video and audio files to the server.

[0587] 5. The server analyzes the received video and audio files and uses speech recognition technology to convert the audio files into text data. For example, the audio data "A new file has been created" is converted into the text data "A new file has been created".

[0588] 6. The server integrates text data and timestamp information with video files to generate a consistent operation manual. The operation manual is output in HTML or PDF format using an HTML template engine or PDF generation software.

[0589] 7. The server will provide users with a download link for the generated operation manual via email or on the dashboard. Users can download and view the operation manual from this link.

[0590] Specific example

[0591] A scene where the user clicks the "Create New Project" button and the process is explained via voice.

[0592] User: "I'll create a new project here."

[0593] The device records audio simultaneously with this operation, along with a timestamp.

[0594] Data transmission and processing after recording stops

[0595] The device sends the recorded data to the server.

[0596] The server uses speech recognition to extract the text "We will create a new project here" and matches it with the video of the operation.

[0597] Manual generation example

[0598] The server organizes the steps on the timeline and generates a page for the step "1. Create a new project," combining corresponding screenshots and text.

[0599] Examples of prompt statements

[0600] 1. "Please record the steps to create a new project using both screen operations and audio."

[0601] 2. "Please continue by explaining the current screen operation aloud."

[0602] 3. "Once recording is complete, please click the 'End' button to send the data to the server."

[0603] The above details describe the embodiments of the present invention. This system enables users to generate highly accurate operation manuals in a short amount of time.

[0604] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0605] Step 1:

[0606] The user starts the system and grants permission for microphone and screen capture.

[0607] Input: Command to start the system, permission request popup.

[0608] Output: Permission granted

[0609] Specific action: When the user starts the system, a pop-up appears asking for permission to use the microphone and screen capture. The user clicks "Allow".

[0610] Step 2:

[0611] The device checks the microphone and screen capture settings and completes the setup.

[0612] Input: User-granted permissions

[0613] Output: Microphone and screen capture ready to use status

[0614] Specific operation: The device automatically checks the settings for the microphone and screen capture tool (e.g., OBS Studio) and displays a "Ready" message.

[0615] Step 3:

[0616] The user clicks the "Start Recording" button.

[0617] Input: User click of the "Start Recording" button.

[0618] Output: Trigger for recording start

[0619] Specific action: The user clicks the "Start Recording" button in the application window.

[0620] Step 4:

[0621] The device records all actions performed on the screen.

[0622] Input: Trigger for recording start

[0623] Output: Video data being recorded

[0624] Specific action: The OBS Studio software begins recording the user's screen activity. For example, it records the user clicking the "Create New Project" button.

[0625] Step 5:

[0626] The device records the user's voice description via the microphone.

[0627] Input: Trigger for recording start, user's voice

[0628] Output: Audio data being recorded

[0629] Specific actions: As the user performs an action, the system provides a voice explanation of that action. For example, it might say, "Here, we will create a new project."

[0630] Step 6:

[0631] The user clicks the "Stop Recording" button.

[0632] Input: User click of the "Stop Recording" button.

[0633] Output: Trigger for recording termination

[0634] Specific action: The user clicks the "Stop Recording" button after completing the operation.

[0635] Step 7:

[0636] The device sends the recorded video file and the recorded audio file to the server.

[0637] Input: Trigger for recording termination, video file, audio file

[0638] Output: Status of successful file transfer to server

[0639] Specific operation: The device uploads video and audio files to the server via the internet connection. For example, it sends files using an API and displays a "transmission complete" notification.

[0640] Step 8:

[0641] The server analyzes the received video and audio files.

[0642] Input: Video file, audio file

[0643] Output: Analyzed audio data, video data

[0644] Specific operation: The server temporarily stores the received data and separates the video file from the audio file.

[0645] Step 9:

[0646] The server converts the audio file into text data.

[0647] Input: Audio file

[0648] Output: Character data (text file)

[0649] Specific operation: Use the Google Speech-to-Text API to convert audio data into text data. For example, the audio "I have created a new file" will be converted into text.

[0650] Step 10:

[0651] The server integrates text and timestamp information with the video file.

[0652] Input: Text data, timestamp, video file

[0653] Output: Operation procedure data, screenshots

[0654] Specific operation: Use OpenCV to obtain appropriate screenshots from video, and generate a consistent instruction manual based on text data and timestamp information.

[0655] Step 11:

[0656] The server generates an operation manual.

[0657] Input: Operation procedure data, screenshots

[0658] Output: Operation manual (PDF or HTML format)

[0659] Specific operation: Use the Jinja2 template engine to generate an operation manual in HTML format, and then convert it to PDF format using wkhtmltopdf.

[0660] Step 12:

[0661] The server provides users with an operation manual.

[0662] Input: Generated operation manual

[0663] Output: Download link, notification email

[0664] Specific operation: The server saves the generated manual and provides a download link to the user. The user can download the manual by displaying the link in an email or on the dashboard.

[0665] (Application Example 1)

[0666] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0667] Creating manuals for operating robots and machinery within a factory is time-consuming and labor-intensive, and accurately conveying advanced operations and complex procedures is difficult. Furthermore, training new workers is lengthy and prone to errors, highlighting the need for improved production efficiency and safety.

[0668] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0669] In this invention, the server includes means for recording video of the screen operated by the user, means for recording the user's voice through a microphone, means for using a visual device to record video of the real environment, means for transmitting the recorded video file and recorded audio file to the server, means for converting the audio file to text on the server, means for generating an operation manual based on the converted text and video file, and means for providing the generated operation manual to the user. This enables accurate and rapid recording of complex procedures and operations within the factory, efficient automatic generation of operation manuals, reduction of worker training time, and prevention of errors.

[0670] "Means for recording video of the user's screen" refers to a function that records all operations performed on the screen of a computer or device being operated by the user in video format.

[0671] "Means of recording user voice via microphone" refers to a function that uses a microphone to record the voice of the user when they give explanations or instructions while operating the device.

[0672] "Means of using visual devices to record images of the real environment" refers to a function that allows users to record images of their surroundings or the object they are operating in real time using devices such as smart glasses or cameras.

[0673] "Means for sending recorded video files and recorded audio files to a server" refers to a function for uploading recorded video data and audio data to a server via the internet or other means.

[0674] "Method for converting audio files to text on the server" refers to a function that uses speech recognition technology on the server side to automatically convert transmitted audio data into text format.

[0675] "A means of generating operation manuals based on converted text and video files" refers to a function that automatically creates consistent operation procedures and manuals by combining audio-to-text data and recorded video data.

[0676] "Means of providing the generated operation manual to the user" refers to a function that outputs the completed operation manual in PDF or HTML format, allowing the user to download or view it.

[0677] This invention is a system that records robot operation and maintenance procedures in a factory in real time and automatically generates an operation manual based on that data. The embodiments of this system are described in detail below.

[0678] System Overview

[0679] 1. Initial Setup

[0680] The user starts the system and grants permission for microphone and screen capture. Additionally, the user wears a visual device (e.g., smart glasses or a camera) and confirms the settings.

[0681] The device checks the settings for microphone, screen capture, and visual devices, and then completes the setup.

[0682] 2. Recording screen operations and audio.

[0683] Recording begins when the user clicks the "Start Recording" button.

[0684] The device records the user's actions on the device screen as a video file, and simultaneously records video of the real environment through a visual device.

[0685] At the same time, the device records the user's voice description through the microphone.

[0686] 3. Sending data

[0687] Once all recording and recording are complete, the user clicks the "Stop Recording" button.

[0688] The terminal sends the generated video and audio files to the server. This transmission takes place over the internet.

[0689] 4. Data Processing

[0690] The server analyzes the received video and audio files and uses speech recognition technology to convert the audio files into text data. For example, the audio data "I have created a new project" is converted to text "I have created a new project".

[0691] 5. Generation of the operation manual

[0692] The server combines the text with the corresponding timestamp information and integrates it with a specific portion (screenshot) of the video file.

[0693] Specifically, text and screenshots are compiled into a consistent manual as a series of steps.

[0694] 6. Distribution of manuals

[0695] The completed operation manual will be generated in PDF or HTML format.

[0696] The server will provide this manual to users by displaying a download link via email or on the dashboard.

[0697] Users can download and view the manual.

[0698] Specific example

[0699] Consider a scenario where factory workers wear smart glasses to perform robot maintenance procedures. For example, they might record the procedure for replacing a hydraulic cylinder.

[0700] The user performs the task while explaining verbally, "Here, we will remove the hydraulic cylinder," and both the operation and the audio are recorded simultaneously.

[0701] After recording stops, the terminal sends the recorded data to the server, converts the audio data into text, and generates a manual.

[0702] Example of a prompt

[0703] "Use video and audio capture to record robot operation procedures within the factory. After recording is complete, convert the audio data into text and generate an operation manual along with the video frames. For example, if you perform a task while explaining, 'Here we remove the hydraulic cylinder,' ensure that this procedure and explanation are integrated into a single procedure in the manual."

[0704] In this way, by using the system of the present invention, complex procedures and operations within a factory can be accurately and quickly recorded, and efficient operation manuals can be automatically generated, thereby reducing worker training time and preventing operational errors.

[0705] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0706] Step 1:

[0707] Initial settings:

[0708] The user starts the system and grants permissions for the microphone, screen capture, and visual devices (3D glasses or camera). The terminal verifies that these devices are working correctly and completes the setup. The input is the permission and setup required for each device, and the output is the result of verifying that the devices are ready for use.

[0709] Step 2:

[0710] Start recording screen operations and audio:

[0711] The recording process begins when the user clicks the "Start Recording" button. The device records all operations performed on the computer screen as a video file using screen capture software (e.g., PyAutoGUI). Simultaneously, audio recording software (e.g., PyAudio) records the user's voice explanations via the microphone. The visual device records the user's viewpoint in real time. Input data consists of user operations, audio, and visual environment, and the output is a recorded file.

[0712] Step 3:

[0713] Sending data:

[0714] Once recording is complete, the user clicks the "Stop Recording" button. The device then sends the recorded video files (screen captures and video from the visual device) and audio files to the server. The transmission takes place over the internet. The input data consists of the recorded video and audio files, and the output is the data uploaded to the server.

[0715] Step 4:

[0716] Data processing:

[0717] The server analyzes the received video and audio files. Audio files are converted to text using speech recognition technology (e.g., Google Cloud Speech-to-Text API). The input is an audio file, and the output is the corresponding text data. Video files are also organized based on timestamps.

[0718] Step 5:

[0719] Generating the operation manual:

[0720] The server combines the converted text data with corresponding timestamp information and integrates it with specific portions (screenshots) of the video file. This generates a consistent operation manual. The input is text data and video files, and the output is an integrated operation manual file.

[0721] Step 6:

[0722] Distribution of manuals:

[0723] The server generates a completed operation manual in PDF or HTML format and makes it available for users to download. The download link is provided via email or on the dashboard. The input is the generated operation manual, and the output is the download link that users can access.

[0724] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0725] This invention provides a system that instantly records user computer screen operations and their explanations, and later automatically generates an operation manual. Furthermore, by combining it with an emotion engine, it enables the recognition and utilization of user emotions. The embodiments thereof will be described in detail below.

[0726] System Overview

[0727] 1. Initial Setup

[0728] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[0729] 2. Recording screen operations and audio.

[0730] The user clicks the "Start Recording" button. This causes the device to begin recording the entire current screen and save all actions as a video file. Simultaneously, the device records the user's voice explanation via the microphone.

[0731] The device activates an emotion engine, which analyzes the user's emotions in real time from their voice and facial expressions, and collects emotional data.

[0732] 3. Sending data

[0733] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and save the generated video and audio files, as well as emotion data, to a temporary directory.

[0734] The device sends these files to the server. The transmission is done using an HTTP POST request, and video files, audio files, and emotion data are uploaded to the server as form data.

[0735] 4. Data Processing

[0736] The server receives uploaded video files, audio files, and emotion data. After receiving the data, it passes the audio files to a speech recognition engine to begin the process of converting them to text. Specifically, speech recognition technology is used to convert the audio data into text data.

[0737] The server also analyzes emotional data and records it along with the corresponding timestamp.

[0738] 5. Generation of the operation manual

[0739] The server integrates text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual.

[0740] The generated operation manual also displays the user's emotions at each operation step. For example, along with a description such as "You have created a new file," it also shows the emotions the user felt at that moment (e.g., joy or confusion).

[0741] 6. Distribution of manuals

[0742] The completed user manual will be generated in PDF or HTML format. In addition to the operating procedures, the generated manual will visually display emotional data, making it clear what emotions the user felt during each operation.

[0743] The server generates a download link to provide this manual to the user. This link is sent to the system dashboard or to the user's email address.

[0744] Users can download and view the operation manual by clicking the provided download link. This operation manual is available for reference as needed and supports the use of the system.

[0745] Specific example

[0746] 1. The user clicks the "Create New Project" button.

[0747] User: "I'll create a new project here."

[0748] The device records audio simultaneously with operation, along with a timestamp.

[0749] At the same time, the emotion engine analyzes the user's emotions and collects emotional data such as "joy" and "anticipation."

[0750] 2. Data transmission and processing after recording stops

[0751] The terminal sends the recorded data to the server. The server uses speech recognition to extract the text "Here we will create a new project" and matches it with the video of the operation.

[0752] Emotional data is also analyzed and recorded along with the corresponding timestamp.

[0753] 3. Example of manual generation

[0754] The server organizes the steps on the timeline and generates a page for the "Create a new project" section, combining the corresponding screenshots, text, and sentiment data at that time.

[0755] The manual displays an emotion such as "Joy" below an explanation like, "Here we will create a new project."

[0756] In this way, the system of the present invention can generate a detailed and intuitive operation manual, including emotional data, based on the user's operations and explanations. This enriches the user's experience and allows for the provision of feedback utilizing emotional information.

[0757] The following describes the processing flow.

[0758] Step 1:

[0759] The user starts the system and accesses the initial setup screen. The terminal displays a dialog box requesting permission for microphone and screen capture. Once the user grants these permissions, the terminal confirms the settings and completes the setup.

[0760] Step 2:

[0761] The user clicks the "Start Recording" button. This causes the device to begin recording the entire screen and save the operation as a video file. Simultaneously, the device records the user's voice through the microphone. The device also activates an emotion engine to analyze the user's emotions in real time from their facial expressions and voice.

[0762] Step 3:

[0763] As the user performs screen operations, the system provides real-time audio explanations of those operations. For example, when the user clicks "New File" and selects it from the "File" menu, the system will say, "Now you will create a new file." The device records these operations and audio along with timestamps. Simultaneously, it analyzes and records sentiment data along with timestamps.

[0764] Step 4:

[0765] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and save the generated video and audio files, as well as emotion data, to a temporary directory.

[0766] Step 5:

[0767] The device sends video files, audio files, and emotion data stored in a temporary directory to the server. The files are sent using an HTTP POST request, and the video files, audio files, and emotion data are uploaded to the server as form data.

[0768] Step 6:

[0769] The server receives the uploaded video files, audio files, and emotion data and begins processing them. First, the audio files are passed to the speech recognition engine to be converted into text. In this process, for example, audio data such as "Here I will create a new file" is converted into text data.

[0770] Step 7:

[0771] The server analyzes the video file and generates frame-by-frame screenshots on the timeline. Based on the video's timestamps, it extracts specific action events (e.g., button clicks or form inputs) and takes screenshots of the moments when these events occurred.

[0772] Step 8:

[0773] The server integrates text, screenshots, and sentiment data to generate an operation manual. The manual sequentially displays text and corresponding screenshots as operating procedures, and also shows the user's sentiment data at each step. For example, under the explanation "Create a new file," the corresponding screenshot and the sentiment data "Joy" are displayed.

[0774] Step 9:

[0775] The server saves the generated operation manual in PDF or HTML format and generates a download link for the user. This link is sent to the system dashboard or to the user's email address.

[0776] Step 10:

[0777] Users can download and view the operation manual by clicking the provided download link. This operation manual visually presents user sentiment data along with the operating procedures, making it easier to understand intuitively.

[0778] (Example 2)

[0779] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0780] Conventional operation manual generation systems could record user instructions and screen operations, but they had the problem of not being able to obtain feedback that included the user's emotions. As a result, feedback based on the user's emotions was not provided, leading to a lack of deeper understanding of how to operate the system.

[0781] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting audio files to text, means for generating an operation manual based on the converted text and video files, means for analyzing the user's emotions in real time and collecting emotion data, and means for integrating the collected emotion data into the operation manual. This makes it possible to generate a more detailed and intuitive operation manual that reflects not only the user's operation instructions but also their emotional information at that time.

[0782] A "user" is the entity that operates the system and performs various inputs.

[0783] "Means of recording screen footage" refers to functions or devices that record the user's screen in video format.

[0784] "Means of recording audio via a microphone" refers to functions or devices for recording the voice spoken by a user.

[0785] "Means of sending to a server" refers to the technology and processes used to upload recorded video files and audio files to a server via a network.

[0786] "Methods for converting audio files to text" refers to the process of converting audio data into text data using speech recognition technology.

[0787] "Means for generating operation manuals" refers to technologies and processes for creating documents that visually show user operation procedures based on acquired audio text and video files.

[0788] "Means of providing operation manuals to users" refers to the processes and functions for providing the generated operation manuals in a user-accessible format (e.g., PDF or HTML).

[0789] "Means of analyzing emotions in real time and collecting emotional data" refers to technologies and devices that analyze a user's facial expressions and voice to identify and record the emotions the user is feeling in real time.

[0790] "Means of integrating emotional data into the operation manual" refers to the techniques and processes used to link collected emotional data to each operation step in the operation manual and ensure consistency.

[0791] The system according to the present invention can instantly record the computer screen operations performed by the user and their explanations, and further recognize, analyze, and utilize the user's emotions by combining them with an emotion engine. Specific embodiments of the present invention will be described in detail below.

[0792] Overall overview

[0793] This system is implemented through a process that records the user's screen operations and their voice explanations via a microphone, sends this data to a server for analysis, and ultimately generates an operation manual. Furthermore, it analyzes the user's emotions using an emotion engine and incorporates this analysis into the operation manual.

[0794] Initial setup

[0795] When the user starts the system, the terminal displays an initial setup screen and a dialog box prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the preparation for recording and audio.

[0796] Screen operation and audio recording

[0797] When the user clicks the "Start Recording" button, the device begins recording the entire current screen and saves all operations as a video file. Simultaneously, the device records the user's voice description via the microphone. The device also activates an emotion engine and uses OpenCV to analyze the user's voice and facial expressions in real time, collecting emotion data.

[0798] Sending data

[0799] When the user clicks the "Stop Recording" button, the device stops recording video and audio, and saves the generated video and audio files, as well as emotion data, to a temporary directory. The device then uses the Python requests library to send these files to the server via an HTTP POST request.

[0800] Data processing

[0801] The server receives uploaded video and audio files, as well as sentiment data. Audio files are converted to text using the Google Cloud Speech-to-Text API. Additionally, the server analyzes the sentiment data and records it along with the corresponding timestamp.

[0802] Generating an operation manual

[0803] The server integrates the converted text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. The manual also visually displays the user's emotions at each operation step. Specifically, it generates a PDF version of the manual using LaTeX and also provides it in HTML format.

[0804] Distribution of manuals

[0805] The completed operation manual is generated in PDF or HTML format, and the server uses Flask to generate download links. Users can download and view the operation manual by clicking the download link provided via the dashboard or their email address.

[0806] Specific examples and prompt statements

[0807] Specific example

[0808] The scenario envisions a user clicking the "Create New Project" button and being instructed, "Here, we will create a new project." The device simultaneously records the user's actions and voice, and uses an emotion engine to analyze the user's emotions, such as "joy" and "expectation." Subsequently, the server generates an operation manual based on the collected data and provides it to the user.

[0809] Example of a prompt

[0810] "Record the steps for creating a new project and generate a manual that includes emotional data along with the operating instructions."

[0811] "Based on the recorded operating procedures, use the emotion engine to analyze the user's emotions and create a detailed operating manual."

[0812] The above describes specific embodiments for carrying out the present invention. By utilizing the present invention, it becomes possible to enrich the user's operational experience and provide valuable feedback that utilizes emotional information.

[0813] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0814] Step 1:

[0815] Initial setup

[0816] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant permission for microphone and screen capture. When the user clicks "Allow," the terminal confirms and saves these settings.

[0817] Input: User actions (system startup, permission granting)

[0818] Data processing: Permission verification and saving of settings

[0819] Output: Microphone and screen capture ready.

[0820] Step 2:

[0821] Screen recording

[0822] The user clicks the "Start Recording" button. The device begins recording the entire current screen and saves all user actions as an MP4 video file. FFmpeg is used as the recording tool.

[0823] Input: User action (click of the "Start Recording" button)

[0824] Data processing: Screen recording

[0825] Output: MP4 video file (recording.mp4)

[0826] Step 3:

[0827] Audio recording

[0828] The device simultaneously starts microphone input and saves the user's voice as a WAV audio file. Pyaudio is used as the audio library.

[0829] Input: User's voice

[0830] Data processing: Audio recording

[0831] Output: WAV format audio file (audio.wav)

[0832] Step 4:

[0833] Emotion analysis

[0834] The device uses OpenCV to analyze camera input, analyzes the user's emotions in real time from their voice and facial expressions, and saves the emotion data in JSON format.

[0835] Input: User's facial expressions, voice

[0836] Data processing: Real-time analysis of emotions

[0837] Output: Emotion data file in JSON format (emotions.json)

[0838] Step 5:

[0839] Sending data

[0840] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The generated video file, audio file, and emotion data are saved to a temporary directory. The device uses the Python requests library to send these files to the server via an HTTP POST request.

[0841] Input: User action (click of the "Stop Recording" button), generated file

[0842] Data processing: Saving and sending files

[0843] Output: File sent to the server

[0844] Step 6:

[0845] Converting audio data to text

[0846] The server sends the received audio file to the Google Cloud Speech-to-Text API, where it converts the audio data into text.

[0847] Input: Audio file (audio.wav)

[0848] Data processing: Text conversion using speech recognition.

[0849] Output: Text data

[0850] Step 7:

[0851] Analysis of emotional data

[0852] The server analyzes the received emotion data file and stores each emotion data item in the database along with its corresponding timestamp.

[0853] Input: Emotion data file (emotions.json)

[0854] Data processing: Analysis of emotional data and its correspondence with timestamps.

[0855] Output: Emotional data stored in the database

[0856] Step 8:

[0857] Generating an operation manual

[0858] The server integrates the received text data and timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. It generates the manual in PDF format using LaTeX, and also in HTML format.

[0859] Input: Text data, timestamp information, video files

[0860] Data processing: Capture and integrate screenshots, generate consistent operation manuals.

[0861] Output: Operation manual in PDF and HTML formats

[0862] Step 9:

[0863] Provision of operation manual

[0864] The server generates a download link for the operation manual using Flask and provides it to the user via their dashboard and email. Users can download and view the operation manual by clicking the download link.

[0865] Input: Generated operation manual

[0866] Data processing: Creating download links

[0867] Output: Providing download links to users

[0868] (Application Example 2)

[0869] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0870] Conventional user manual generation systems lacked the ability to consider user emotions, resulting in a failure to integrate emotional information into operating procedures and explanations. Especially in online learning platforms, there is a demand for learning content that accurately reflects the instructor's emotional state and explanations. Therefore, an intuitive user manual generation system with emotion analysis capabilities is needed to deepen users' understanding of the operations.

[0871] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting emotion data, means for converting audio files into text, and means for generating an operation manual based on the converted text, video files, and emotion data. This makes it possible to provide an operation manual that integrates user operations and explanations with the emotion information at that time.

[0872] "User" refers to the entity that operates this system and uses it to record screen operations and audio.

[0873] "Screen images" refer to the visual information displayed on the monitor of a computer operated by a user.

[0874] "Means of recording" refers to a device or software that captures and records video of the screen being operated by the user.

[0875] A "microphone" refers to a device that converts sound into electrical signals and records the user's voice.

[0876] "Means of recording audio" refers to a device or software that uses a microphone to record the user's voice.

[0877] "Means for analyzing user facial expressions" refers to devices or software that capture the user's facial movements and expressions, analyze them, and obtain emotional data.

[0878] "Emotional data" refers to emotional information analyzed from the user's facial expressions and voice.

[0879] A "server" refers to a computer system used to store, process, and generate operation manuals for recorded data.

[0880] A "video file" refers to a digital file containing recorded data of the user's screen.

[0881] An "audio file" refers to a digital file containing user voice data recorded via a microphone.

[0882] "Speech recognition technology" refers to algorithms and techniques for converting speech data into text data.

[0883] An "operation manual" refers to a set of instructions or manuals that include user actions, explanations, and emotional data.

[0884] "Means of provision" refers to devices or software for distributing or displaying the generated operation manual to the user.

[0885] "PDF format" is an abbreviation for Portable Document Format, and refers to a digital file format that allows users to view and print operation manuals.

[0886] "HTML format" is an abbreviation for HyperText Markup Language, and refers to a file format used to display instruction manuals on the web.

[0887] System Overview

[0888] Initial setup

[0889] The user starts the system and accesses the initial setup screen. At this point, the terminal displays a dialog box prompting the user to grant permission for microphone and screen capture. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[0890] Screen operation and audio recording

[0891] When the user clicks the "Start Recording" button, the device begins recording the entire current screen and saves all operations as a video file. Simultaneously, the device records the user's voice explanation via the microphone. In addition, the device captures the user's facial expressions and analyzes and collects emotion data in real time.

[0892] Sending data

[0893] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The generated video file, audio file, and emotion data are then saved to a temporary directory and sent to the server. The transmission is performed using an HTTP POST request, and the video file, audio file, and emotion data are uploaded to the server as form data.

[0894] Data processing

[0895] The server receives uploaded video files, audio files, and sentiment data. After receiving the data, the server passes the audio files to a speech recognition engine to begin the process of converting them to text. Specifically, speech recognition technology is used to convert the audio data into text data. The server also analyzes the sentiment data and records it along with the corresponding timestamp.

[0896] Generating an operation manual

[0897] The server integrates text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. The generated operation manual also displays the user's emotions at each operation step.

[0898] Distribution of manuals

[0899] The completed user manual is generated in PDF or HTML format. In addition to the operating procedures, the generated manual visually displays sentiment data, clearly showing what emotions the user experienced during each operation. The server generates a download link to provide this manual to the user. This link is sent to the system dashboard or to the user's email address. The user can download and view the user manual by clicking the provided download link.

[0900] Program Processing Overview

[0901] Hardware and software to be used

[0902] Hardware:

[0903] Webcam: Used for video and sentiment analysis.

[0904] Microphone: Used for voice recording

[0905] software:

[0906] OpenCV: Used for video capture and saving

[0907] pyaudio: Used for audio recording

[0908] wave: For saving audio files

[0909] Request: For sending files via HTTP POST request.

[0910] EmotionRecognizer: For emotion analysis

[0911] speech_recognition: For speech recognition

[0912] Processing details

[0913] The terminal simultaneously performs video and audio recording and emotional data analysis. The recorded video files, audio files, and emotional data are sent to the server. The server converts the audio data to text and integrates the emotional data with timestamps to generate an operation manual. Finally, the generated manual is distributed to the user in PDF or HTML format.

[0914] Specific example

[0915] If the system is started: "start_recording('Lesson Start', 'Audio Explanation')"

[0916] When the operation is finished: "analyze_emotions('lesson_audio.wav', 'lesson_video.avi')"

[0917] Generating a lesson operation manual: "generate_manual('generated text', 'emotion data', 'lesson_manual.txt')"

[0918] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0919] Step 1:

[0920] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant permission for microphone and screen capture. The input is the user's actions, and the output is the state of readiness for completing the microphone and screen capture permission settings. Specifically, the terminal verifies the user's permission and configures the microphone and screen capture settings.

[0921] Step 2:

[0922] When the user clicks the "Start Recording" button, the device begins recording the entire current screen. The input is the click of the "Start Recording" button, and the output is the start of recording the video file. Specifically, the device uses OpenCV to capture and save the screen video.

[0923] Step 3:

[0924] The device simultaneously records the user's voice description via the microphone. The input is the user's voice, and the output is an audio file. Specifically, the device uses pyaudio to record the audio and saves it using the wave library.

[0925] Step 4:

[0926] The device captures the user's facial expressions and analyzes and collects emotional data in real time. The input is a video of the user's face, and the output is emotional data. Specifically, the device uses its EmotionRecognizer to analyze facial expressions and saves them as emotional data.

[0927] Step 5:

[0928] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The input is the click of the "Stop Recording" button, and the output is the completion of collecting video and audio files, as well as emotion data. Specifically, it stops recording video and audio and saves the data to a temporary directory.

[0929] Step 6:

[0930] The device sends recorded video files, recorded audio files, and emotion data to the server. The inputs are video files, audio files, and emotion data, and the output is the uploading of this data to the server. Specifically, the device uses the requests library to send an HTTP POST request and upload the data to the server.

[0931] Step 7:

[0932] The server receives uploaded video files, audio files, and emotion data. The input is the transmitted data, and the output is the completion of data reception on the server side. Specifically, the server saves the files to storage.

[0933] Step 8:

[0934] The server passes the audio file to the speech recognition engine and begins the process of converting it to text. The input is an audio file, and the output is text data. Specifically, the server uses the speech_recognition library to convert the audio to text.

[0935] Step 9:

[0936] The server analyzes sentiment data and records it along with a timestamp. The input is sentiment data, and the output is the sentiment analysis result with a timestamp. Specifically, the server analyzes the sentiment data and integrates it with the corresponding time information.

[0937] Step 10:

[0938] The server generates operation manuals based on text, video files, and sentiment data. Inputs are text data, screenshots, and sentiment data, while output is the operation manual. Specifically, the server organizes this data into a consistent set of operating procedures and creates the manual in PDF or HTML format.

[0939] Step 11:

[0940] The server provides the user with a completed operation manual. The input is the generated operation manual, and the output is a download link. Specifically, the server creates a download link and sends it to the system dashboard or the user's email address.

[0941] Example of a prompt

[0942] If the system is started: "start_recording('Lesson Start', 'Audio Explanation')"

[0943] When the operation is finished: "analyze_emotions('lesson_audio.wav', 'lesson_video.avi')"

[0944] Generating a lesson operation manual: "generate_manual('generated text', 'emotion data', 'lesson_manual.txt')"

[0945] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0946] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0947] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0948] [Third Embodiment]

[0949] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0950] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0951] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0952] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0953] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0954] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0955] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0956] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0957] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0958] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0959] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0960] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0961] This invention is a system that instantly records the computer screen operations performed by a user and their explanations, and later automatically generates an operation manual. The embodiments thereof are described in detail below.

[0962] System Overview

[0963] 1. Initial Setup

[0964] The user starts the system and grants permission for microphone and screen capture.

[0965] The device checks the microphone and screen capture settings and completes the setup.

[0966] 2. Recording screen operations and audio.

[0967] Recording begins when the user clicks the "Start Recording" button.

[0968] The device records all actions performed on the screen as video files.

[0969] At the same time, the device records the user's voice description through the microphone.

[0970] As a concrete example, consider a scenario where a user clicks "New File" and says aloud, "I have created a new file." This action and the accompanying audio are recorded simultaneously.

[0971] 3. Sending data

[0972] Once all recording and recording are complete, the user clicks the "Stop Recording" button.

[0973] The terminal sends the generated video and audio files to the server. This transmission takes place over the internet.

[0974] 4. Data Processing

[0975] The server analyzes the received video and audio files.

[0976] The server uses speech recognition technology to convert audio files into text data. For example, the audio file "A new file has been created" is converted to text "A new file has been created".

[0977] 5. Generation of the operation manual

[0978] The server combines the text with the corresponding timestamp information and integrates it with a specific portion (screenshot) of the video file.

[0979] Specifically, text and screenshots are compiled into a consistent manual as a series of steps.

[0980] 6. Distribution of manuals

[0981] The completed operation manual will be generated in PDF or HTML format.

[0982] The server will provide this manual to users by displaying a download link via email or on the dashboard.

[0983] Users can download and view the manual.

[0984] Specific example

[0985] 1. A scene where the user clicks the "Create New Project" button and the process is explained verbally.

[0986] User: "I'll create a new project here."

[0987] The device records audio simultaneously with operation, along with a timestamp.

[0988] 2. Data transmission and processing after recording stops

[0989] The device sends the recorded data to the server.

[0990] The server uses speech recognition to extract the text "We will create a new project here" and matches it with the video of the operation.

[0991] 3. Example of manual generation

[0992] The server organizes the steps on the timeline and generates a page for the "1. Create a new project" section, combining corresponding screenshots and text.

[0993] In this way, by using the system of the present invention, an efficient and accurate operation manual is automatically generated based on the user's operations and explanations. The implementation of this invention significantly reduces the effort required to create manuals and enables the provision of highly accurate documentation.

[0994] The following describes the processing flow.

[0995] Step 1:

[0996] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[0997] Step 2:

[0998] The user clicks the "Start Recording" button. This causes the device to begin recording the entire current screen, saving all actions as a video file. Simultaneously, the device records the user's voice via the microphone. Specifically, the screen capture and audio recording processes are executed at the same time.

[0999] Step 3:

[1000] As the user performs screen operations, the system provides real-time audio explanations of those operations. For example, when the user clicks "New File" and selects it from the "File" menu, the system will provide the audio explanation, "Now you will create a new file." The device records these operations and audio, along with a timestamp.

[1001] Step 4:

[1002] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and saves the generated video and audio files to a temporary directory.

[1003] Step 5:

[1004] The terminal sends video and audio files stored in a temporary directory to the server. The files are sent using an HTTP POST request. Specifically, the video and audio files are uploaded to the server as form data.

[1005] Step 6:

[1006] The server receives the uploaded video and audio files. After receiving them, it passes the audio files to a speech recognition engine to begin the process of converting them to text. Here, speech recognition technology (e.g., a speech recognition API) is used to convert the audio data into text data.

[1007] Step 7:

[1008] The server analyzes the video file and generates frame-by-frame screenshots. Based on the video timeline, it extracts specific action events (e.g., button clicks or form inputs) and takes screenshots of the moments when these events occur.

[1009] Step 8:

[1010] The server integrates the converted text and screenshots to generate an operation manual. Based on the text and timestamp information, the screenshots are arranged as a sequence of steps to create a consistent operation manual. The generated operation manual is in PDF or HTML format.

[1011] Step 9:

[1012] The server generates a download link to provide the user with the completed operation manual. This link is sent to the system dashboard or to the user's email address.

[1013] Step 10:

[1014] Users can download and view the operation manual by clicking the provided download link. This operation manual is available for reference as needed and supports the use of the system.

[1015] (Example 1)

[1016] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1017] Traditional methods for creating user manuals require users to manually record each operation step and then transcribe the information, which is time-consuming and labor-intensive. Furthermore, the operation and its explanation may not always be synchronized, potentially leading to the creation of inaccurate manuals. Therefore, there is a need for a system that allows users to efficiently and accurately generate user manuals without having to spend time on incidental tasks.

[1018] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1019] In this invention, the server includes means for recording video of the screen operated by the user, means for recording the user's voice through a microphone, means for transmitting the recorded video file and the recorded audio file from the device to an information processing device, means for converting the audio file into text data on the information processing device, means for generating an operation manual based on the converted text data and video file, and means for providing the generated operation manual to the user. This makes it possible for the user to generate an operation manual simply and efficiently.

[1020] A "user" refers to anyone who intends to use this system to generate an operation manual.

[1021] "Screen images" refers to all visual information displayed on the screen of a device operated by the user.

[1022] A "microphone" is an audio input device used to record the user's voice.

[1023] A "recorded video file" is a digital file created by recording the video of the screen being operated by the user.

[1024] "Recorded audio files" refer to digital files generated by recording the user's voice.

[1025] A "device" refers to electronic devices such as computers and smartphones that are operated by the user.

[1026] An "information processing device" refers to a server or cloud computer used to process video and audio files transmitted from a device.

[1027] "Speech recognition technology" is a technology that converts audio files into text data; it analyzes audio data and turns it into text.

[1028] "Text data" refers to text information converted by speech recognition technology.

[1029] An "operation manual" refers to a document that summarizes the user's operating procedures and their explanations, and is generated based on video files and text data.

[1030] "Electronic document format" refers to digital document formats such as PDF and HTML.

[1031] "Web page format" refers to a web-based document that can be viewed on an internet browser.

[1032] This invention provides a system that instantly records user-performed computer screen operations and their descriptions, and automatically generates an operation manual later. This system allows users to efficiently and accurately generate operation manuals.

[1033] Hardware and software to be used

[1034] Terminal: A device used by the user to perform operations (e.g., personal computer, smartphone)

[1035] Microphone: A voice input device for recording the user's voice.

[1036] Screen capture tool: Software that records screen activity (e.g., OBS Studio)

[1037] Server: An information processing device that performs data processing and manual generation (e.g., cloud computing service).

[1038] Speech recognition technology: A technology that converts speech data into text data (e.g., Google Speech-to-Text API).

[1039] Image processing software: Technology for automatically generating screenshots from video (e.g., OpenCV)

[1040] HTML template engine: A technology for generating operation manuals in HTML format (e.g., Jinja2).

[1041] PDF generation software: Technology for generating manuals in PDF format (e.g., wkhtmltopdf).

[1042] Process Overview

[1043] 1. The user starts the system and grants permission for microphone and screen capture.

[1044] 2. The device checks the screen capture and microphone settings and completes the setup.

[1045] 3. When the user clicks the "Start Recording" button, the device simultaneously records the video and audio of the screen. For example, the user clicks "New File" and says aloud, "A new file has been created."

[1046] 4. When the user clicks the "Stop Recording" button, the device sends the recorded video and audio files to the server.

[1047] 5. The server analyzes the received video and audio files and uses speech recognition technology to convert the audio files into text data. For example, the audio data "A new file has been created" is converted into the text data "A new file has been created".

[1048] 6. The server integrates text data and timestamp information with video files to generate a consistent operation manual. The operation manual is output in HTML or PDF format using an HTML template engine or PDF generation software.

[1049] 7. The server will provide users with a download link for the generated operation manual via email or on the dashboard. Users can download and view the operation manual from this link.

[1050] Specific example

[1051] A scene where the user clicks the "Create New Project" button and the process is explained via voice.

[1052] User: "I'll create a new project here."

[1053] The device records audio simultaneously with this operation, along with a timestamp.

[1054] Data transmission and processing after recording stops

[1055] The device sends the recorded data to the server.

[1056] The server uses speech recognition to extract the text "We will create a new project here" and matches it with the video of the operation.

[1057] Manual generation example

[1058] The server organizes the steps on the timeline and generates a page for the step "1. Create a new project," combining corresponding screenshots and text.

[1059] Examples of prompt statements

[1060] 1. "Please record the steps to create a new project using both screen operations and audio."

[1061] 2. "Please continue by explaining the current screen operation aloud."

[1062] 3. "Once recording is complete, please click the 'End' button to send the data to the server."

[1063] The above details describe the embodiments of the present invention. This system enables users to generate highly accurate operation manuals in a short amount of time.

[1064] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1065] Step 1:

[1066] The user starts the system and grants permission for microphone and screen capture.

[1067] Input: Command to start the system, permission request popup.

[1068] Output: Permission granted

[1069] Specific action: When the user starts the system, a pop-up appears asking for permission to use the microphone and screen capture. The user clicks "Allow".

[1070] Step 2:

[1071] The device checks the microphone and screen capture settings and completes the setup.

[1072] Input: User-granted permissions

[1073] Output: Microphone and screen capture ready to use status

[1074] Specific operation: The device automatically checks the settings for the microphone and screen capture tool (e.g., OBS Studio) and displays a "Ready" message.

[1075] Step 3:

[1076] The user clicks the "Start Recording" button.

[1077] Input: User click of the "Start Recording" button.

[1078] Output: Trigger for recording start

[1079] Specific action: The user clicks the "Start Recording" button in the application window.

[1080] Step 4:

[1081] The device records all actions performed on the screen.

[1082] Input: Trigger for recording start

[1083] Output: Video data being recorded

[1084] Specific action: The OBS Studio software begins recording the user's screen activity. For example, it records the user clicking the "Create New Project" button.

[1085] Step 5:

[1086] The device records the user's voice description via the microphone.

[1087] Input: Trigger for recording start, user's voice

[1088] Output: Audio data being recorded

[1089] Specific actions: As the user performs an action, the system provides a voice explanation of that action. For example, it might say, "Here, we will create a new project."

[1090] Step 6:

[1091] The user clicks the "Stop Recording" button.

[1092] Input: User click of the "Stop Recording" button.

[1093] Output: Trigger for recording termination

[1094] Specific action: The user clicks the "Stop Recording" button after completing the operation.

[1095] Step 7:

[1096] The device sends the recorded video file and the recorded audio file to the server.

[1097] Input: Trigger for recording termination, video file, audio file

[1098] Output: Status of successful file transfer to server

[1099] Specific operation: The device uploads video and audio files to the server via the internet connection. For example, it sends files using an API and displays a "transmission complete" notification.

[1100] Step 8:

[1101] The server analyzes the received video and audio files.

[1102] Input: Video file, audio file

[1103] Output: Analyzed audio data, video data

[1104] Specific operation: The server temporarily stores the received data and separates the video file from the audio file.

[1105] Step 9:

[1106] The server converts the audio file into text data.

[1107] Input: Audio file

[1108] Output: Character data (text file)

[1109] Specific operation: Use the Google Speech-to-Text API to convert audio data into text data. For example, the audio "I have created a new file" will be converted into text.

[1110] Step 10:

[1111] The server integrates text and timestamp information with the video file.

[1112] Input: Text data, timestamp, video file

[1113] Output: Operation procedure data, screenshots

[1114] Specific operation: Use OpenCV to obtain appropriate screenshots from video, and generate a consistent instruction manual based on text data and timestamp information.

[1115] Step 11:

[1116] The server generates an operation manual.

[1117] Input: Operation procedure data, screenshots

[1118] Output: Operation manual (PDF or HTML format)

[1119] Specific operation: Use the Jinja2 template engine to generate an operation manual in HTML format, and then convert it to PDF format using wkhtmltopdf.

[1120] Step 12:

[1121] The server provides users with an operation manual.

[1122] Input: Generated operation manual

[1123] Output: Download link, notification email

[1124] Specific operation: The server saves the generated manual and provides a download link to the user. The user can download the manual by displaying the link in an email or on the dashboard.

[1125] (Application Example 1)

[1126] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1127] Creating manuals for operating robots and machinery within a factory is time-consuming and labor-intensive, and accurately conveying advanced operations and complex procedures is difficult. Furthermore, training new workers is lengthy and prone to errors, highlighting the need for improved production efficiency and safety.

[1128] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1129] In this invention, the server includes means for recording video of the screen operated by the user, means for recording the user's voice through a microphone, means for using a visual device to record video of the real environment, means for transmitting the recorded video file and recorded audio file to the server, means for converting the audio file to text on the server, means for generating an operation manual based on the converted text and video file, and means for providing the generated operation manual to the user. This enables accurate and rapid recording of complex procedures and operations within the factory, efficient automatic generation of operation manuals, reduction of worker training time, and prevention of errors.

[1130] "Means for recording video of the user's screen" refers to a function that records all operations performed on the screen of a computer or device being operated by the user in video format.

[1131] "Means of recording user voice via microphone" refers to a function that uses a microphone to record the voice of the user when they give explanations or instructions while operating the device.

[1132] "Means of using visual devices to record images of the real environment" refers to a function that allows users to record images of their surroundings or the object they are operating in real time using devices such as smart glasses or cameras.

[1133] "Means for sending recorded video files and recorded audio files to a server" refers to a function for uploading recorded video data and audio data to a server via the internet or other means.

[1134] "Method for converting audio files to text on the server" refers to a function that uses speech recognition technology on the server side to automatically convert transmitted audio data into text format.

[1135] "A means of generating operation manuals based on converted text and video files" refers to a function that automatically creates consistent operation procedures and manuals by combining audio-to-text data and recorded video data.

[1136] "Means of providing the generated operation manual to the user" refers to a function that outputs the completed operation manual in PDF or HTML format, allowing the user to download or view it.

[1137] This invention is a system that records robot operation and maintenance procedures in a factory in real time and automatically generates an operation manual based on that data. The embodiments of this system are described in detail below.

[1138] System Overview

[1139] 1. Initial Setup

[1140] The user starts the system and grants permission for microphone and screen capture. Additionally, the user wears a visual device (e.g., smart glasses or a camera) and confirms the settings.

[1141] The device checks the settings for microphone, screen capture, and visual devices, and then completes the setup.

[1142] 2. Recording screen operations and audio.

[1143] Recording begins when the user clicks the "Start Recording" button.

[1144] The device records the user's actions on the device screen as a video file, and simultaneously records video of the real environment through a visual device.

[1145] At the same time, the device records the user's voice description through the microphone.

[1146] 3. Sending data

[1147] Once all recording and recording are complete, the user clicks the "Stop Recording" button.

[1148] The terminal sends the generated video and audio files to the server. This transmission takes place over the internet.

[1149] 4. Data Processing

[1150] The server analyzes the received video and audio files and uses speech recognition technology to convert the audio files into text data. For example, the audio data "I have created a new project" is converted to text "I have created a new project".

[1151] 5. Generation of the operation manual

[1152] The server combines the text with the corresponding timestamp information and integrates it with a specific portion (screenshot) of the video file.

[1153] Specifically, text and screenshots are compiled into a consistent manual as a series of steps.

[1154] 6. Distribution of manuals

[1155] The completed operation manual will be generated in PDF or HTML format.

[1156] The server will provide this manual to users by displaying a download link via email or on the dashboard.

[1157] Users can download and view the manual.

[1158] Specific example

[1159] Consider a scenario where factory workers wear smart glasses to perform robot maintenance procedures. For example, they might record the procedure for replacing a hydraulic cylinder.

[1160] The user performs the task while explaining verbally, "Here, we will remove the hydraulic cylinder," and both the operation and the audio are recorded simultaneously.

[1161] After recording stops, the terminal sends the recorded data to the server, converts the audio data into text, and generates a manual.

[1162] Example of a prompt

[1163] "Use video and audio capture to record robot operation procedures within the factory. After recording is complete, convert the audio data into text and generate an operation manual along with the video frames. For example, if you perform a task while explaining, 'Here we remove the hydraulic cylinder,' ensure that this procedure and explanation are integrated into a single procedure in the manual."

[1164] In this way, by using the system of the present invention, complex procedures and operations within a factory can be accurately and quickly recorded, and efficient operation manuals can be automatically generated, thereby reducing worker training time and preventing operational errors.

[1165] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1166] Step 1:

[1167] Initial settings:

[1168] The user starts the system and grants permissions for the microphone, screen capture, and visual devices (3D glasses or camera). The terminal verifies that these devices are working correctly and completes the setup. The input is the permission and setup required for each device, and the output is the result of verifying that the devices are ready for use.

[1169] Step 2:

[1170] Start recording screen operations and audio:

[1171] The recording process begins when the user clicks the "Start Recording" button. The device records all operations performed on the computer screen as a video file using screen capture software (e.g., PyAutoGUI). Simultaneously, audio recording software (e.g., PyAudio) records the user's voice explanations via the microphone. The visual device records the user's viewpoint in real time. Input data consists of user operations, audio, and visual environment, and the output is a recorded file.

[1172] Step 3:

[1173] Sending data:

[1174] Once recording is complete, the user clicks the "Stop Recording" button. The device then sends the recorded video files (screen captures and video from the visual device) and audio files to the server. The transmission takes place over the internet. The input data consists of the recorded video and audio files, and the output is the data uploaded to the server.

[1175] Step 4:

[1176] Data processing:

[1177] The server analyzes the received video and audio files. Audio files are converted to text using speech recognition technology (e.g., Google Cloud Speech-to-Text API). The input is an audio file, and the output is the corresponding text data. Video files are also organized based on timestamps.

[1178] Step 5:

[1179] Generating the operation manual:

[1180] The server combines the converted text data with corresponding timestamp information and integrates it with specific portions (screenshots) of the video file. This generates a consistent operation manual. The input is text data and video files, and the output is an integrated operation manual file.

[1181] Step 6:

[1182] Distribution of manuals:

[1183] The server generates a completed operation manual in PDF or HTML format and makes it available for users to download. The download link is provided via email or on the dashboard. The input is the generated operation manual, and the output is the download link that users can access.

[1184] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1185] This invention provides a system that instantly records user computer screen operations and their explanations, and later automatically generates an operation manual. Furthermore, by combining it with an emotion engine, it enables the recognition and utilization of user emotions. The embodiments thereof will be described in detail below.

[1186] System Overview

[1187] 1. Initial Setup

[1188] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[1189] 2. Recording screen operations and audio.

[1190] The user clicks the "Start Recording" button. This causes the device to begin recording the entire current screen and save all actions as a video file. Simultaneously, the device records the user's voice explanation via the microphone.

[1191] The device activates an emotion engine, which analyzes the user's emotions in real time from their voice and facial expressions, and collects emotional data.

[1192] 3. Sending data

[1193] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and save the generated video and audio files, as well as emotion data, to a temporary directory.

[1194] The device sends these files to the server. The transmission is done using an HTTP POST request, and video files, audio files, and emotion data are uploaded to the server as form data.

[1195] 4. Data Processing

[1196] The server receives uploaded video files, audio files, and emotion data. After receiving the data, it passes the audio files to a speech recognition engine to begin the process of converting them to text. Specifically, speech recognition technology is used to convert the audio data into text data.

[1197] The server also analyzes emotional data and records it along with the corresponding timestamp.

[1198] 5. Generation of the operation manual

[1199] The server integrates text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual.

[1200] The generated operation manual also displays the user's emotions at each operation step. For example, along with a description such as "You have created a new file," it also shows the emotions the user felt at that moment (e.g., joy or confusion).

[1201] 6. Distribution of manuals

[1202] The completed user manual will be generated in PDF or HTML format. In addition to the operating procedures, the generated manual will visually display emotional data, making it clear what emotions the user felt during each operation.

[1203] The server generates a download link to provide this manual to the user. This link is sent to the system dashboard or to the user's email address.

[1204] Users can download and view the operation manual by clicking the provided download link. This operation manual is available for reference as needed and supports the use of the system.

[1205] Specific example

[1206] 1. The user clicks the "Create New Project" button.

[1207] User: "I'll create a new project here."

[1208] The device records audio simultaneously with operation, along with a timestamp.

[1209] At the same time, the emotion engine analyzes the user's emotions and collects emotional data such as "joy" and "anticipation."

[1210] 2. Data transmission and processing after recording stops

[1211] The terminal sends the recorded data to the server. The server uses speech recognition to extract the text "Here we will create a new project" and matches it with the video of the operation.

[1212] Emotional data is also analyzed and recorded along with the corresponding timestamp.

[1213] 3. Example of manual generation

[1214] The server organizes the steps on the timeline and generates a page for the "Create a new project" section, combining the corresponding screenshots, text, and sentiment data at that time.

[1215] The manual displays an emotion such as "Joy" below an explanation like, "Here we will create a new project."

[1216] In this way, the system of the present invention can generate a detailed and intuitive operation manual, including emotional data, based on the user's operations and explanations. This enriches the user's experience and allows for the provision of feedback utilizing emotional information.

[1217] The following describes the processing flow.

[1218] Step 1:

[1219] The user starts the system and accesses the initial setup screen. The terminal displays a dialog box requesting permission for microphone and screen capture. Once the user grants these permissions, the terminal confirms the settings and completes the setup.

[1220] Step 2:

[1221] The user clicks the "Start Recording" button. This causes the device to begin recording the entire screen and save the operation as a video file. Simultaneously, the device records the user's voice through the microphone. The device also activates an emotion engine to analyze the user's emotions in real time from their facial expressions and voice.

[1222] Step 3:

[1223] As the user performs screen operations, the system provides real-time audio explanations of those operations. For example, when the user clicks "New File" and selects it from the "File" menu, the system will say, "Now you will create a new file." The device records these operations and audio along with timestamps. Simultaneously, it analyzes and records sentiment data along with timestamps.

[1224] Step 4:

[1225] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and save the generated video and audio files, as well as emotion data, to a temporary directory.

[1226] Step 5:

[1227] The device sends video files, audio files, and emotion data stored in a temporary directory to the server. The files are sent using an HTTP POST request, and the video files, audio files, and emotion data are uploaded to the server as form data.

[1228] Step 6:

[1229] The server receives the uploaded video files, audio files, and emotion data and begins processing them. First, the audio files are passed to the speech recognition engine to be converted into text. In this process, for example, audio data such as "Here I will create a new file" is converted into text data.

[1230] Step 7:

[1231] The server analyzes the video file and generates frame-by-frame screenshots on the timeline. Based on the video's timestamps, it extracts specific action events (e.g., button clicks or form inputs) and takes screenshots of the moments when these events occurred.

[1232] Step 8:

[1233] The server integrates text, screenshots, and sentiment data to generate an operation manual. The manual sequentially displays text and corresponding screenshots as operating procedures, and also shows the user's sentiment data at each step. For example, under the explanation "Create a new file," the corresponding screenshot and the sentiment data "Joy" are displayed.

[1234] Step 9:

[1235] The server saves the generated operation manual in PDF or HTML format and generates a download link for the user. This link is sent to the system dashboard or to the user's email address.

[1236] Step 10:

[1237] Users can download and view the operation manual by clicking the provided download link. This operation manual visually presents user sentiment data along with the operating procedures, making it easier to understand intuitively.

[1238] (Example 2)

[1239] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1240] Conventional operation manual generation systems could record user instructions and screen operations, but they had the problem of not being able to obtain feedback that included the user's emotions. As a result, feedback based on the user's emotions was not provided, leading to a lack of deeper understanding of how to operate the system.

[1241] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting audio files to text, means for generating an operation manual based on the converted text and video files, means for analyzing the user's emotions in real time and collecting emotion data, and means for integrating the collected emotion data into the operation manual. This makes it possible to generate a more detailed and intuitive operation manual that reflects not only the user's operation instructions but also their emotional information at that time.

[1242] A "user" is the entity that operates the system and performs various inputs.

[1243] "Means of recording screen footage" refers to functions or devices that record the user's screen in video format.

[1244] "Means of recording audio via a microphone" refers to functions or devices for recording the voice spoken by a user.

[1245] "Means of sending to a server" refers to the technology and processes used to upload recorded video files and audio files to a server via a network.

[1246] "Methods for converting audio files to text" refers to the process of converting audio data into text data using speech recognition technology.

[1247] "Means for generating operation manuals" refers to technologies and processes for creating documents that visually show user operation procedures based on acquired audio text and video files.

[1248] "Means of providing operation manuals to users" refers to the processes and functions for providing the generated operation manuals in a user-accessible format (e.g., PDF or HTML).

[1249] "Means of analyzing emotions in real time and collecting emotional data" refers to technologies and devices that analyze a user's facial expressions and voice to identify and record the emotions the user is feeling in real time.

[1250] "Means of integrating emotional data into the operation manual" refers to the techniques and processes used to link collected emotional data to each operation step in the operation manual and ensure consistency.

[1251] The system according to the present invention can instantly record the computer screen operations performed by the user and their explanations, and further recognize, analyze, and utilize the user's emotions by combining them with an emotion engine. Specific embodiments of the present invention will be described in detail below.

[1252] Overall overview

[1253] This system is implemented through a process that records the user's screen operations and their voice explanations via a microphone, sends this data to a server for analysis, and ultimately generates an operation manual. Furthermore, it analyzes the user's emotions using an emotion engine and incorporates this analysis into the operation manual.

[1254] Initial setup

[1255] When the user starts the system, the terminal displays an initial setup screen and a dialog box prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the preparation for recording and audio.

[1256] Screen operation and audio recording

[1257] When the user clicks the "Start Recording" button, the device begins recording the entire current screen and saves all operations as a video file. Simultaneously, the device records the user's voice description via the microphone. The device also activates an emotion engine and uses OpenCV to analyze the user's voice and facial expressions in real time, collecting emotion data.

[1258] Sending data

[1259] When the user clicks the "Stop Recording" button, the device stops recording video and audio, and saves the generated video and audio files, as well as emotion data, to a temporary directory. The device then uses the Python requests library to send these files to the server via an HTTP POST request.

[1260] Data processing

[1261] The server receives uploaded video and audio files, as well as sentiment data. Audio files are converted to text using the Google Cloud Speech-to-Text API. Additionally, the server analyzes the sentiment data and records it along with the corresponding timestamp.

[1262] Generating an operation manual

[1263] The server integrates the converted text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. The manual also visually displays the user's emotions at each operation step. Specifically, it generates a PDF version of the manual using LaTeX and also provides it in HTML format.

[1264] Distribution of manuals

[1265] The completed operation manual is generated in PDF or HTML format, and the server uses Flask to generate download links. Users can download and view the operation manual by clicking the download link provided via the dashboard or their email address.

[1266] Specific examples and prompt statements

[1267] Specific example

[1268] The scenario envisions a user clicking the "Create New Project" button and being instructed, "Here, we will create a new project." The device simultaneously records the user's actions and voice, and uses an emotion engine to analyze the user's emotions, such as "joy" and "expectation." Subsequently, the server generates an operation manual based on the collected data and provides it to the user.

[1269] Example of a prompt

[1270] "Record the steps for creating a new project and generate a manual that includes emotional data along with the operating instructions."

[1271] "Based on the recorded operating procedures, use the emotion engine to analyze the user's emotions and create a detailed operating manual."

[1272] The above describes specific embodiments for carrying out the present invention. By utilizing the present invention, it becomes possible to enrich the user's operational experience and provide valuable feedback that utilizes emotional information.

[1273] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1274] Step 1:

[1275] Initial setup

[1276] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant permission for microphone and screen capture. When the user clicks "Allow," the terminal confirms and saves these settings.

[1277] Input: User actions (system startup, permission granting)

[1278] Data processing: Permission verification and saving of settings

[1279] Output: Microphone and screen capture ready.

[1280] Step 2:

[1281] Screen recording

[1282] The user clicks the "Start Recording" button. The device begins recording the entire current screen and saves all user actions as an MP4 video file. FFmpeg is used as the recording tool.

[1283] Input: User action (click of the "Start Recording" button)

[1284] Data processing: Screen recording

[1285] Output: MP4 video file (recording.mp4)

[1286] Step 3:

[1287] Audio recording

[1288] The device simultaneously starts microphone input and saves the user's voice as a WAV audio file. Pyaudio is used as the audio library.

[1289] Input: User's voice

[1290] Data processing: Audio recording

[1291] Output: WAV format audio file (audio.wav)

[1292] Step 4:

[1293] Emotion analysis

[1294] The device uses OpenCV to analyze camera input, analyzes the user's emotions in real time from their voice and facial expressions, and saves the emotion data in JSON format.

[1295] Input: User's facial expressions, voice

[1296] Data processing: Real-time analysis of emotions

[1297] Output: Emotion data file in JSON format (emotions.json)

[1298] Step 5:

[1299] Sending data

[1300] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The generated video file, audio file, and emotion data are saved to a temporary directory. The device uses the Python requests library to send these files to the server via an HTTP POST request.

[1301] Input: User action (click of the "Stop Recording" button), generated file

[1302] Data processing: Saving and sending files

[1303] Output: File sent to the server

[1304] Step 6:

[1305] Converting audio data to text

[1306] The server sends the received audio file to the Google Cloud Speech-to-Text API, where it converts the audio data into text.

[1307] Input: Audio file (audio.wav)

[1308] Data processing: Text conversion using speech recognition.

[1309] Output: Text data

[1310] Step 7:

[1311] Analysis of emotional data

[1312] The server analyzes the received emotion data file and stores each emotion data item in the database along with its corresponding timestamp.

[1313] Input: Emotion data file (emotions.json)

[1314] Data processing: Analysis of emotional data and its correspondence with timestamps.

[1315] Output: Emotional data stored in the database

[1316] Step 8:

[1317] Generating an operation manual

[1318] The server integrates the received text data and timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. It generates the manual in PDF format using LaTeX, and also in HTML format.

[1319] Input: Text data, timestamp information, video files

[1320] Data processing: Capture and integrate screenshots, generate consistent operation manuals.

[1321] Output: Operation manual in PDF and HTML formats

[1322] Step 9:

[1323] Provision of operation manual

[1324] The server generates a download link for the operation manual using Flask and provides it to the user via their dashboard and email. Users can download and view the operation manual by clicking the download link.

[1325] Input: Generated operation manual

[1326] Data processing: Creating download links

[1327] Output: Providing download links to users

[1328] (Application Example 2)

[1329] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1330] Conventional user manual generation systems lacked the ability to consider user emotions, resulting in a failure to integrate emotional information into operating procedures and explanations. Especially in online learning platforms, there is a demand for learning content that accurately reflects the instructor's emotional state and explanations. Therefore, an intuitive user manual generation system with emotion analysis capabilities is needed to deepen users' understanding of the operations.

[1331] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting emotion data, means for converting audio files into text, and means for generating an operation manual based on the converted text, video files, and emotion data. This makes it possible to provide an operation manual that integrates user operations and explanations with the emotion information at that time.

[1332] "User" refers to the entity that operates this system and uses it to record screen operations and audio.

[1333] "Screen images" refer to the visual information displayed on the monitor of a computer operated by a user.

[1334] "Means of recording" refers to a device or software that captures and records video of the screen being operated by the user.

[1335] A "microphone" refers to a device that converts sound into electrical signals and records the user's voice.

[1336] "Means of recording audio" refers to a device or software that uses a microphone to record the user's voice.

[1337] "Means for analyzing user facial expressions" refers to devices or software that capture the user's facial movements and expressions, analyze them, and obtain emotional data.

[1338] "Emotional data" refers to emotional information analyzed from the user's facial expressions and voice.

[1339] A "server" refers to a computer system used to store, process, and generate operation manuals for recorded data.

[1340] A "video file" refers to a digital file containing recorded data of the user's screen.

[1341] An "audio file" refers to a digital file containing user voice data recorded via a microphone.

[1342] "Speech recognition technology" refers to algorithms and techniques for converting speech data into text data.

[1343] An "operation manual" refers to a set of instructions or manuals that include user actions, explanations, and emotional data.

[1344] "Means of provision" refers to devices or software for distributing or displaying the generated operation manual to the user.

[1345] "PDF format" is an abbreviation for Portable Document Format, and refers to a digital file format that allows users to view and print operation manuals.

[1346] "HTML format" is an abbreviation for HyperText Markup Language, and refers to a file format used to display instruction manuals on the web.

[1347] System Overview

[1348] Initial setup

[1349] The user starts the system and accesses the initial setup screen. At this point, the terminal displays a dialog box prompting the user to grant permission for microphone and screen capture. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[1350] Screen operation and audio recording

[1351] When the user clicks the "Start Recording" button, the device begins recording the entire current screen and saves all operations as a video file. Simultaneously, the device records the user's voice explanation via the microphone. In addition, the device captures the user's facial expressions and analyzes and collects emotion data in real time.

[1352] Sending data

[1353] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The generated video file, audio file, and emotion data are then saved to a temporary directory and sent to the server. The transmission is performed using an HTTP POST request, and the video file, audio file, and emotion data are uploaded to the server as form data.

[1354] Data processing

[1355] The server receives uploaded video files, audio files, and sentiment data. After receiving the data, the server passes the audio files to a speech recognition engine to begin the process of converting them to text. Specifically, speech recognition technology is used to convert the audio data into text data. The server also analyzes the sentiment data and records it along with the corresponding timestamp.

[1356] Generating an operation manual

[1357] The server integrates text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. The generated operation manual also displays the user's emotions at each operation step.

[1358] Distribution of manuals

[1359] The completed user manual is generated in PDF or HTML format. In addition to the operating procedures, the generated manual visually displays sentiment data, clearly showing what emotions the user experienced during each operation. The server generates a download link to provide this manual to the user. This link is sent to the system dashboard or to the user's email address. The user can download and view the user manual by clicking the provided download link.

[1360] Program Processing Overview

[1361] Hardware and software to be used

[1362] Hardware:

[1363] Webcam: Used for video and sentiment analysis.

[1364] Microphone: Used for voice recording

[1365] software:

[1366] OpenCV: Used for video capture and saving

[1367] pyaudio: Used for audio recording

[1368] wave: For saving audio files

[1369] Request: For sending files via HTTP POST request.

[1370] EmotionRecognizer: For emotion analysis

[1371] speech_recognition: For speech recognition

[1372] Processing details

[1373] The terminal simultaneously performs video and audio recording and emotional data analysis. The recorded video files, audio files, and emotional data are sent to the server. The server converts the audio data to text and integrates the emotional data with timestamps to generate an operation manual. Finally, the generated manual is distributed to the user in PDF or HTML format.

[1374] Specific example

[1375] If the system is started: "start_recording('Lesson Start', 'Audio Explanation')"

[1376] When the operation is finished: "analyze_emotions('lesson_audio.wav', 'lesson_video.avi')"

[1377] Generating a lesson operation manual: "generate_manual('generated text', 'emotion data', 'lesson_manual.txt')"

[1378] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1379] Step 1:

[1380] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant permission for microphone and screen capture. The input is the user's actions, and the output is the state of readiness for completing the microphone and screen capture permission settings. Specifically, the terminal verifies the user's permission and configures the microphone and screen capture settings.

[1381] Step 2:

[1382] When the user clicks the "Start Recording" button, the device begins recording the entire current screen. The input is the click of the "Start Recording" button, and the output is the start of recording the video file. Specifically, the device uses OpenCV to capture and save the screen video.

[1383] Step 3:

[1384] The device simultaneously records the user's voice description via the microphone. The input is the user's voice, and the output is an audio file. Specifically, the device uses pyaudio to record the audio and saves it using the wave library.

[1385] Step 4:

[1386] The device captures the user's facial expressions and analyzes and collects emotional data in real time. The input is a video of the user's face, and the output is emotional data. Specifically, the device uses its EmotionRecognizer to analyze facial expressions and saves them as emotional data.

[1387] Step 5:

[1388] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The input is the click of the "Stop Recording" button, and the output is the completion of collecting video and audio files, as well as emotion data. Specifically, it stops recording video and audio and saves the data to a temporary directory.

[1389] Step 6:

[1390] The device sends recorded video files, recorded audio files, and emotion data to the server. The inputs are video files, audio files, and emotion data, and the output is the uploading of this data to the server. Specifically, the device uses the requests library to send an HTTP POST request and upload the data to the server.

[1391] Step 7:

[1392] The server receives uploaded video files, audio files, and emotion data. The input is the transmitted data, and the output is the completion of data reception on the server side. Specifically, the server saves the files to storage.

[1393] Step 8:

[1394] The server passes the audio file to the speech recognition engine and begins the process of converting it to text. The input is an audio file, and the output is text data. Specifically, the server uses the speech_recognition library to convert the audio to text.

[1395] Step 9:

[1396] The server analyzes sentiment data and records it along with a timestamp. The input is sentiment data, and the output is the sentiment analysis result with a timestamp. Specifically, the server analyzes the sentiment data and integrates it with the corresponding time information.

[1397] Step 10:

[1398] The server generates operation manuals based on text, video files, and sentiment data. Inputs are text data, screenshots, and sentiment data, while output is the operation manual. Specifically, the server organizes this data into a consistent set of operating procedures and creates the manual in PDF or HTML format.

[1399] Step 11:

[1400] The server provides the user with a completed operation manual. The input is the generated operation manual, and the output is a download link. Specifically, the server creates a download link and sends it to the system dashboard or the user's email address.

[1401] Example of a prompt

[1402] If the system is started: "start_recording('Lesson Start', 'Audio Explanation')"

[1403] When the operation is finished: "analyze_emotions('lesson_audio.wav', 'lesson_video.avi')"

[1404] Generating a lesson operation manual: "generate_manual('generated text', 'emotion data', 'lesson_manual.txt')"

[1405] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1406] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1408] [Fourth Embodiment]

[1409] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1410] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1412] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1416] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1417] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1418] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1419] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1420] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1421] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1422] This invention is a system that instantly records the computer screen operations performed by a user and their explanations, and later automatically generates an operation manual. The embodiments thereof are described in detail below.

[1423] System Overview

[1424] 1. Initial Setup

[1425] The user starts the system and grants permission for microphone and screen capture.

[1426] The device checks the microphone and screen capture settings and completes the setup.

[1427] 2. Recording screen operations and audio.

[1428] Recording begins when the user clicks the "Start Recording" button.

[1429] The device records all actions performed on the screen as video files.

[1430] At the same time, the device records the user's voice description through the microphone.

[1431] As a concrete example, consider a scenario where a user clicks "New File" and says aloud, "I have created a new file." This action and the accompanying audio are recorded simultaneously.

[1432] 3. Sending data

[1433] Once all recording and recording are complete, the user clicks the "Stop Recording" button.

[1434] The terminal sends the generated video and audio files to the server. This transmission takes place over the internet.

[1435] 4. Data Processing

[1436] The server analyzes the received video and audio files.

[1437] The server uses speech recognition technology to convert audio files into text data. For example, the audio file "A new file has been created" is converted to text "A new file has been created".

[1438] 5. Generation of the operation manual

[1439] The server combines the text with the corresponding timestamp information and integrates it with a specific portion (screenshot) of the video file.

[1440] Specifically, text and screenshots are compiled into a consistent manual as a series of steps.

[1441] 6. Distribution of manuals

[1442] The completed operation manual will be generated in PDF or HTML format.

[1443] The server will provide this manual to users by displaying a download link via email or on the dashboard.

[1444] Users can download and view the manual.

[1445] Specific example

[1446] 1. A scene where the user clicks the "Create New Project" button and the process is explained verbally.

[1447] User: "I'll create a new project here."

[1448] The device records audio simultaneously with operation, along with a timestamp.

[1449] 2. Data transmission and processing after recording stops

[1450] The device sends the recorded data to the server.

[1451] The server uses speech recognition to extract the text "We will create a new project here" and matches it with the video of the operation.

[1452] 3. Example of manual generation

[1453] The server organizes the steps on the timeline and generates a page for the "1. Create a new project" section, combining corresponding screenshots and text.

[1454] In this way, by using the system of the present invention, an efficient and accurate operation manual is automatically generated based on the user's operations and explanations. The implementation of this invention significantly reduces the effort required to create manuals and enables the provision of highly accurate documentation.

[1455] The following describes the processing flow.

[1456] Step 1:

[1457] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[1458] Step 2:

[1459] The user clicks the "Start Recording" button. This causes the device to begin recording the entire current screen, saving all actions as a video file. Simultaneously, the device records the user's voice via the microphone. Specifically, the screen capture and audio recording processes are executed at the same time.

[1460] Step 3:

[1461] As the user performs screen operations, the system provides real-time audio explanations of those operations. For example, when the user clicks "New File" and selects it from the "File" menu, the system will provide the audio explanation, "Now you will create a new file." The device records these operations and audio, along with a timestamp.

[1462] Step 4:

[1463] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and saves the generated video and audio files to a temporary directory.

[1464] Step 5:

[1465] The terminal sends video and audio files stored in a temporary directory to the server. The files are sent using an HTTP POST request. Specifically, the video and audio files are uploaded to the server as form data.

[1466] Step 6:

[1467] The server receives the uploaded video and audio files. After receiving them, it passes the audio files to a speech recognition engine to begin the process of converting them to text. Here, speech recognition technology (e.g., a speech recognition API) is used to convert the audio data into text data.

[1468] Step 7:

[1469] The server analyzes the video file and generates frame-by-frame screenshots. Based on the video timeline, it extracts specific action events (e.g., button clicks or form inputs) and takes screenshots of the moments when these events occur.

[1470] Step 8:

[1471] The server integrates the converted text and screenshots to generate an operation manual. Based on the text and timestamp information, the screenshots are arranged as a sequence of steps to create a consistent operation manual. The generated operation manual is in PDF or HTML format.

[1472] Step 9:

[1473] The server generates a download link to provide the user with the completed operation manual. This link is sent to the system dashboard or to the user's email address.

[1474] Step 10:

[1475] Users can download and view the operation manual by clicking the provided download link. This operation manual is available for reference as needed and supports the use of the system.

[1476] (Example 1)

[1477] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1478] Traditional methods for creating user manuals require users to manually record each operation step and then transcribe the information, which is time-consuming and labor-intensive. Furthermore, the operation and its explanation may not always be synchronized, potentially leading to the creation of inaccurate manuals. Therefore, there is a need for a system that allows users to efficiently and accurately generate user manuals without having to spend time on incidental tasks.

[1479] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1480] In this invention, the server includes means for recording video of the screen operated by the user, means for recording the user's voice through a microphone, means for transmitting the recorded video file and the recorded audio file from the device to an information processing device, means for converting the audio file into text data on the information processing device, means for generating an operation manual based on the converted text data and video file, and means for providing the generated operation manual to the user. This makes it possible for the user to generate an operation manual simply and efficiently.

[1481] A "user" refers to anyone who intends to use this system to generate an operation manual.

[1482] "Screen images" refers to all visual information displayed on the screen of a device operated by the user.

[1483] A "microphone" is an audio input device used to record the user's voice.

[1484] A "recorded video file" is a digital file created by recording the video of the screen being operated by the user.

[1485] "Recorded audio files" refer to digital files generated by recording the user's voice.

[1486] A "device" refers to electronic devices such as computers and smartphones that are operated by the user.

[1487] An "information processing device" refers to a server or cloud computer used to process video and audio files transmitted from a device.

[1488] "Speech recognition technology" is a technology that converts audio files into text data; it analyzes audio data and turns it into text.

[1489] "Text data" refers to text information converted by speech recognition technology.

[1490] An "operation manual" refers to a document that summarizes the user's operating procedures and their explanations, and is generated based on video files and text data.

[1491] "Electronic document format" refers to digital document formats such as PDF and HTML.

[1492] "Web page format" refers to a web-based document that can be viewed on an internet browser.

[1493] This invention provides a system that instantly records user-performed computer screen operations and their descriptions, and automatically generates an operation manual later. This system allows users to efficiently and accurately generate operation manuals.

[1494] Hardware and software to be used

[1495] Terminal: A device used by the user to perform operations (e.g., personal computer, smartphone)

[1496] Microphone: A voice input device for recording the user's voice.

[1497] Screen capture tool: Software that records screen activity (e.g., OBS Studio)

[1498] Server: An information processing device that performs data processing and manual generation (e.g., cloud computing service).

[1499] Speech recognition technology: A technology that converts speech data into text data (e.g., Google Speech-to-Text API).

[1500] Image processing software: Technology for automatically generating screenshots from video (e.g., OpenCV)

[1501] HTML template engine: A technology for generating operation manuals in HTML format (e.g., Jinja2).

[1502] PDF generation software: Technology for generating manuals in PDF format (e.g., wkhtmltopdf).

[1503] Process Overview

[1504] 1. The user starts the system and grants permission for microphone and screen capture.

[1505] 2. The device checks the screen capture and microphone settings and completes the setup.

[1506] 3. When the user clicks the "Start Recording" button, the device simultaneously records the video and audio of the screen. For example, the user clicks "New File" and says aloud, "A new file has been created."

[1507] 4. When the user clicks the "Stop Recording" button, the device sends the recorded video and audio files to the server.

[1508] 5. The server analyzes the received video and audio files and uses speech recognition technology to convert the audio files into text data. For example, the audio data "A new file has been created" is converted into the text data "A new file has been created".

[1509] 6. The server integrates text data and timestamp information with video files to generate a consistent operation manual. The operation manual is output in HTML or PDF format using an HTML template engine or PDF generation software.

[1510] 7. The server will provide users with a download link for the generated operation manual via email or on the dashboard. Users can download and view the operation manual from this link.

[1511] Specific example

[1512] A scene where the user clicks the "Create New Project" button and the process is explained via voice.

[1513] User: "I'll create a new project here."

[1514] The device records audio simultaneously with this operation, along with a timestamp.

[1515] Data transmission and processing after recording stops

[1516] The device sends the recorded data to the server.

[1517] The server uses speech recognition to extract the text "We will create a new project here" and matches it with the video of the operation.

[1518] Manual generation example

[1519] The server organizes the steps on the timeline and generates a page for the step "1. Create a new project," combining corresponding screenshots and text.

[1520] Examples of prompt statements

[1521] 1. "Please record the steps to create a new project using both screen operations and audio."

[1522] 2. "Please continue by explaining the current screen operation aloud."

[1523] 3. "Once recording is complete, please click the 'End' button to send the data to the server."

[1524] The above details describe the embodiments of the present invention. This system enables users to generate highly accurate operation manuals in a short amount of time.

[1525] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1526] Step 1:

[1527] The user starts the system and grants permission for microphone and screen capture.

[1528] Input: Command to start the system, permission request popup.

[1529] Output: Permission granted

[1530] Specific action: When the user starts the system, a pop-up appears asking for permission to use the microphone and screen capture. The user clicks "Allow".

[1531] Step 2:

[1532] The device checks the microphone and screen capture settings and completes the setup.

[1533] Input: User-granted permissions

[1534] Output: Microphone and screen capture ready to use status

[1535] Specific operation: The device automatically checks the settings for the microphone and screen capture tool (e.g., OBS Studio) and displays a "Ready" message.

[1536] Step 3:

[1537] The user clicks the "Start Recording" button.

[1538] Input: User click of the "Start Recording" button.

[1539] Output: Trigger for recording start

[1540] Specific action: The user clicks the "Start Recording" button in the application window.

[1541] Step 4:

[1542] The device records all actions performed on the screen.

[1543] Input: Trigger for recording start

[1544] Output: Video data being recorded

[1545] Specific action: The OBS Studio software begins recording the user's screen activity. For example, it records the user clicking the "Create New Project" button.

[1546] Step 5:

[1547] The device records the user's voice description via the microphone.

[1548] Input: Trigger for recording start, user's voice

[1549] Output: Audio data being recorded

[1550] Specific actions: As the user performs an action, the system provides a voice explanation of that action. For example, it might say, "Here, we will create a new project."

[1551] Step 6:

[1552] The user clicks the "Stop Recording" button.

[1553] Input: User click of the "Stop Recording" button.

[1554] Output: Trigger for recording termination

[1555] Specific action: The user clicks the "Stop Recording" button after completing the operation.

[1556] Step 7:

[1557] The device sends the recorded video file and the recorded audio file to the server.

[1558] Input: Trigger for recording termination, video file, audio file

[1559] Output: Status of successful file transfer to server

[1560] Specific operation: The device uploads video and audio files to the server via the internet connection. For example, it sends files using an API and displays a "transmission complete" notification.

[1561] Step 8:

[1562] The server analyzes the received video and audio files.

[1563] Input: Video file, audio file

[1564] Output: Analyzed audio data, video data

[1565] Specific operation: The server temporarily stores the received data and separates the video file from the audio file.

[1566] Step 9:

[1567] The server converts the audio file into text data.

[1568] Input: Audio file

[1569] Output: Character data (text file)

[1570] Specific operation: Use the Google Speech-to-Text API to convert audio data into text data. For example, the audio "I have created a new file" will be converted into text.

[1571] Step 10:

[1572] The server integrates text and timestamp information with the video file.

[1573] Input: Text data, timestamp, video file

[1574] Output: Operation procedure data, screenshots

[1575] Specific operation: Use OpenCV to obtain appropriate screenshots from video, and generate a consistent instruction manual based on text data and timestamp information.

[1576] Step 11:

[1577] The server generates an operation manual.

[1578] Input: Operation procedure data, screenshots

[1579] Output: Operation manual (PDF or HTML format)

[1580] Specific operation: Use the Jinja2 template engine to generate an operation manual in HTML format, and then convert it to PDF format using wkhtmltopdf.

[1581] Step 12:

[1582] The server provides users with an operation manual.

[1583] Input: Generated operation manual

[1584] Output: Download link, notification email

[1585] Specific operation: The server saves the generated manual and provides a download link to the user. The user can download the manual by displaying the link in an email or on the dashboard.

[1586] (Application Example 1)

[1587] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1588] Creating manuals for operating robots and machinery within a factory is time-consuming and labor-intensive, and accurately conveying advanced operations and complex procedures is difficult. Furthermore, training new workers is lengthy and prone to errors, highlighting the need for improved production efficiency and safety.

[1589] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1590] In this invention, the server includes means for recording video of the screen operated by the user, means for recording the user's voice through a microphone, means for using a visual device to record video of the real environment, means for transmitting the recorded video file and recorded audio file to the server, means for converting the audio file to text on the server, means for generating an operation manual based on the converted text and video file, and means for providing the generated operation manual to the user. This enables accurate and rapid recording of complex procedures and operations within the factory, efficient automatic generation of operation manuals, reduction of worker training time, and prevention of errors.

[1591] "Means for recording video of the user's screen" refers to a function that records all operations performed on the screen of a computer or device being operated by the user in video format.

[1592] "Means of recording user voice via microphone" refers to a function that uses a microphone to record the voice of the user when they give explanations or instructions while operating the device.

[1593] "Means of using visual devices to record images of the real environment" refers to a function that allows users to record images of their surroundings or the object they are operating in real time using devices such as smart glasses or cameras.

[1594] "Means for sending recorded video files and recorded audio files to a server" refers to a function for uploading recorded video data and audio data to a server via the internet or other means.

[1595] "Method for converting audio files to text on the server" refers to a function that uses speech recognition technology on the server side to automatically convert transmitted audio data into text format.

[1596] "A means of generating operation manuals based on converted text and video files" refers to a function that automatically creates consistent operation procedures and manuals by combining audio-to-text data and recorded video data.

[1597] "Means of providing the generated operation manual to the user" refers to a function that outputs the completed operation manual in PDF or HTML format, allowing the user to download or view it.

[1598] This invention is a system that records robot operation and maintenance procedures in a factory in real time and automatically generates an operation manual based on that data. The embodiments of this system are described in detail below.

[1599] System Overview

[1600] 1. Initial Setup

[1601] The user starts the system and grants permission for microphone and screen capture. Additionally, the user wears a visual device (e.g., smart glasses or a camera) and confirms the settings.

[1602] The device checks the settings for microphone, screen capture, and visual devices, and then completes the setup.

[1603] 2. Recording screen operations and audio.

[1604] Recording begins when the user clicks the "Start Recording" button.

[1605] The device records the user's actions on the device screen as a video file, and simultaneously records video of the real environment through a visual device.

[1606] At the same time, the device records the user's voice description through the microphone.

[1607] 3. Sending data

[1608] Once all recording and recording are complete, the user clicks the "Stop Recording" button.

[1609] The terminal sends the generated video and audio files to the server. This transmission takes place over the internet.

[1610] 4. Data Processing

[1611] The server analyzes the received video and audio files and uses speech recognition technology to convert the audio files into text data. For example, the audio data "I have created a new project" is converted to text "I have created a new project".

[1612] 5. Generation of the operation manual

[1613] The server combines the text with the corresponding timestamp information and integrates it with a specific portion (screenshot) of the video file.

[1614] Specifically, text and screenshots are compiled into a consistent manual as a series of steps.

[1615] 6. Distribution of manuals

[1616] The completed operation manual will be generated in PDF or HTML format.

[1617] The server will provide this manual to users by displaying a download link via email or on the dashboard.

[1618] Users can download and view the manual.

[1619] Specific example

[1620] Consider a scenario where factory workers wear smart glasses to perform robot maintenance procedures. For example, they might record the procedure for replacing a hydraulic cylinder.

[1621] The user performs the task while explaining verbally, "Here, we will remove the hydraulic cylinder," and both the operation and the audio are recorded simultaneously.

[1622] After recording stops, the terminal sends the recorded data to the server, converts the audio data into text, and generates a manual.

[1623] Example of a prompt

[1624] "Use video and audio capture to record robot operation procedures within the factory. After recording is complete, convert the audio data into text and generate an operation manual along with the video frames. For example, if you perform a task while explaining, 'Here we remove the hydraulic cylinder,' ensure that this procedure and explanation are integrated into a single procedure in the manual."

[1625] In this way, by using the system of the present invention, complex procedures and operations within a factory can be accurately and quickly recorded, and efficient operation manuals can be automatically generated, thereby reducing worker training time and preventing operational errors.

[1626] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1627] Step 1:

[1628] Initial settings:

[1629] The user starts the system and grants permissions for the microphone, screen capture, and visual devices (3D glasses or camera). The terminal verifies that these devices are working correctly and completes the setup. The input is the permission and setup required for each device, and the output is the result of verifying that the devices are ready for use.

[1630] Step 2:

[1631] Start recording screen operations and audio:

[1632] The recording process begins when the user clicks the "Start Recording" button. The device records all operations performed on the computer screen as a video file using screen capture software (e.g., PyAutoGUI). Simultaneously, audio recording software (e.g., PyAudio) records the user's voice explanations via the microphone. The visual device records the user's viewpoint in real time. Input data consists of user operations, audio, and visual environment, and the output is a recorded file.

[1633] Step 3:

[1634] Sending data:

[1635] Once recording is complete, the user clicks the "Stop Recording" button. The device then sends the recorded video files (screen captures and video from the visual device) and audio files to the server. The transmission takes place over the internet. The input data consists of the recorded video and audio files, and the output is the data uploaded to the server.

[1636] Step 4:

[1637] Data processing:

[1638] The server analyzes the received video and audio files. Audio files are converted to text using speech recognition technology (e.g., Google Cloud Speech-to-Text API). The input is an audio file, and the output is the corresponding text data. Video files are also organized based on timestamps.

[1639] Step 5:

[1640] Generating the operation manual:

[1641] The server combines the converted text data with corresponding timestamp information and integrates it with specific portions (screenshots) of the video file. This generates a consistent operation manual. The input is text data and video files, and the output is an integrated operation manual file.

[1642] Step 6:

[1643] Distribution of manuals:

[1644] The server generates a completed operation manual in PDF or HTML format and makes it available for users to download. The download link is provided via email or on the dashboard. The input is the generated operation manual, and the output is the download link that users can access.

[1645] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1646] This invention provides a system that instantly records user computer screen operations and their explanations, and later automatically generates an operation manual. Furthermore, by combining it with an emotion engine, it enables the recognition and utilization of user emotions. The embodiments thereof will be described in detail below.

[1647] System Overview

[1648] 1. Initial Setup

[1649] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[1650] 2. Recording screen operations and audio.

[1651] The user clicks the "Start Recording" button. This causes the device to begin recording the entire current screen and save all actions as a video file. Simultaneously, the device records the user's voice explanation via the microphone.

[1652] The device activates an emotion engine, which analyzes the user's emotions in real time from their voice and facial expressions, and collects emotional data.

[1653] 3. Sending data

[1654] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and save the generated video and audio files, as well as emotion data, to a temporary directory.

[1655] The device sends these files to the server. The transmission is done using an HTTP POST request, and video files, audio files, and emotion data are uploaded to the server as form data.

[1656] 4. Data Processing

[1657] The server receives uploaded video files, audio files, and emotion data. After receiving the data, it passes the audio files to a speech recognition engine to begin the process of converting them to text. Specifically, speech recognition technology is used to convert the audio data into text data.

[1658] The server also analyzes emotional data and records it along with the corresponding timestamp.

[1659] 5. Generation of the operation manual

[1660] The server integrates text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual.

[1661] The generated operation manual also displays the user's emotions at each operation step. For example, along with a description such as "You have created a new file," it also shows the emotions the user felt at that moment (e.g., joy or confusion).

[1662] 6. Distribution of manuals

[1663] The completed user manual will be generated in PDF or HTML format. In addition to the operating procedures, the generated manual will visually display emotional data, making it clear what emotions the user felt during each operation.

[1664] The server generates a download link to provide this manual to the user. This link is sent to the system dashboard or to the user's email address.

[1665] Users can download and view the operation manual by clicking the provided download link. This operation manual is available for reference as needed and supports the use of the system.

[1666] Specific example

[1667] 1. The user clicks the "Create New Project" button.

[1668] User: "I'll create a new project here."

[1669] The device records audio simultaneously with operation, along with a timestamp.

[1670] At the same time, the emotion engine analyzes the user's emotions and collects emotional data such as "joy" and "anticipation."

[1671] 2. Data transmission and processing after recording stops

[1672] The terminal sends the recorded data to the server. The server uses speech recognition to extract the text "Here we will create a new project" and matches it with the video of the operation.

[1673] Emotional data is also analyzed and recorded along with the corresponding timestamp.

[1674] 3. Example of manual generation

[1675] The server organizes the steps on the timeline and generates a page for the "Create a new project" section, combining the corresponding screenshots, text, and sentiment data at that time.

[1676] The manual displays an emotion such as "Joy" below an explanation like, "Here we will create a new project."

[1677] In this way, the system of the present invention can generate a detailed and intuitive operation manual, including emotional data, based on the user's operations and explanations. This enriches the user's experience and allows for the provision of feedback utilizing emotional information.

[1678] The following describes the processing flow.

[1679] Step 1:

[1680] The user starts the system and accesses the initial setup screen. The terminal displays a dialog box requesting permission for microphone and screen capture. Once the user grants these permissions, the terminal confirms the settings and completes the setup.

[1681] Step 2:

[1682] The user clicks the "Start Recording" button. This causes the device to begin recording the entire screen and save the operation as a video file. Simultaneously, the device records the user's voice through the microphone. The device also activates an emotion engine to analyze the user's emotions in real time from their facial expressions and voice.

[1683] Step 3:

[1684] As the user performs screen operations, the system provides real-time audio explanations of those operations. For example, when the user clicks "New File" and selects it from the "File" menu, the system will say, "Now you will create a new file." The device records these operations and audio along with timestamps. Simultaneously, it analyzes and records sentiment data along with timestamps.

[1685] Step 4:

[1686] The user clicks the "Stop Recording" button. This causes the device to stop recording video and audio, and save the generated video and audio files, as well as emotion data, to a temporary directory.

[1687] Step 5:

[1688] The device sends video files, audio files, and emotion data stored in a temporary directory to the server. The files are sent using an HTTP POST request, and the video files, audio files, and emotion data are uploaded to the server as form data.

[1689] Step 6:

[1690] The server receives the uploaded video files, audio files, and emotion data and begins processing them. First, the audio files are passed to the speech recognition engine to be converted into text. In this process, for example, audio data such as "Here I will create a new file" is converted into text data.

[1691] Step 7:

[1692] The server analyzes the video file and generates frame-by-frame screenshots on the timeline. Based on the video's timestamps, it extracts specific action events (e.g., button clicks or form inputs) and takes screenshots of the moments when these events occurred.

[1693] Step 8:

[1694] The server integrates text, screenshots, and sentiment data to generate an operation manual. The manual sequentially displays text and corresponding screenshots as operating procedures, and also shows the user's sentiment data at each step. For example, under the explanation "Create a new file," the corresponding screenshot and the sentiment data "Joy" are displayed.

[1695] Step 9:

[1696] The server saves the generated operation manual in PDF or HTML format and generates a download link for the user. This link is sent to the system dashboard or to the user's email address.

[1697] Step 10:

[1698] Users can download and view the operation manual by clicking the provided download link. This operation manual visually presents user sentiment data along with the operating procedures, making it easier to understand intuitively.

[1699] (Example 2)

[1700] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1701] Conventional operation manual generation systems could record user instructions and screen operations, but they had the problem of not being able to obtain feedback that included the user's emotions. As a result, feedback based on the user's emotions was not provided, leading to a lack of deeper understanding of how to operate the system.

[1702] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting audio files to text, means for generating an operation manual based on the converted text and video files, means for analyzing the user's emotions in real time and collecting emotion data, and means for integrating the collected emotion data into the operation manual. This makes it possible to generate a more detailed and intuitive operation manual that reflects not only the user's operation instructions but also their emotional information at that time.

[1703] A "user" is the entity that operates the system and performs various inputs.

[1704] "Means of recording screen footage" refers to functions or devices that record the user's screen in video format.

[1705] "Means of recording audio via a microphone" refers to functions or devices for recording the voice spoken by a user.

[1706] "Means of sending to a server" refers to the technology and processes used to upload recorded video files and audio files to a server via a network.

[1707] "Methods for converting audio files to text" refers to the process of converting audio data into text data using speech recognition technology.

[1708] "Means for generating operation manuals" refers to technologies and processes for creating documents that visually show user operation procedures based on acquired audio text and video files.

[1709] "Means of providing operation manuals to users" refers to the processes and functions for providing the generated operation manuals in a user-accessible format (e.g., PDF or HTML).

[1710] "Means of analyzing emotions in real time and collecting emotional data" refers to technologies and devices that analyze a user's facial expressions and voice to identify and record the emotions the user is feeling in real time.

[1711] "Means of integrating emotional data into the operation manual" refers to the techniques and processes used to link collected emotional data to each operation step in the operation manual and ensure consistency.

[1712] The system according to the present invention can instantly record the computer screen operations performed by the user and their explanations, and further recognize, analyze, and utilize the user's emotions by combining them with an emotion engine. Specific embodiments of the present invention will be described in detail below.

[1713] Overall overview

[1714] This system is implemented through a process that records the user's screen operations and their voice explanations via a microphone, sends this data to a server for analysis, and ultimately generates an operation manual. Furthermore, it analyzes the user's emotions using an emotion engine and incorporates this analysis into the operation manual.

[1715] Initial setup

[1716] When the user starts the system, the terminal displays an initial setup screen and a dialog box prompting the user to grant microphone and screen capture permissions. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the preparation for recording and audio.

[1717] Screen operation and audio recording

[1718] When the user clicks the "Start Recording" button, the device begins recording the entire current screen and saves all operations as a video file. Simultaneously, the device records the user's voice description via the microphone. The device also activates an emotion engine and uses OpenCV to analyze the user's voice and facial expressions in real time, collecting emotion data.

[1719] Sending data

[1720] When the user clicks the "Stop Recording" button, the device stops recording video and audio, and saves the generated video and audio files, as well as emotion data, to a temporary directory. The device then uses the Python requests library to send these files to the server via an HTTP POST request.

[1721] Data processing

[1722] The server receives uploaded video and audio files, as well as sentiment data. Audio files are converted to text using the Google Cloud Speech-to-Text API. Additionally, the server analyzes the sentiment data and records it along with the corresponding timestamp.

[1723] Generating an operation manual

[1724] The server integrates the converted text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. The manual also visually displays the user's emotions at each operation step. Specifically, it generates a PDF version of the manual using LaTeX and also provides it in HTML format.

[1725] Distribution of manuals

[1726] The completed operation manual is generated in PDF or HTML format, and the server uses Flask to generate download links. Users can download and view the operation manual by clicking the download link provided via the dashboard or their email address.

[1727] Specific examples and prompt statements

[1728] Specific example

[1729] The scenario envisions a user clicking the "Create New Project" button and being instructed, "Here, we will create a new project." The device simultaneously records the user's actions and voice, and uses an emotion engine to analyze the user's emotions, such as "joy" and "expectation." Subsequently, the server generates an operation manual based on the collected data and provides it to the user.

[1730] Example of a prompt

[1731] "Record the steps for creating a new project and generate a manual that includes emotional data along with the operating instructions."

[1732] "Based on the recorded operating procedures, use the emotion engine to analyze the user's emotions and create a detailed operating manual."

[1733] The above describes specific embodiments for carrying out the present invention. By utilizing the present invention, it becomes possible to enrich the user's operational experience and provide valuable feedback that utilizes emotional information.

[1734] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1735] Step 1:

[1736] Initial setup

[1737] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant permission for microphone and screen capture. When the user clicks "Allow," the terminal confirms and saves these settings.

[1738] Input: User actions (system startup, permission granting)

[1739] Data processing: Permission verification and saving of settings

[1740] Output: Microphone and screen capture ready.

[1741] Step 2:

[1742] Screen recording

[1743] The user clicks the "Start Recording" button. The device begins recording the entire current screen and saves all user actions as an MP4 video file. FFmpeg is used as the recording tool.

[1744] Input: User action (click of the "Start Recording" button)

[1745] Data processing: Screen recording

[1746] Output: MP4 video file (recording.mp4)

[1747] Step 3:

[1748] Audio recording

[1749] The device simultaneously starts microphone input and saves the user's voice as a WAV audio file. Pyaudio is used as the audio library.

[1750] Input: User's voice

[1751] Data processing: Audio recording

[1752] Output: WAV format audio file (audio.wav)

[1753] Step 4:

[1754] Emotion analysis

[1755] The device uses OpenCV to analyze camera input, analyzes the user's emotions in real time from their voice and facial expressions, and saves the emotion data in JSON format.

[1756] Input: User's facial expressions, voice

[1757] Data processing: Real-time analysis of emotions

[1758] Output: Emotion data file in JSON format (emotions.json)

[1759] Step 5:

[1760] Sending data

[1761] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The generated video file, audio file, and emotion data are saved to a temporary directory. The device uses the Python requests library to send these files to the server via an HTTP POST request.

[1762] Input: User action (click of the "Stop Recording" button), generated file

[1763] Data processing: Saving and sending files

[1764] Output: File sent to the server

[1765] Step 6:

[1766] Converting audio data to text

[1767] The server sends the received audio file to the Google Cloud Speech-to-Text API, where it converts the audio data into text.

[1768] Input: Audio file (audio.wav)

[1769] Data processing: Text conversion using speech recognition.

[1770] Output: Text data

[1771] Step 7:

[1772] Analysis of emotional data

[1773] The server analyzes the received emotion data file and stores each emotion data item in the database along with its corresponding timestamp.

[1774] Input: Emotion data file (emotions.json)

[1775] Data processing: Analysis of emotional data and its correspondence with timestamps.

[1776] Output: Emotional data stored in the database

[1777] Step 8:

[1778] Generating an operation manual

[1779] The server integrates the received text data and timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. It generates the manual in PDF format using LaTeX, and also in HTML format.

[1780] Input: Text data, timestamp information, video files

[1781] Data processing: Capture and integrate screenshots, generate consistent operation manuals.

[1782] Output: Operation manual in PDF and HTML formats

[1783] Step 9:

[1784] Provision of operation manual

[1785] The server generates a download link for the operation manual using Flask and provides it to the user via their dashboard and email. Users can download and view the operation manual by clicking the download link.

[1786] Input: Generated operation manual

[1787] Data processing: Creating download links

[1788] Output: Providing download links to users

[1789] (Application Example 2)

[1790] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1791] Conventional user manual generation systems lacked the ability to consider user emotions, resulting in a failure to integrate emotional information into operating procedures and explanations. Especially in online learning platforms, there is a demand for learning content that accurately reflects the instructor's emotional state and explanations. Therefore, an intuitive user manual generation system with emotion analysis capabilities is needed to deepen users' understanding of the operations.

[1792] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting emotion data, means for converting audio files into text, and means for generating an operation manual based on the converted text, video files, and emotion data. This makes it possible to provide an operation manual that integrates user operations and explanations with the emotion information at that time.

[1793] "User" refers to the entity that operates this system and uses it to record screen operations and audio.

[1794] "Screen images" refer to the visual information displayed on the monitor of a computer operated by a user.

[1795] "Means of recording" refers to a device or software that captures and records video of the screen being operated by the user.

[1796] A "microphone" refers to a device that converts sound into electrical signals and records the user's voice.

[1797] "Means of recording audio" refers to a device or software that uses a microphone to record the user's voice.

[1798] "Means for analyzing user facial expressions" refers to devices or software that capture the user's facial movements and expressions, analyze them, and obtain emotional data.

[1799] "Emotional data" refers to emotional information analyzed from the user's facial expressions and voice.

[1800] A "server" refers to a computer system used to store, process, and generate operation manuals for recorded data.

[1801] A "video file" refers to a digital file containing recorded data of the user's screen.

[1802] An "audio file" refers to a digital file containing user voice data recorded via a microphone.

[1803] "Speech recognition technology" refers to algorithms and techniques for converting speech data into text data.

[1804] An "operation manual" refers to a set of instructions or manuals that include user actions, explanations, and emotional data.

[1805] "Means of provision" refers to devices or software for distributing or displaying the generated operation manual to the user.

[1806] "PDF format" is an abbreviation for Portable Document Format, and refers to a digital file format that allows users to view and print operation manuals.

[1807] "HTML format" is an abbreviation for HyperText Markup Language, and refers to a file format used to display instruction manuals on the web.

[1808] System Overview

[1809] Initial setup

[1810] The user starts the system and accesses the initial setup screen. At this point, the terminal displays a dialog box prompting the user to grant permission for microphone and screen capture. Once the user grants these permissions, the terminal verifies the microphone and screen capture settings and completes the setup.

[1811] Screen operation and audio recording

[1812] When the user clicks the "Start Recording" button, the device begins recording the entire current screen and saves all operations as a video file. Simultaneously, the device records the user's voice explanation via the microphone. In addition, the device captures the user's facial expressions and analyzes and collects emotion data in real time.

[1813] Sending data

[1814] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The generated video file, audio file, and emotion data are then saved to a temporary directory and sent to the server. The transmission is performed using an HTTP POST request, and the video file, audio file, and emotion data are uploaded to the server as form data.

[1815] Data processing

[1816] The server receives uploaded video files, audio files, and sentiment data. After receiving the data, the server passes the audio files to a speech recognition engine to begin the process of converting them to text. Specifically, speech recognition technology is used to convert the audio data into text data. The server also analyzes the sentiment data and records it along with the corresponding timestamp.

[1817] Generating an operation manual

[1818] The server integrates text and corresponding timestamp information with specific parts of the video file (screenshots) to generate a consistent operation manual. The generated operation manual also displays the user's emotions at each operation step.

[1819] Distribution of manuals

[1820] The completed user manual is generated in PDF or HTML format. In addition to the operating procedures, the generated manual visually displays sentiment data, clearly showing what emotions the user experienced during each operation. The server generates a download link to provide this manual to the user. This link is sent to the system dashboard or to the user's email address. The user can download and view the user manual by clicking the provided download link.

[1821] Program Processing Overview

[1822] Hardware and software to be used

[1823] Hardware:

[1824] Webcam: Used for video and sentiment analysis.

[1825] Microphone: Used for voice recording

[1826] software:

[1827] OpenCV: Used for video capture and saving

[1828] pyaudio: Used for audio recording

[1829] wave: For saving audio files

[1830] Request: For sending files via HTTP POST request.

[1831] EmotionRecognizer: For emotion analysis

[1832] speech_recognition: For speech recognition

[1833] Processing details

[1834] The terminal simultaneously performs video and audio recording and emotional data analysis. The recorded video files, audio files, and emotional data are sent to the server. The server converts the audio data to text and integrates the emotional data with timestamps to generate an operation manual. Finally, the generated manual is distributed to the user in PDF or HTML format.

[1835] Specific example

[1836] If the system is started: "start_recording('Lesson Start', 'Audio Explanation')"

[1837] When the operation is finished: "analyze_emotions('lesson_audio.wav', 'lesson_video.avi')"

[1838] Generating a lesson operation manual: "generate_manual('generated text', 'emotion data', 'lesson_manual.txt')"

[1839] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1840] Step 1:

[1841] The user starts the system and accesses the initial setup screen. The terminal displays a dialog prompting the user to grant permission for microphone and screen capture. The input is the user's actions, and the output is the state of readiness for completing the microphone and screen capture permission settings. Specifically, the terminal verifies the user's permission and configures the microphone and screen capture settings.

[1842] Step 2:

[1843] When the user clicks the "Start Recording" button, the device begins recording the entire current screen. The input is the click of the "Start Recording" button, and the output is the start of recording the video file. Specifically, the device uses OpenCV to capture and save the screen video.

[1844] Step 3:

[1845] The device simultaneously records the user's voice description via the microphone. The input is the user's voice, and the output is an audio file. Specifically, the device uses pyaudio to record the audio and saves it using the wave library.

[1846] Step 4:

[1847] The device captures the user's facial expressions and analyzes and collects emotional data in real time. The input is a video of the user's face, and the output is emotional data. Specifically, the device uses its EmotionRecognizer to analyze facial expressions and saves them as emotional data.

[1848] Step 5:

[1849] When the user clicks the "Stop Recording" button, the device stops recording video and audio. The input is the click of the "Stop Recording" button, and the output is the completion of collecting video and audio files, as well as emotion data. Specifically, it stops recording video and audio and saves the data to a temporary directory.

[1850] Step 6:

[1851] The device sends recorded video files, recorded audio files, and emotion data to the server. The inputs are video files, audio files, and emotion data, and the output is the uploading of this data to the server. Specifically, the device uses the requests library to send an HTTP POST request and upload the data to the server.

[1852] Step 7:

[1853] The server receives uploaded video files, audio files, and emotion data. The input is the transmitted data, and the output is the completion of data reception on the server side. Specifically, the server saves the files to storage.

[1854] Step 8:

[1855] The server passes the audio file to the speech recognition engine and begins the process of converting it to text. The input is an audio file, and the output is text data. Specifically, the server uses the speech_recognition library to convert the audio to text.

[1856] Step 9:

[1857] The server analyzes sentiment data and records it along with a timestamp. The input is sentiment data, and the output is the sentiment analysis result with a timestamp. Specifically, the server analyzes the sentiment data and integrates it with the corresponding time information.

[1858] Step 10:

[1859] The server generates operation manuals based on text, video files, and sentiment data. Inputs are text data, screenshots, and sentiment data, while output is the operation manual. Specifically, the server organizes this data into a consistent set of operating procedures and creates the manual in PDF or HTML format.

[1860] Step 11:

[1861] The server provides the user with a completed operation manual. The input is the generated operation manual, and the output is a download link. Specifically, the server creates a download link and sends it to the system dashboard or the user's email address.

[1862] Example of a prompt

[1863] If the system is started: "start_recording('Lesson Start', 'Audio Explanation')"

[1864] When the operation is finished: "analyze_emotions('lesson_audio.wav', 'lesson_video.avi')"

[1865] Generating a lesson operation manual: "generate_manual('generated text', 'emotion data', 'lesson_manual.txt')"

[1866] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1867] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1868] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1869] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1870] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1871] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1872] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1873] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1874] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1875] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1876] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1877] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1878] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1879] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1880] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1881] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1882] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1883] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1884] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1885] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1886] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1887] The following is further disclosed regarding the embodiments described above.

[1888] (Claim 1)

[1889] A means of recording video of the screen being operated by the user,

[1890] A means of recording the user's voice through a microphone,

[1891] A means of sending recorded video files and recorded audio files to a server,

[1892] A means of converting audio files to text on a server,

[1893] A means of generating an operation manual based on converted text and video files,

[1894] A means of providing the generated operation manual to the user,

[1895] A system that includes this.

[1896] (Claim 2)

[1897] The system according to claim 1, which uses speech recognition technology to convert an audio file into text.

[1898] (Claim 3)

[1899] The system according to claim 1, which provides an operation manual in PDF or HTML format.

[1900] "Example 1"

[1901] (Claim 1)

[1902] A means of recording video of the screen being operated by the user,

[1903] A means of recording the user's voice through a microphone,

[1904] A means for transmitting recorded video files and recorded audio files from a device to an information processing device,

[1905] A means for converting audio files into text data on an information processing device,

[1906] A means of generating an operation manual based on converted text data and video files,

[1907] A means of providing the generated operation manual to the user,

[1908] A system that includes this.

[1909] (Claim 2)

[1910] The system according to claim 1, which uses speech recognition technology to convert an audio file into text data.

[1911] (Claim 3)

[1912] The system according to claim 1, which provides an operation manual in electronic document format or web page format.

[1913] "Application Example 1"

[1914] (Claim 1)

[1915] A means of recording video of the screen being operated by the user,

[1916] A means of recording the user's voice through a microphone,

[1917] Means of using visual devices to record images of the real environment,

[1918] A means of sending recorded video files and recorded audio files to a server,

[1919] A means of converting audio files to text on a server,

[1920] A means of generating an operation manual based on converted text and video files,

[1921] A means of providing the generated operation manual to the user,

[1922] A system that includes this.

[1923] (Claim 2)

[1924] The system according to claim 1, which uses speech recognition technology to convert an audio file into text.

[1925] (Claim 3)

[1926] The system according to claim 1, which provides an operation manual in PDF or HTML format.

[1927] "Example 2 of combining an emotion engine"

[1928] (Claim 1)

[1929] A means of recording video of the screen being operated by the user,

[1930] A means of recording the user's voice through a microphone,

[1931] A means of sending recorded video files and recorded audio files to a server,

[1932] A means of converting audio files to text on a server,

[1933] A means of generating an operation manual based on converted text and video files,

[1934] A means of providing the generated operation manual to the user,

[1935] A means of analyzing user emotions in real time and collecting emotional data,

[1936] A means of integrating collected emotional data into the operation manual,

[1937] A system that includes this.

[1938] (Claim 2)

[1939] The system according to claim 1, which uses speech recognition technology to convert an audio file into text.

[1940] (Claim 3)

[1941] The system according to claim 1, which provides an operation manual in PDF or HTML format.

[1942] (Claim 4)

[1943] The system according to claim 1, which uses emotion analysis technology to analyze the user's emotions.

[1944] "Application example 2 when combining with an emotional engine"

[1945] (Claim 1)

[1946] A means of recording video of the screen being operated by the user,

[1947] A means of recording the user's voice through a microphone,

[1948] A method for analyzing users' facial expressions and collecting emotional data,

[1949] A means for transmitting recorded video files, recorded audio files, and emotional data to a server,

[1950] A means of converting audio files to text on a server,

[1951] A means for generating an operation manual based on converted text, video files, and emotion data,

[1952] A means of providing the generated operation manual to the user,

[1953] A system that includes this.

[1954] (Claim 2)

[1955] The system according to claim 1, which uses speech recognition technology to convert an audio file into text.

[1956] (Claim 3)

[1957] The system according to claim 1, which provides an operation manual in PDF or HTML format. [Explanation of Symbols]

[1958] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of recording video of the screen being operated by the user, A means of recording the user's voice through a microphone, A means of sending recorded video files and recorded audio files to a server, A means of converting audio files to text on a server, A means of generating an operation manual based on converted text and video files, A means of providing the generated operation manual to the user, A system that includes this.

2. The system according to claim 1, which uses speech recognition technology to convert an audio file into text.

3. The system according to claim 1, which provides an operation manual in PDF or HTML format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A