system

A system that automates business manual creation and task execution using voice and screen data analysis improves productivity and resource efficiency in small enterprises by standardizing processes and reducing manual labor.

JP2026069137APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

In small and medium-sized enterprises, creating and maintaining business manuals is labor-intensive, prone to human errors, and leads to decreased productivity and competitiveness due to procedural inconsistencies and inefficient resource utilization.

Method used

A system that analyzes voice input and screen operations to automatically generate business manuals, standardize processes, and execute tasks efficiently using a cloud-based platform, incorporating speech recognition and generative AI models.

Benefits of technology

Optimizes resource use, reduces manual effort, and enhances productivity by providing reproducible business processes and efficient task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069137000001_ABST
    Figure 2026069137000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A processing means for acquiring voice input and converting that voice into text data, A processing means for monitoring user screen operations and acquiring screen captures at important moments, A processing means that analyzes business procedures based on converted text data and acquired screen captures, and automatically generates a business manual. A storage method for saving and managing the generated business manuals on the cloud, From the next time onward, a control means will automatically execute tasks based on the saved business manual in response to user requests, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In an information processing apparatus, especially in an environment with limited resources such as small and medium-sized enterprises, it is difficult to improve business efficiency because creating and maintaining business manuals requires labor and time. Also, human errors and procedural inconsistencies may occur due to manual business handover. As a result, there is a problem that business productivity decreases, which in turn leads to a decrease in competitiveness.

Means for Solving the Problems

[0005] This invention provides a technology for analyzing business procedures by converting voice input into text data and monitoring user screen operations to acquire screen captures. Based on the analyzed procedures, a business manual is automatically generated, stored and managed in the cloud, thereby establishing a reproducible business process. Furthermore, by automatically executing subsequent tasks, the technology achieves standardization and efficiency. This optimizes the resources necessary for task execution and supports the smooth execution of business processes, even in environments with severe labor shortages.

[0006] "Voice input" refers to a method in which a user provides instructions or information to a system verbally via a microphone or similar device.

[0007] "Text data" refers to data in audio or other formats that is represented as written information.

[0008] "Screen operation" refers to a series of actions performed by a user through a computer interface.

[0009] "Screen capture" refers to obtaining an image of a computer screen at a specific point in time while a user is operating it.

[0010] A "business procedure" is a series of steps or processes necessary to complete a specific task.

[0011] A "business manual" is a document that describes the procedures and methods for performing a specific task.

[0012] "Storage means" refers to a method or apparatus for recording data in a specific location and making it accessible at a later date.

[0013] "Managed in the cloud" means utilizing external servers via the internet to store and control data and services.

[0014] An "automatically executed control mechanism" is a mechanism that allows a system to perform tasks based on pre-set procedures without explicit user intervention.

[0015] A "speech recognition engine" is software or hardware technology used to convert speech data into text format.

[0016] A "generative AI model" is an artificial intelligence algorithm that learns from data and automatically identifies and applies specific patterns and procedures. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the language used in the following description will be explained.

[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] This invention provides a system that automatically generates work manuals using voice and screen operation data to streamline operations, and automatically executes tasks based on those manuals. This system is provided on a cloud-based platform, and users can access and use it through their own devices.

[0039] For the system to be implemented, it is a prerequisite that the user terminal is equipped with hardware (microphone) and software (voice recognition function) that enables voice input. When performing a task, the user will explain the procedure for which they wish to create a manual by voice. This explanation will be captured by the user's terminal and sent to the server.

[0040] The server analyzes the audio data using a dedicated speech recognition engine and converts it into text data. This text data forms the basis for the subsequent process of generating work manuals. Additionally, the user terminal monitors the screen as the user performs their tasks and takes screenshots at critical operation steps. This image information is also sent to the server.

[0041] The server analyzes the received text and screen capture data to construct work procedures. A generative AI model is used for the analysis process, automatically recognizing the procedures based on the data and generating a work manual based on the extracted information. The generated manual includes detailed procedures and corresponding images, and is securely stored in the cloud.

[0042] For example, if a user explains the expense reimbursement process verbally and then performs the actual operations on a terminal, the server will generate a manual for the expense reimbursement process based on this verbal explanation and operation. This manual will clearly explain everything from how to start the software to the procedure for entering the necessary information and the verification process.

[0043] In subsequent instances, when a user performs a similar task, their terminal will request automated execution of the task. The server will refer to the task manual stored in the cloud, and the generated AI model will automatically complete the task according to the procedure. As a result, task standardization will progress, and more efficient work execution will be possible.

[0044] The introduction of this system will reduce the need for manual manual creation and work handover, thereby optimizing human resources and improving productivity. This technology will particularly contribute to the effective use of resources in small and medium-sized enterprises and provides a new solution for business operations.

[0045] The following describes the processing flow.

[0046] Step 1:

[0047] The user speaks voice instructions related to their work into the terminal. The terminal acquires this voice in real time and temporarily saves it as a digital audio file.

[0048] Step 2:

[0049] The terminal sends the acquired voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. This converted text accurately reflects the work instructions given by the user.

[0050] Step 3:

[0051] When a user performs a task, the terminal monitors the user's screen operations. It then automatically captures a screen image when a specific operation trigger is detected. This capture contains important operational details necessary for creating procedure manuals.

[0052] Step 4:

[0053] The terminal sends the acquired screen capture data to the server. The server uses a generating AI model to analyze the text data and screen captures, and automatically extracts the components of the business procedure.

[0054] Step 5:

[0055] The server automatically generates an operational manual based on the extracted information. This manual combines screen captures corresponding to instructions with detailed operational procedures. The generated manual is securely stored in cloud storage.

[0056] Step 6:

[0057] From the next time onward, the user will request the same task to be performed automatically from the terminal. The server will read the task manual stored in the cloud, and the generated AI model will accurately reproduce the steps required for automated execution.

[0058] Step 7:

[0059] After the server completes its task, the terminal notifies the user of the results. If the process was successful, or if an error occurred, the user is informed of the details.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] Traditional methods of creating operational manuals and automating business processes were inefficient and required significant human resources. In particular, in small organizations, a lack of standardization and insufficient knowledge sharing led to decreased productivity. Furthermore, manual manual updates were infrequent, making it difficult to maintain consistency in work procedures.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means equipped with an information processing device that acquires voice data and converts that data into text data; means equipped with a device that monitors the user's operation procedures and collects visual data for specific actions; and means equipped with a device that analyzes the business flow based on the converted text data and acquired visual data and automatically generates business guidelines. This enables the automatic generation of business manuals and the automatic execution of business processes.

[0065] "Audio data" refers to information expressed through the medium of sound, represented in digital format.

[0066] "Text data" refers to character information represented in digital format, and includes information converted from audio data.

[0067] An "information processing device" is a device that receives data and performs processing such as analysis, conversion, and storage.

[0068] "Operating procedure" refers to a sequence of actions or instructions that a user takes to achieve a specific objective.

[0069] "Visual data" refers to image information, such as screen captures, represented in digital format.

[0070] A "business process flow" is a diagram that shows the sequence of procedures and steps involved in carrying out a business task.

[0071] A "work guidelines" is a document that outlines the necessary procedures and points to note when carrying out a task.

[0072] "Equipment" refers to process equipment or devices that have a specific function and are designed to perform that function.

[0073] An "intelligent construction model" is a model created by applying artificial intelligence technology to automate various tasks, including data analysis and prediction.

[0074] A "remote database" is a data storage location that is physically distant but accessible via a network such as the internet.

[0075] "Storage" refers to the process of securely preserving data and making it accessible for later use.

[0076] This invention provides a system that automatically generates work guidelines using voice and visual data to streamline operations, and then automatically executes those guidelines. The system is cloud-based and can be accessed and used by users through their own devices.

[0077] The user's terminal is equipped with a microphone and speech recognition software for acquiring voice input. The user verbally explains the business workflow, and the audio is acquired and converted by the terminal. Examples of speech recognition software include available speech recognition APIs. The terminal also monitors the user's actions and takes screen captures at specific times. This visual data is later sent to a server and contributes to the formation of the business workflow.

[0078] The server uses a dedicated information processing device to convert audio data into text data. This text data is analyzed using a generative AI model and forms the basis for a process that recognizes and extracts business flows. The server also analyzes this data in conjunction with acquired visual data to automatically generate detailed business guidelines. These guidelines visually reproduce the business flow and assist in understanding the business processes. The generated business guidelines are securely stored in a remote database in the cloud.

[0079] For example, if a user explains the expense reimbursement process verbally and performs the corresponding operations on their device, the server will generate expense reimbursement guidelines based on this data. These guidelines will detail everything from how to launch the software to the procedures for entering necessary information and the verification process.

[0080] In subsequent uses, when a user performs a similar task, the terminal will refer to the task guidelines stored on the server and automatically proceed with the task using a generated AI model. An example of a prompt message that can be entered is, "Please explain and instruct me on the expense reimbursement process using voice." This invention promotes the standardization and efficiency of tasks, enabling users to perform tasks quickly and accurately.

[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0082] Step 1:

[0083] The user explains the work procedure by voice using their device. During this process, voice data is input via the device's microphone. This voice data is passed to speech recognition software, which converts it into text data in real time. This converted text data is then prepared for transmission to the server.

[0084] Step 2:

[0085] The device monitors the user's screen activity in the background. When important operations or screen transitions occur, the device automatically takes a screen capture. This captured image is collected as the user's activity history and prepared as a dataset for transmission to the server.

[0086] Step 3:

[0087] The server receives audio data and screen captures sent from the terminal. The input audio data is analyzed again on the server using a high-precision speech recognition engine and converted into text data. The input screen captures are processed as information for visually analyzing the business flow. At this stage, the foundational data for the entire business process is complete.

[0088] Step 4:

[0089] The server applies a generating AI model using the converted text data and acquired screen captures. The model appropriately recognizes the business flow from the text and visual data and extracts detailed business guidelines. The output business guidelines include detailed descriptions of business procedures and corresponding images. These business guidelines are securely stored in a remote database in the cloud.

[0090] Step 5:

[0091] The next time the user performs a similar task, the user terminal will request the server to automatically execute the task via a prompt message. The server will refer to the task guidelines stored in the cloud, and the generated AI model will perform the task according to the automated instructions. This process allows the user to complete the task smoothly without any intervention.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] Improving work efficiency and standardization are crucial challenges, especially in manufacturing. Currently, work procedures are often performed manually based on manuals, which is time-consuming and labor-intensive, and can lead to variations in quality. Furthermore, processes that rely on worker experience increase the risk of errors and require training new employees. There is a need for solutions to these problems and ensure efficient and consistent work quality.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for acquiring voice input and converting the voice into text data, means for monitoring the user's information processing device and acquiring images at important moments, means for analyzing the operation procedure based on the converted text data and acquired images and automatically generating work instructions, and analysis means for analyzing the work procedure using an artificial structure and optimizing the work instructions. This enables the rapid automatic generation and optimization of work procedures.

[0097] "Voice input" refers to the process of acquiring a user's speech as a digital signal.

[0098] "Text data" refers to string information obtained by analyzing voice input.

[0099] An "information processing device" is an electronic device operated by a user to process and monitor data.

[0100] An "image" is a visual record that captures a specific scene on an information processing device.

[0101] An "operating procedure" is a set of steps necessary to complete a specific task or process.

[0102] A "work instruction sheet" is a guideline that specifically outlines the flow and procedures of a task.

[0103] An "artificial structure" is a digital architecture used for generation and analysis using AI models.

[0104] "Analysis means" refers to methods or techniques for analyzing data and processing it into meaningful information.

[0105] A "generative information processing model" is a framework of AI trained to process data and automate specific tasks.

[0106] To implement this invention, it is necessary to accurately acquire voice input and process that information as text data. Therefore, the server should utilize a speech recognition engine and use services such as Google® Speech-to-Text API. Voice input is captured by the terminal via a microphone and converted into text data. In addition, the user's information processing device should use a camera to acquire images at important scenes in the work procedure and an image processing library such as OpenCV.

[0107] Text and image data acquired by the information processing device are sent to a server. The server then uses an AI model to analyze this data and automatically generates standardized work instructions. These work instructions contain detailed descriptions of specific operating procedures and can be used for the automated execution of subsequent tasks.

[0108] Once the work instructions are generated, the server saves them to a data storage device, making them available via the cloud when requested by the user in the future. This process improves work efficiency and standardization, particularly facilitating the automation of robotic tasks within factories.

[0109] As a concrete example, considering the assembly of parts in a factory, the user can give a voice command such as "Screw part A into part B," and a camera simultaneously records the operation. The server can then use this data to enable automated assembly in the future. An example of a prompt message would be, "Please tell me the process of assembling the toy according to the following steps."

[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0111] Step 1:

[0112] The user inputs work instructions by voice into the device. The device receives this voice through its microphone and converts the voice data into text data using the Google Speech-to-Text API. The input is voice, and the output is text data.

[0113] Step 2:

[0114] The terminal monitors the user's actions and uses the camera to capture images at critical moments. Input is information about the user's actions, and output is the captured image. The image processing library OpenCV is used to capture the necessary scenes.

[0115] Step 3:

[0116] The terminal sends the acquired text data and captured images to the server. The input consists of text data and images, and the terminal performs the operation of transferring these together to the server.

[0117] Step 4:

[0118] The server uses a generative AI model to analyze the received text data and images. Through this analysis, it recognizes specific operating procedures and automatically generates work instructions. The input is text data and images, and the output is work instructions. This process involves structured analysis of the data.

[0119] Step 5:

[0120] The server saves the generated work instructions to an information storage device. The input is the work instruction data, and the output is the saved digital file. This process establishes the functionality necessary for future use.

[0121] Step 6:

[0122] The next time the user requests automated execution of a task, the server will refer to the saved work order and send data to the terminal to perform the requested operation. The inputs are the user's request and the saved work order, and the output is the automated operation. This supports the efficient execution of tasks.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention is a system that combines a voice input analysis function with an emotion engine that recognizes user emotions, in order to significantly improve work efficiency. In addition to automating work procedures, this system supports more flexible and effective work execution by understanding the user's emotional state in real time and adjusting work accordingly.

[0125] This system begins by acquiring voice input through a user terminal. Users input the details of their work as instructions via voice, and can also convey emotions and intentions included in those instructions. The terminal acquires this voice data in real time and sends it to the server.

[0126] The server uses a dedicated speech recognition engine to convert speech data into natural language text data. Simultaneously, an emotion engine analyzes the user's emotional state from the speech data. The results of this analysis function as a crucial element in generating business procedures and subsequent task execution.

[0127] The user's terminal monitors the user's screen operations and takes screen captures at specific stages of the operation. Based on these captures and audio data, the server uses a generated AI model to analyze the work procedure. The analyzed procedure will also include dynamic adjustments based on the user's emotions.

[0128] For example, when a user is preparing for an important meeting, if they include feelings of tension or stress while giving voice instructions, the emotion engine will detect that emotional state. This information will then be automatically incorporated into the work manual to provide additional assistance and tips to facilitate meeting preparation.

[0129] In the next task execution, automated execution will be possible in response to user requests. The server will use a generated AI model to automatically execute the procedures based on the stored task manual. During this process, the user's emotional state may change in real time, so the emotion engine will detect these changes and adjust the execution flow as needed.

[0130] This system allows users to enjoy optimal work processes tailored to their individual needs and emotions, contributing to improved work efficiency and a better work environment. As a result, it enables human-centered, flexible work design and smarter business operations.

[0131] The following describes the processing flow.

[0132] Step 1:

[0133] The user speaks work-related instructions and associated emotions into the device. The device captures this audio in real time and saves it as a digital audio file.

[0134] Step 2:

[0135] The device sends the acquired audio data to the server. The server uses a speech recognition engine to convert the audio data into natural language text data.

[0136] Step 3:

[0137] The server uses an emotion engine to analyze the user's emotional state from voice data and extracts emotional data. This emotional data is used when generating business procedures.

[0138] Step 4:

[0139] As the user performs their tasks, the terminal continuously monitors the user's screen operations and takes screen captures at critical operation steps.

[0140] Step 5:

[0141] The terminal sends the acquired screen capture data to the server. The server uses a generated AI model to analyze the work procedure based on the text data, screen captures, and sentiment data.

[0142] Step 6:

[0143] The server automatically generates dynamically adjusted operational manuals based on the analysis results. These manuals include emotionally responsive tips and assistance. The generated manuals are stored in the cloud.

[0144] Step 7:

[0145] From the next time onward, when a user requests automated task execution, the server will use the saved task manual as a basis, and the generated AI model will automatically execute the task according to the procedure.

[0146] Step 8:

[0147] During task execution, the server uses an emotion engine to monitor the user's real-time emotional changes and adjust the workflow as needed.

[0148] Step 9:

[0149] After completing a task, the terminal notifies the user of the results and provides relevant feedback if there have been any changes in emotions.

[0150] (Example 2)

[0151] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0152] Conventional business execution systems faced challenges in efficiently processing work instructions entered via user voice input and in dynamically adjusting tasks while considering the user's emotional state. This resulted in a lack of flexibility and efficiency in operations, and the need for additional confirmation by the user.

[0153] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0154] In this invention, the server includes means for converting speech into text information, means for analyzing the user's emotional state and dynamically adjusting the work operation procedure, and means for automatically executing tasks based on work instructions. This enables flexible and efficient work execution that takes the user's emotions into consideration.

[0155] "Voice input" refers to instructions or information that a user gives via voice, and is the data that the system acquires and processes from that voice.

[0156] "Textual information" refers to data in the form of a string of characters that has been converted from voice input through a speech recognition device.

[0157] "Display device operation" refers to a series of operations and controls performed by a user using a display device.

[0158] "Image information" refers to visual data obtained as screen captures of specific important scenes during a user's operation of their display device.

[0159] "Business operation procedures" refer to information that describes the steps and procedures necessary to perform a specific task.

[0160] A "work instruction sheet" is a document that is automatically generated based on work procedures and is used as a guide for carrying out work.

[0161] A "remote storage device" is a data storage device installed in an external facility such as the cloud, and is used to store and manage work instructions and related data.

[0162] "Emotional state" refers to the emotions a user experiences in relation to their work, as captured through analysis, and includes feelings such as joy, stress, and tension.

[0163] A "generative AI model" refers to an algorithm or process that automatically executes tasks based on work instructions, and is used to dynamically adjust tasks according to user requirements.

[0164] This invention is a voice input analysis system aimed at improving work efficiency, and it begins with the user giving work instructions via voice.

[0165] The user provides voice input through the device's microphone. This voice data is acquired by the device in real time. The device then sends the voice data to a server. The server uses speech recognition software to convert the voice data into natural language text. This process utilizes common speech recognition technologies.

[0166] The server uses an emotion analysis engine to analyze the user's emotional state from the converted text information. This is necessary to analyze emotions such as stress and joy. The emotion analysis engine determines the user's emotions by analyzing the tone of voice and word choice.

[0167] Based on the sentiment analysis results and the converted text information, the server uses a generative AI model to generate work operation procedures. These procedures incorporate dynamic adjustments based on the user's emotional state. The generated work instructions are stored in a remote storage device in the cloud.

[0168] For example, if the emotion analysis engine detects tension when a user provides voice input while giving instructions for meeting preparation, the generating AI model will use the prompt "Adjust the meeting procedure based on the user's instructions and emotion data" to generate a work procedure that includes helpful hints and reminders.

[0169] This process allows users to experience the flexible execution of tasks based on voice input and automated analysis. Work instructions that adjust according to emotional changes enable efficient and human-centered operations.

[0170] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0171] Step 1:

[0172] This is the stage where the terminal acquires voice input. The user speaks the necessary instructions for the task into the terminal's microphone. The terminal temporarily stores the acquired voice data and then prepares to send it to the server. In this step, voice data is received as input and the terminal enters a state of waiting to send to the server.

[0173] Step 2:

[0174] The terminal processes the acquired audio data and sends it to the server. The server receives the transmitted audio data and uses speech recognition software to convert it into natural language text. Here, the input is raw audio data, and the output is text information. A speech recognition algorithm is used for this conversion.

[0175] Step 3:

[0176] The server uses an emotion analysis engine to analyze the user's emotional state based on the converted text information. This analysis includes techniques to identify emotions such as stress and tension by evaluating word choice and voice characteristics. The input is the converted text information, and the output is the user's emotional state data.

[0177] Step 4:

[0178] The server uses emotional state data and textual information to input prompt statements into a generative AI model, which then generates business operation procedures. These prompt statements might be something like, "Create the optimal business procedure based on the user's instructions." The input for this step is emotional state data and textual information, and the output is business procedure data. The generative AI model then executes this process.

[0179] Step 5:

[0180] The server saves the generated work procedure data to remote storage. The saved work instructions are managed on the cloud and can be reused later. The input is work procedure data, and the output is saved work instructions.

[0181] Step 6:

[0182] For subsequent task executions, the server will automatically execute tasks based on user requests, referencing saved task instructions. A generative AI model is used here, enabling dynamic adjustments. The input is the saved task instructions, and the output is the task process to be executed. This allows users to enjoy flexible tasks that respond to their emotions.

[0183] (Application Example 2)

[0184] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0185] In today's work environment, there is a demand for both improved work efficiency and worker comfort. However, conventional systems have difficulty flexibly adjusting tasks while taking into account the emotional state of workers, which can result in increased worker stress and fatigue, potentially leading to decreased productivity.

[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0187] In this invention, the server includes means for acquiring voice input and converting the voice into text data, means for recognizing the user's emotions and dynamically adjusting work procedures based on those emotions, and means for analyzing the work procedures and optimizing work instructions based on the converted text data and emotion data. This enables optimal work adjustment according to the worker's emotional state and an automated, efficient work process.

[0188] "Voice input" is the process of capturing the sound emitted by the user as a digital signal into the system.

[0189] "Text data" refers to digital information that represents voice input as a string of characters.

[0190] "Recognizing emotions" means analyzing and identifying the emotional state of a user from their voice or input information.

[0191] "Business procedures" refer to a series of steps or processes required to perform a specific task.

[0192] "Dynamic adjustment" means changing the work content and processes in real time in response to changes in circumstances and conditions.

[0193] "Emotional data" refers to information that indicates a user's emotional state, obtained during the process of recognizing emotions.

[0194] "Optimizing work instructions" means restructuring work procedures and processes to maximize work efficiency and effectiveness.

[0195] "Storage means" refers to a method or apparatus for storing data or information over a long period of time.

[0196] "Control means" refers to technologies and devices used to operate and manage equipment and systems.

[0197] A "generative AI model" is a model that automatically creates algorithms and information from data based on machine learning technology.

[0198] This invention is a system that combines voice input and emotion recognition, and the system for realizing this includes the following program.

[0199] The server receives voice data from the user's terminal over the network to acquire voice input. This voice data is converted into text data using the Google Cloud Speech-to-Text API. Simultaneously, the server recognizes the user's emotions from the voice using the Microsoft® Azure® Emotion API. This emotion data is an important element for the user's work, and the server uses it to analyze and optimize work procedures. Generative AI models are used for analysis and optimization.

[0200] The generated work instructions are stored in the cloud and made available for download to the user when needed. This allows workers to obtain the optimal work procedures tailored to their emotional state at the time. When the user performs the work again, the server refers to the stored data and automatically executes the tasks. This process includes dynamic adjustments in response to changes in emotions, enabling work to proceed in real time while considering the user's feelings.

[0201] As a concrete example, consider a scenario where a factory worker uses smart glasses and gives a voice command such as "Pick up the part for the next process." In this case, the server analyzes the emotional data from the voice command and adjusts the work pace if fatigue is detected. It can also display a message on the screen saying, "Please take a break."

[0202] The generative AI model generates instructions in real time that respond to the user's emotions each time it receives a voice command. An example of a prompt message is: "When the user says 'remove the part for the next step,' analyze their emotions and, if they are feeling fatigued, generate a message to inform the user of this."

[0203] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0204] Step 1:

[0205] The user wears smart glasses and initiates work by voice input. The voice data is captured by the terminal and transmitted to the server via the network. The input is the user's voice, and the output is the transmission of voice data to the server.

[0206] Step 2:

[0207] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. Audio data is input to the server, and the output is the converted text data. Data processing is performed to analyze the audio waveform information and represent it as string information.

[0208] Step 3:

[0209] The server uses Microsoft Azure's Emotion API to diagnose emotions from voice data. The input is voice data, and the output is data indicating the user's emotional state. It performs data calculations to analyze the tone and intonation of the voice and extract emotional parameters.

[0210] Step 4:

[0211] The server uses a generative AI model with converted text data and sentiment data to optimize work instructions. The input is text data and sentiment data, and the output is optimized work instructions. The generative AI model is given prompt text and executes an algorithm that generates the optimal work procedure.

[0212] Step 5:

[0213] Optimized work instructions are stored on a cloud server and managed so that users can refer to them in the future. The input is work instruction data, and the output is the stored work instruction data.

[0214] Step 6:

[0215] When the user performs the task again, the server automatically executes the work process based on the saved work instructions. The instructions are dynamically adjusted according to the user's real-time sentiment data. The input is real-time sentiment data, and the output is the adjusted work instructions.

[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0217] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0219] [Second Embodiment]

[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0232] This invention provides a system that automatically generates work manuals using voice and screen operation data to streamline operations, and automatically executes tasks based on those manuals. This system is provided on a cloud-based platform, and users can access and use it through their own devices.

[0233] For the system to be implemented, it is a prerequisite that the user terminal is equipped with hardware (microphone) and software (voice recognition function) that enables voice input. When performing a task, the user will explain the procedure for which they wish to create a manual by voice. This explanation will be captured by the user's terminal and sent to the server.

[0234] The server analyzes the audio data using a dedicated speech recognition engine and converts it into text data. This text data forms the basis for the subsequent process of generating work manuals. Additionally, the user terminal monitors the screen as the user performs their tasks and takes screenshots at critical operation steps. This image information is also sent to the server.

[0235] The server analyzes the received text and screen capture data to construct work procedures. A generative AI model is used for the analysis process, automatically recognizing the procedures based on the data and generating a work manual based on the extracted information. The generated manual includes detailed procedures and corresponding images, and is securely stored in the cloud.

[0236] For example, if a user explains the expense reimbursement process verbally and then performs the actual operations on a terminal, the server will generate a manual for the expense reimbursement process based on this verbal explanation and operation. This manual will clearly explain everything from how to start the software to the procedure for entering the necessary information and the verification process.

[0237] In subsequent instances, when a user performs a similar task, their terminal will request automated execution of the task. The server will refer to the task manual stored in the cloud, and the generated AI model will automatically complete the task according to the procedure. As a result, task standardization will progress, and more efficient work execution will be possible.

[0238] The introduction of this system will reduce the need for manual manual creation and work handover, thereby optimizing human resources and improving productivity. This technology will particularly contribute to the effective use of resources in small and medium-sized enterprises and provides a new solution for business operations.

[0239] The following describes the processing flow.

[0240] Step 1:

[0241] The user speaks voice instructions related to their work into the terminal. The terminal acquires this voice in real time and temporarily saves it as a digital audio file.

[0242] Step 2:

[0243] The terminal sends the acquired voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. This converted text accurately reflects the work instructions given by the user.

[0244] Step 3:

[0245] When a user performs a task, the terminal monitors the user's screen operations. It then automatically captures a screen image when a specific operation trigger is detected. This capture contains important operational details necessary for creating procedure manuals.

[0246] Step 4:

[0247] The terminal sends the acquired screen capture data to the server. The server uses a generating AI model to analyze the text data and screen captures, and automatically extracts the components of the business procedure.

[0248] Step 5:

[0249] The server automatically generates an operational manual based on the extracted information. This manual combines screen captures corresponding to instructions with detailed operational procedures. The generated manual is securely stored in cloud storage.

[0250] Step 6:

[0251] From the next time onward, the user will request the same task to be performed automatically from the terminal. The server will read the task manual stored in the cloud, and the generated AI model will accurately reproduce the steps required for automated execution.

[0252] Step 7:

[0253] After the server completes its task, the terminal notifies the user of the results. If the process was successful, or if an error occurred, the user is informed of the details.

[0254] (Example 1)

[0255] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0256] Traditional methods of creating operational manuals and automating business processes were inefficient and required significant human resources. In particular, in small organizations, a lack of standardization and insufficient knowledge sharing led to decreased productivity. Furthermore, manual manual updates were infrequent, making it difficult to maintain consistency in work procedures.

[0257] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0258] In this invention, the server includes means equipped with an information processing device that acquires voice data and converts that data into text data; means equipped with a device that monitors the user's operation procedures and collects visual data for specific actions; and means equipped with a device that analyzes the business flow based on the converted text data and acquired visual data and automatically generates business guidelines. This enables the automatic generation of business manuals and the automatic execution of business processes.

[0259] "Audio data" refers to information expressed through the medium of sound, represented in digital format.

[0260] "Text data" refers to character information represented in digital format, and includes information converted from audio data.

[0261] An "information processing device" is a device that receives data and performs processing such as analysis, conversion, and storage.

[0262] "Operating procedure" refers to a sequence of actions or instructions that a user takes to achieve a specific objective.

[0263] "Visual data" refers to image information, such as screen captures, represented in digital format.

[0264] A "business process flow" is a diagram that shows the sequence of procedures and steps involved in carrying out a business task.

[0265] A "work guidelines" is a document that outlines the necessary procedures and points to note when carrying out a task.

[0266] "Equipment" refers to process equipment or devices that have a specific function and are designed to perform that function.

[0267] An "intelligent construction model" is a model created by applying artificial intelligence technology to automate various tasks, including data analysis and prediction.

[0268] A "remote database" is a data storage location that is physically distant but accessible via a network such as the internet.

[0269] "Storage" refers to the process of securely preserving data and making it accessible for later use.

[0270] This invention provides a system that automatically generates work guidelines using voice and visual data to streamline operations, and then automatically executes those guidelines. The system is cloud-based and can be accessed and used by users through their own devices.

[0271] The user's terminal is equipped with a microphone and speech recognition software for acquiring voice input. The user verbally explains the business workflow, and the audio is acquired and converted by the terminal. Examples of speech recognition software include available speech recognition APIs. The terminal also monitors the user's actions and takes screen captures at specific times. This visual data is later sent to a server and contributes to the formation of the business workflow.

[0272] The server uses a dedicated information processing device to convert audio data into text data. This text data is analyzed using a generative AI model and forms the basis for a process that recognizes and extracts business flows. The server also analyzes this data in conjunction with acquired visual data to automatically generate detailed business guidelines. These guidelines visually reproduce the business flow and assist in understanding the business processes. The generated business guidelines are securely stored in a remote database in the cloud.

[0273] For example, if a user explains the expense reimbursement process verbally and performs the corresponding operations on their device, the server will generate expense reimbursement guidelines based on this data. These guidelines will detail everything from how to launch the software to the procedures for entering necessary information and the verification process.

[0274] In subsequent uses, when a user performs a similar task, the terminal will refer to the task guidelines stored on the server and automatically proceed with the task using a generated AI model. An example of a prompt message that can be entered is, "Please explain and instruct me on the expense reimbursement process using voice." This invention promotes the standardization and efficiency of tasks, enabling users to perform tasks quickly and accurately.

[0275] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0276] Step 1:

[0277] The user uses their own terminal to verbally explain the business procedures. At this time, voice data is input through the microphone of the terminal. The voice data is passed to the voice recognition software and converted into text data in real time. This converted text data is prepared to be sent to the server.

[0278] Step 2:

[0279] The terminal monitors the user's screen operations in the background. When important operations or screen transitions occur, the terminal automatically captures the screen. This captured image is collected as the user's operation history and prepared as a dataset to be sent to the server.

[0280] Step 3:

[0281] The server receives the voice data and screen captures sent from the terminal. The input voice data is analyzed again by a high-precision voice recognition engine on the server and converted into text data. Also, the input screen captures are processed as information for visually analyzing the business process. At this stage, the basic data for the entire business is complete.

[0282] Step 4:

[0283] The server applies the generated AI model using the converted text data and the captured screen captures. The model appropriately recognizes the business process from the text data and visual data and extracts detailed business guidelines. The output business guidelines include a detailed description of the business procedures and corresponding images. These business guidelines are securely stored in the cloud remote database.

[0284] Step 5:

[0285] Next, when the user performs similar operations, the user terminal requests the server to automatically execute the operations through a prompt message. The server refers to the operation guidelines stored on the cloud, and the generated AI model performs the operations according to the automatic instructions. Through this process, the user can smoothly complete the operations without intervention.

[0286] (Application Example 1)

[0287] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0288] The improvement of operation efficiency and standardization is an important issue, especially in the manufacturing site. Currently, the operation procedures are often manually performed based on manuals, which is time-consuming and labor-intensive, and there may be variations in quality. Furthermore, processes that rely on the experience of workers increase the risk of new worker education and operation errors. There is a need for means to solve such problems and ensure efficient and stable operation quality.

[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0290] In this invention, the server includes means for acquiring voice input and converting the voice into text data, means for monitoring the user's information processing device and acquiring images in important scenes, means for analyzing operation procedures based on the converted text data and the acquired images and automatically generating work instructions, and analysis means for analyzing operation procedures using artificial structures and optimizing work instructions. This enables the rapid automatic generation and optimization of operation procedures.

[0291] "Voice input" refers to the process of acquiring the user's speech as a digital signal.

[0292] "Text data" refers to the string information obtained by analyzing voice input.

[0293] An "information processing device" is an electronic device operated by a user to process and monitor data.

[0294] An "image" is a visual record that captures a specific scene on an information processing device.

[0295] An "operating procedure" is a set of steps necessary to complete a specific task or process.

[0296] A "work instruction sheet" is a guideline that specifically outlines the flow and procedures of a task.

[0297] An "artificial structure" is a digital architecture used for generation and analysis using AI models.

[0298] "Analysis means" refers to methods or techniques for analyzing data and processing it into meaningful information.

[0299] A "generative information processing model" is a framework of AI trained to process data and automate specific tasks.

[0300] To implement this invention, it is necessary to accurately acquire voice input and process that information as text data. Therefore, the server should utilize a speech recognition engine and use a service such as the Google Speech-to-Text API. Voice input is captured by the terminal via a microphone and converted into text data. In addition, the user's information processing device should use a camera to acquire images at important scenes in the work procedure and an image processing library such as OpenCV.

[0301] The text data and image data acquired by the information processing device are sent to the server. Then, based on these data, the server analyzes the operation procedures using the generated AI model and automatically generates a standardized work instruction. This work instruction details specific operation procedures and can be utilized for the subsequent automatic execution of business operations.

[0302] When the generation of the work instruction is completed, the server saves it on the information storage device and makes it available via the cloud for use when requested by the user from the next time onwards. Through this process, the efficiency and standardization of work are realized, and in particular, the automation of robot operations within the factory is promoted.

[0303] As a specific example, considering component assembly within a factory, when the user gives an instruction such as "Screw component A into B" verbally and at the same time the camera records the operation, the server can enable automatic assembly from the next time onwards based on this. As an example of a prompt sentence, an instruction can be given in the form of "Please teach me the process of assembling a toy according to the following operation procedure."

[0304] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0305] Step 1:

[0306] The user inputs a work instruction verbally towards the terminal. The terminal acquires this voice through the microphone and uses the Google Speech-to-Text API to convert the voice data into text data. The input is voice, and the output is text data.

[0307] Step 2:

[0308] The terminal monitors the user's operation status and acquires an image using the camera in important scenes. The input is the status information of the operation, and the output is a captured image. The OpenCV, an image processing library, is used to perform the operation of capturing necessary scenes.

[0309] Step 3:

[0310] The terminal sends the acquired text data and captured images to the server. The input consists of text data and images, and the terminal performs the operation of transferring these together to the server.

[0311] Step 4:

[0312] The server uses a generative AI model to analyze the received text data and images. Through this analysis, it recognizes specific operating procedures and automatically generates work instructions. The input is text data and images, and the output is work instructions. This process involves structured analysis of the data.

[0313] Step 5:

[0314] The server saves the generated work instructions to an information storage device. The input is the work instruction data, and the output is the saved digital file. This process establishes the functionality necessary for future use.

[0315] Step 6:

[0316] The next time the user requests automated execution of a task, the server will refer to the saved work order and send data to the terminal to perform the requested operation. The inputs are the user's request and the saved work order, and the output is the automated operation. This supports the efficient execution of tasks.

[0317] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0318] This invention is a system that combines a voice input analysis function with an emotion engine that recognizes user emotions, in order to significantly improve work efficiency. In addition to automating work procedures, this system supports more flexible and effective work execution by understanding the user's emotional state in real time and adjusting work accordingly.

[0319] This system begins by acquiring voice input through a user terminal. Users input the details of their work as instructions via voice, and can also convey emotions and intentions included in those instructions. The terminal acquires this voice data in real time and sends it to the server.

[0320] The server uses a dedicated speech recognition engine to convert speech data into natural language text data. Simultaneously, an emotion engine analyzes the user's emotional state from the speech data. The results of this analysis function as a crucial element in generating business procedures and subsequent task execution.

[0321] The user's terminal monitors the user's screen operations and takes screen captures at specific stages of the operation. Based on these captures and audio data, the server uses a generated AI model to analyze the work procedure. The analyzed procedure will also include dynamic adjustments based on the user's emotions.

[0322] For example, when a user is preparing for an important meeting, if they include feelings of tension or stress while giving voice instructions, the emotion engine will detect that emotional state. This information will then be automatically incorporated into the work manual to provide additional assistance and tips to facilitate meeting preparation.

[0323] In the next task execution, automated execution will be possible in response to user requests. The server will use a generated AI model to automatically execute the procedures based on the stored task manual. During this process, the user's emotional state may change in real time, so the emotion engine will detect these changes and adjust the execution flow as needed.

[0324] This system allows users to enjoy optimal work processes tailored to their individual needs and emotions, contributing to improved work efficiency and a better work environment. As a result, it enables human-centered, flexible work design and smarter business operations.

[0325] The following describes the processing flow.

[0326] Step 1:

[0327] The user speaks work-related instructions and associated emotions into the device. The device captures this audio in real time and saves it as a digital audio file.

[0328] Step 2:

[0329] The device sends the acquired audio data to the server. The server uses a speech recognition engine to convert the audio data into natural language text data.

[0330] Step 3:

[0331] The server uses an emotion engine to analyze the user's emotional state from voice data and extracts emotional data. This emotional data is used when generating business procedures.

[0332] Step 4:

[0333] As the user performs their tasks, the terminal continuously monitors the user's screen operations and takes screen captures at critical operation steps.

[0334] Step 5:

[0335] The terminal sends the acquired screen capture data to the server. The server uses a generated AI model to analyze the work procedure based on the text data, screen captures, and sentiment data.

[0336] Step 6:

[0337] The server automatically generates dynamically adjusted operational manuals based on the analysis results. These manuals include emotionally responsive tips and assistance. The generated manuals are stored in the cloud.

[0338] Step 7:

[0339] From the next time onward, when a user requests automated task execution, the server will use the saved task manual as a basis, and the generated AI model will automatically execute the task according to the procedure.

[0340] Step 8:

[0341] During task execution, the server uses an emotion engine to monitor the user's real-time emotional changes and adjust the workflow as needed.

[0342] Step 9:

[0343] After completing a task, the terminal notifies the user of the results and provides relevant feedback if there have been any changes in emotions.

[0344] (Example 2)

[0345] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0346] Conventional business execution systems faced challenges in efficiently processing work instructions entered via user voice input and in dynamically adjusting tasks while considering the user's emotional state. This resulted in a lack of flexibility and efficiency in operations, and the need for additional confirmation by the user.

[0347] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0348] In this invention, the server includes means for converting speech into text information, means for analyzing the user's emotional state and dynamically adjusting the work operation procedure, and means for automatically executing tasks based on work instructions. This enables flexible and efficient work execution that takes the user's emotions into consideration.

[0349] "Voice input" refers to instructions or information that a user gives via voice, and is the data that the system acquires and processes from that voice.

[0350] "Textual information" refers to data in the form of a string of characters that has been converted from voice input through a speech recognition device.

[0351] "Display device operation" refers to a series of operations and controls performed by a user using a display device.

[0352] "Image information" refers to visual data obtained as screen captures of specific important scenes during a user's operation of their display device.

[0353] "Business operation procedures" refer to information that describes the steps and procedures necessary to perform a specific task.

[0354] A "work instruction sheet" is a document that is automatically generated based on work procedures and is used as a guide for carrying out work.

[0355] A "remote storage device" is a data storage device installed in an external facility such as the cloud, and is used to store and manage work instructions and related data.

[0356] "Emotional state" refers to the emotions a user experiences in relation to their work, as captured through analysis, and includes feelings such as joy, stress, and tension.

[0357] A "generative AI model" refers to an algorithm or process that automatically executes tasks based on work instructions, and is used to dynamically adjust tasks according to user requirements.

[0358] This invention is a voice input analysis system aimed at improving work efficiency, and it begins with the user giving work instructions via voice.

[0359] The user provides voice input through the device's microphone. This voice data is acquired by the device in real time. The device then sends the voice data to a server. The server uses speech recognition software to convert the voice data into natural language text. This process utilizes common speech recognition technologies.

[0360] The server uses an emotion analysis engine to analyze the user's emotional state from the converted text information. This is necessary to analyze emotions such as stress and joy. The emotion analysis engine determines the user's emotions by analyzing the tone of voice and word choice.

[0361] Based on the sentiment analysis results and the converted text information, the server uses a generative AI model to generate work operation procedures. These procedures incorporate dynamic adjustments based on the user's emotional state. The generated work instructions are stored in a remote storage device in the cloud.

[0362] For example, if the emotion analysis engine detects tension when a user provides voice input while giving instructions for meeting preparation, the generating AI model will use the prompt "Adjust the meeting procedure based on the user's instructions and emotion data" to generate a work procedure that includes helpful hints and reminders.

[0363] This process allows users to experience the flexible execution of tasks based on voice input and automated analysis. Work instructions that adjust according to emotional changes enable efficient and human-centered operations.

[0364] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0365] Step 1:

[0366] This is the stage where the terminal acquires voice input. The user speaks the necessary instructions for the task into the terminal's microphone. The terminal temporarily stores the acquired voice data and then prepares to send it to the server. In this step, voice data is received as input and the terminal enters a state of waiting to send to the server.

[0367] Step 2:

[0368] The terminal processes the acquired audio data and sends it to the server. The server receives the transmitted audio data and uses speech recognition software to convert it into natural language text. Here, the input is raw audio data, and the output is text information. A speech recognition algorithm is used for this conversion.

[0369] Step 3:

[0370] The server uses an emotion analysis engine to analyze the user's emotional state based on the converted text information. This analysis includes techniques to identify emotions such as stress and tension by evaluating word choice and voice characteristics. The input is the converted text information, and the output is the user's emotional state data.

[0371] Step 4:

[0372] The server uses emotional state data and textual information to input prompt statements into a generative AI model, which then generates business operation procedures. These prompt statements might be something like, "Create the optimal business procedure based on the user's instructions." The input for this step is emotional state data and textual information, and the output is business procedure data. The generative AI model then executes this process.

[0373] Step 5:

[0374] The server saves the generated work procedure data to remote storage. The saved work instructions are managed on the cloud and can be reused later. The input is work procedure data, and the output is saved work instructions.

[0375] Step 6:

[0376] For subsequent task executions, the server will automatically execute tasks based on user requests, referencing saved task instructions. A generative AI model is used here, enabling dynamic adjustments. The input is the saved task instructions, and the output is the task process to be executed. This allows users to enjoy flexible tasks that respond to their emotions.

[0377] (Application Example 2)

[0378] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0379] In today's work environment, there is a demand for both improved work efficiency and worker comfort. However, conventional systems have difficulty flexibly adjusting tasks while taking into account the emotional state of workers, which can result in increased worker stress and fatigue, potentially leading to decreased productivity.

[0380] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0381] In this invention, the server includes means for acquiring voice input and converting the voice into text data, means for recognizing the user's emotions and dynamically adjusting work procedures based on those emotions, and means for analyzing the work procedures and optimizing work instructions based on the converted text data and emotion data. This enables optimal work adjustment according to the worker's emotional state and an automated, efficient work process.

[0382] "Voice input" is the process of capturing the sound emitted by the user as a digital signal into the system.

[0383] "Text data" refers to digital information that represents voice input as a string of characters.

[0384] "Recognizing emotions" means analyzing and identifying the emotional state of a user from their voice or input information.

[0385] "Business procedures" refer to a series of steps or processes required to perform a specific task.

[0386] "Dynamic adjustment" means changing the work content and processes in real time in response to changes in circumstances and conditions.

[0387] "Emotional data" refers to information that indicates a user's emotional state, obtained during the process of recognizing emotions.

[0388] "Optimizing work instructions" means restructuring work procedures and processes to maximize work efficiency and effectiveness.

[0389] "Storage means" refers to a method or apparatus for storing data or information over a long period of time.

[0390] "Control means" refers to technologies and devices used to operate and manage equipment and systems.

[0391] A "generative AI model" is a model that automatically creates algorithms and information from data based on machine learning technology.

[0392] This invention is a system that combines voice input and emotion recognition, and the system for realizing this includes the following program.

[0393] The server receives voice data from the user's terminal over the network to acquire voice input. This voice data is converted into text data using the Google Cloud Speech-to-Text API. Simultaneously, the server recognizes the user's emotions from the voice using the Microsoft Azure Emotion API. This emotion data is an important element for the user's work, and the server uses it to analyze and optimize work procedures. Generative AI models are used for analysis and optimization.

[0394] The generated work instructions are stored in the cloud and made available for download to the user when needed. This allows workers to obtain the optimal work procedures tailored to their emotional state at the time. When the user performs the work again, the server refers to the stored data and automatically executes the tasks. This process includes dynamic adjustments in response to changes in emotions, enabling work to proceed in real time while considering the user's feelings.

[0395] As a concrete example, consider a scenario where a factory worker uses smart glasses and gives a voice command such as "Pick up the part for the next process." In this case, the server analyzes the emotional data from the voice command and adjusts the work pace if fatigue is detected. It can also display a message on the screen saying, "Please take a break."

[0396] The generative AI model generates instructions in real time that respond to the user's emotions each time it receives a voice command. An example of a prompt message is: "When the user says 'remove the part for the next step,' analyze their emotions and, if they are feeling fatigued, generate a message to inform the user of this."

[0397] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0398] Step 1:

[0399] The user wears smart glasses and initiates work by voice input. The voice data is captured by the terminal and transmitted to the server via the network. The input is the user's voice, and the output is the transmission of voice data to the server.

[0400] Step 2:

[0401] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. Audio data is input to the server, and the output is the converted text data. Data processing is performed to analyze the audio waveform information and represent it as string information.

[0402] Step 3:

[0403] The server uses Microsoft Azure's Emotion API to diagnose emotions from voice data. The input is voice data, and the output is data indicating the user's emotional state. It performs data calculations to analyze the tone and intonation of the voice and extract emotional parameters.

[0404] Step 4:

[0405] The server uses a generative AI model with converted text data and sentiment data to optimize work instructions. The input is text data and sentiment data, and the output is optimized work instructions. The generative AI model is given prompt text and executes an algorithm that generates the optimal work procedure.

[0406] Step 5:

[0407] Optimized work instructions are stored on a cloud server and managed so that users can refer to them in the future. The input is work instruction data, and the output is the stored work instruction data.

[0408] Step 6:

[0409] When the user performs the task again, the server automatically executes the work process based on the saved work instructions. The instructions are dynamically adjusted according to the user's real-time sentiment data. The input is real-time sentiment data, and the output is the adjusted work instructions.

[0410] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0411] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0412] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0413] [Third Embodiment]

[0414] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0415] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0416] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0417] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0418] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0420] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0421] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0422] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0423] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0424] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0425] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0426] This invention provides a system that automatically generates work manuals using voice and screen operation data to streamline operations, and automatically executes tasks based on those manuals. This system is provided on a cloud-based platform, and users can access and use it through their own devices.

[0427] For the system to be implemented, it is a prerequisite that the user terminal is equipped with hardware (microphone) and software (voice recognition function) that enables voice input. When performing a task, the user will explain the procedure for which they wish to create a manual by voice. This explanation will be captured by the user's terminal and sent to the server.

[0428] The server analyzes the audio data using a dedicated speech recognition engine and converts it into text data. This text data forms the basis for the subsequent process of generating work manuals. Additionally, the user terminal monitors the screen as the user performs their tasks and takes screenshots at critical operation steps. This image information is also sent to the server.

[0429] The server analyzes the received text and screen capture data to construct work procedures. A generative AI model is used for the analysis process, automatically recognizing the procedures based on the data and generating a work manual based on the extracted information. The generated manual includes detailed procedures and corresponding images, and is securely stored in the cloud.

[0430] For example, if a user explains the expense reimbursement process verbally and then performs the actual operations on a terminal, the server will generate a manual for the expense reimbursement process based on this verbal explanation and operation. This manual will clearly explain everything from how to start the software to the procedure for entering the necessary information and the verification process.

[0431] In subsequent instances, when a user performs a similar task, their terminal will request automated execution of the task. The server will refer to the task manual stored in the cloud, and the generated AI model will automatically complete the task according to the procedure. As a result, task standardization will progress, and more efficient work execution will be possible.

[0432] The introduction of this system will reduce the need for manual manual creation and work handover, thereby optimizing human resources and improving productivity. This technology will particularly contribute to the effective use of resources in small and medium-sized enterprises and provides a new solution for business operations.

[0433] The following describes the processing flow.

[0434] Step 1:

[0435] The user speaks voice instructions related to their work into the terminal. The terminal acquires this voice in real time and temporarily saves it as a digital audio file.

[0436] Step 2:

[0437] The terminal sends the acquired voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. This converted text accurately reflects the work instructions given by the user.

[0438] Step 3:

[0439] When a user performs a task, the terminal monitors the user's screen operations. It then automatically captures a screen image when a specific operation trigger is detected. This capture contains important operational details necessary for creating procedure manuals.

[0440] Step 4:

[0441] The terminal sends the acquired screen capture data to the server. The server uses a generating AI model to analyze the text data and screen captures, and automatically extracts the components of the business procedure.

[0442] Step 5:

[0443] The server automatically generates an operational manual based on the extracted information. This manual combines screen captures corresponding to instructions with detailed operational procedures. The generated manual is securely stored in cloud storage.

[0444] Step 6:

[0445] From the next time onward, the user will request the same task to be performed automatically from the terminal. The server will read the task manual stored in the cloud, and the generated AI model will accurately reproduce the steps required for automated execution.

[0446] Step 7:

[0447] After the server completes its task, the terminal notifies the user of the results. If the process was successful, or if an error occurred, the user is informed of the details.

[0448] (Example 1)

[0449] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0450] Traditional methods of creating operational manuals and automating business processes were inefficient and required significant human resources. In particular, in small organizations, a lack of standardization and insufficient knowledge sharing led to decreased productivity. Furthermore, manual manual updates were infrequent, making it difficult to maintain consistency in work procedures.

[0451] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0452] In this invention, the server includes means equipped with an information processing device that acquires voice data and converts that data into text data; means equipped with a device that monitors the user's operation procedures and collects visual data for specific actions; and means equipped with a device that analyzes the business flow based on the converted text data and acquired visual data and automatically generates business guidelines. This enables the automatic generation of business manuals and the automatic execution of business processes.

[0453] "Audio data" refers to information expressed through the medium of sound, represented in digital format.

[0454] "Text data" refers to character information represented in digital format, and includes information converted from audio data.

[0455] An "information processing device" is a device that receives data and performs processing such as analysis, conversion, and storage.

[0456] "Operating procedure" refers to a sequence of actions or instructions that a user takes to achieve a specific objective.

[0457] "Visual data" refers to image information, such as screen captures, represented in digital format.

[0458] A "business process flow" is a diagram that shows the sequence of procedures and steps involved in carrying out a business task.

[0459] A "work guidelines" is a document that outlines the necessary procedures and points to note when carrying out a task.

[0460] "Equipment" refers to process equipment or devices that have a specific function and are designed to perform that function.

[0461] An "intelligent construction model" is a model created by applying artificial intelligence technology to automate various tasks, including data analysis and prediction.

[0462] A "remote database" is a data storage location that is physically distant but accessible via a network such as the internet.

[0463] "Storage" refers to the process of securely preserving data and making it accessible for later use.

[0464] This invention provides a system that automatically generates work guidelines using voice and visual data to streamline operations, and then automatically executes those guidelines. The system is cloud-based and can be accessed and used by users through their own devices.

[0465] The user's terminal is equipped with a microphone and speech recognition software for acquiring voice input. The user verbally explains the business workflow, and the audio is acquired and converted by the terminal. Examples of speech recognition software include available speech recognition APIs. The terminal also monitors the user's actions and takes screen captures at specific times. This visual data is later sent to a server and contributes to the formation of the business workflow.

[0466] The server uses a dedicated information processing device to convert audio data into text data. This text data is analyzed using a generative AI model and forms the basis for a process that recognizes and extracts business flows. The server also analyzes this data in conjunction with acquired visual data to automatically generate detailed business guidelines. These guidelines visually reproduce the business flow and assist in understanding the business processes. The generated business guidelines are securely stored in a remote database in the cloud.

[0467] For example, if a user explains the expense reimbursement process verbally and performs the corresponding operations on their device, the server will generate expense reimbursement guidelines based on this data. These guidelines will detail everything from how to launch the software to the procedures for entering necessary information and the verification process.

[0468] In subsequent uses, when a user performs a similar task, the terminal will refer to the task guidelines stored on the server and automatically proceed with the task using a generated AI model. An example of a prompt message that can be entered is, "Please explain and instruct me on the expense reimbursement process using voice." This invention promotes the standardization and efficiency of tasks, enabling users to perform tasks quickly and accurately.

[0469] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0470] Step 1:

[0471] The user explains the work procedure by voice using their device. During this process, voice data is input via the device's microphone. This voice data is passed to speech recognition software, which converts it into text data in real time. This converted text data is then prepared for transmission to the server.

[0472] Step 2:

[0473] The device monitors the user's screen activity in the background. When important operations or screen transitions occur, the device automatically takes a screen capture. This captured image is collected as the user's activity history and prepared as a dataset for transmission to the server.

[0474] Step 3:

[0475] The server receives audio data and screen captures sent from the terminal. The input audio data is analyzed again on the server using a high-precision speech recognition engine and converted into text data. The input screen captures are processed as information for visually analyzing the business flow. At this stage, the foundational data for the entire business process is complete.

[0476] Step 4:

[0477] The server applies a generating AI model using the converted text data and acquired screen captures. The model appropriately recognizes the business flow from the text and visual data and extracts detailed business guidelines. The output business guidelines include detailed descriptions of business procedures and corresponding images. These business guidelines are securely stored in a remote database in the cloud.

[0478] Step 5:

[0479] The next time the user performs a similar task, the user terminal will request the server to automatically execute the task via a prompt message. The server will refer to the task guidelines stored in the cloud, and the generated AI model will perform the task according to the automated instructions. This process allows the user to complete the task smoothly without any intervention.

[0480] (Application Example 1)

[0481] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0482] Improving work efficiency and standardization are crucial challenges, especially in manufacturing. Currently, work procedures are often performed manually based on manuals, which is time-consuming and labor-intensive, and can lead to variations in quality. Furthermore, processes that rely on worker experience increase the risk of errors and require training new employees. There is a need for solutions to these problems and ensure efficient and consistent work quality.

[0483] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0484] In this invention, the server includes means for acquiring voice input and converting the voice into text data, means for monitoring the user's information processing device and acquiring images at important moments, means for analyzing the operation procedure based on the converted text data and acquired images and automatically generating work instructions, and analysis means for analyzing the work procedure using an artificial structure and optimizing the work instructions. This enables the rapid automatic generation and optimization of work procedures.

[0485] "Voice input" refers to the process of acquiring a user's speech as a digital signal.

[0486] "Text data" refers to string information obtained by analyzing voice input.

[0487] An "information processing device" is an electronic device operated by a user to process and monitor data.

[0488] An "image" is a visual record that captures a specific scene on an information processing device.

[0489] An "operating procedure" is a set of steps necessary to complete a specific task or process.

[0490] A "work instruction sheet" is a guideline that specifically outlines the flow and procedures of a task.

[0491] An "artificial structure" is a digital architecture used for generation and analysis using AI models.

[0492] "Analysis means" refers to methods or techniques for analyzing data and processing it into meaningful information.

[0493] A "generative information processing model" is a framework of AI trained to process data and automate specific tasks.

[0494] To implement this invention, it is necessary to accurately acquire voice input and process that information as text data. Therefore, the server should utilize a speech recognition engine and use a service such as the Google Speech-to-Text API. Voice input is captured by the terminal via a microphone and converted into text data. In addition, the user's information processing device should use a camera to acquire images at important scenes in the work procedure and an image processing library such as OpenCV.

[0495] Text and image data acquired by the information processing device are sent to a server. The server then uses an AI model to analyze this data and automatically generates standardized work instructions. These work instructions contain detailed descriptions of specific operating procedures and can be used for the automated execution of subsequent tasks.

[0496] Once the work instructions are generated, the server saves them to a data storage device, making them available via the cloud when requested by the user in the future. This process improves work efficiency and standardization, particularly facilitating the automation of robotic tasks within factories.

[0497] As a concrete example, considering the assembly of parts in a factory, the user can give a voice command such as "Screw part A into part B," and a camera simultaneously records the operation. The server can then use this data to enable automated assembly in the future. An example of a prompt message would be, "Please tell me the process of assembling the toy according to the following steps."

[0498] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0499] Step 1:

[0500] The user inputs work instructions by voice into the device. The device receives this voice through its microphone and converts the voice data into text data using the Google Speech-to-Text API. The input is voice, and the output is text data.

[0501] Step 2:

[0502] The terminal monitors the user's actions and uses the camera to capture images at critical moments. Input is information about the user's actions, and output is the captured image. The image processing library OpenCV is used to capture the necessary scenes.

[0503] Step 3:

[0504] The terminal sends the acquired text data and captured images to the server. The input consists of text data and images, and the terminal performs the operation of transferring these together to the server.

[0505] Step 4:

[0506] The server uses a generative AI model to analyze the received text data and images. Through this analysis, it recognizes specific operating procedures and automatically generates work instructions. The input is text data and images, and the output is work instructions. This process involves structured analysis of the data.

[0507] Step 5:

[0508] The server saves the generated work instructions to an information storage device. The input is the work instruction data, and the output is the saved digital file. This process establishes the functionality necessary for future use.

[0509] Step 6:

[0510] The next time the user requests automated execution of a task, the server will refer to the saved work order and send data to the terminal to perform the requested operation. The inputs are the user's request and the saved work order, and the output is the automated operation. This supports the efficient execution of tasks.

[0511] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0512] This invention is a system that combines a voice input analysis function with an emotion engine that recognizes user emotions, in order to significantly improve work efficiency. In addition to automating work procedures, this system supports more flexible and effective work execution by understanding the user's emotional state in real time and adjusting work accordingly.

[0513] This system begins by acquiring voice input through a user terminal. Users input the details of their work as instructions via voice, and can also convey emotions and intentions included in those instructions. The terminal acquires this voice data in real time and sends it to the server.

[0514] The server uses a dedicated speech recognition engine to convert speech data into natural language text data. Simultaneously, an emotion engine analyzes the user's emotional state from the speech data. The results of this analysis function as a crucial element in generating business procedures and subsequent task execution.

[0515] The user's terminal monitors the user's screen operations and takes screen captures at specific stages of the operation. Based on these captures and audio data, the server uses a generated AI model to analyze the work procedure. The analyzed procedure will also include dynamic adjustments based on the user's emotions.

[0516] For example, when a user is preparing for an important meeting, if they include feelings of tension or stress while giving voice instructions, the emotion engine will detect that emotional state. This information will then be automatically incorporated into the work manual to provide additional assistance and tips to facilitate meeting preparation.

[0517] In the next task execution, automated execution will be possible in response to user requests. The server will use a generated AI model to automatically execute the procedures based on the stored task manual. During this process, the user's emotional state may change in real time, so the emotion engine will detect these changes and adjust the execution flow as needed.

[0518] This system allows users to enjoy optimal work processes tailored to their individual needs and emotions, contributing to improved work efficiency and a better work environment. As a result, it enables human-centered, flexible work design and smarter business operations.

[0519] The following describes the processing flow.

[0520] Step 1:

[0521] The user speaks work-related instructions and associated emotions into the device. The device captures this audio in real time and saves it as a digital audio file.

[0522] Step 2:

[0523] The device sends the acquired audio data to the server. The server uses a speech recognition engine to convert the audio data into natural language text data.

[0524] Step 3:

[0525] The server uses an emotion engine to analyze the user's emotional state from voice data and extracts emotional data. This emotional data is used when generating business procedures.

[0526] Step 4:

[0527] As the user performs their tasks, the terminal continuously monitors the user's screen operations and takes screen captures at critical operation steps.

[0528] Step 5:

[0529] The terminal sends the acquired screen capture data to the server. The server uses a generated AI model to analyze the work procedure based on the text data, screen captures, and sentiment data.

[0530] Step 6:

[0531] The server automatically generates dynamically adjusted operational manuals based on the analysis results. These manuals include emotionally responsive tips and assistance. The generated manuals are stored in the cloud.

[0532] Step 7:

[0533] From the next time onward, when a user requests automated task execution, the server will use the saved task manual as a basis, and the generated AI model will automatically execute the task according to the procedure.

[0534] Step 8:

[0535] During task execution, the server uses an emotion engine to monitor the user's real-time emotional changes and adjust the workflow as needed.

[0536] Step 9:

[0537] After completing a task, the terminal notifies the user of the results and provides relevant feedback if there have been any changes in emotions.

[0538] (Example 2)

[0539] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0540] Conventional business execution systems faced challenges in efficiently processing work instructions entered via user voice input and in dynamically adjusting tasks while considering the user's emotional state. This resulted in a lack of flexibility and efficiency in operations, and the need for additional confirmation by the user.

[0541] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0542] In this invention, the server includes means for converting speech into text information, means for analyzing the user's emotional state and dynamically adjusting the work operation procedure, and means for automatically executing tasks based on work instructions. This enables flexible and efficient work execution that takes the user's emotions into consideration.

[0543] "Voice input" refers to instructions or information that a user gives via voice, and is the data that the system acquires and processes from that voice.

[0544] "Textual information" refers to data in the form of a string of characters that has been converted from voice input through a speech recognition device.

[0545] "Display device operation" refers to a series of operations and controls performed by a user using a display device.

[0546] "Image information" refers to visual data obtained as screen captures of specific important scenes during a user's operation of their display device.

[0547] "Business operation procedures" refer to information that describes the steps and procedures necessary to perform a specific task.

[0548] A "work instruction sheet" is a document that is automatically generated based on work procedures and is used as a guide for carrying out work.

[0549] A "remote storage device" is a data storage device installed in an external facility such as the cloud, and is used to store and manage work instructions and related data.

[0550] "Emotional state" refers to the emotions a user experiences in relation to their work, as captured through analysis, and includes feelings such as joy, stress, and tension.

[0551] A "generative AI model" refers to an algorithm or process that automatically executes tasks based on work instructions, and is used to dynamically adjust tasks according to user requirements.

[0552] This invention is a voice input analysis system aimed at improving work efficiency, and it begins with the user giving work instructions via voice.

[0553] The user provides voice input through the device's microphone. This voice data is acquired by the device in real time. The device then sends the voice data to a server. The server uses speech recognition software to convert the voice data into natural language text. This process utilizes common speech recognition technologies.

[0554] The server uses an emotion analysis engine to analyze the user's emotional state from the converted text information. This is necessary to analyze emotions such as stress and joy. The emotion analysis engine determines the user's emotions by analyzing the tone of voice and word choice.

[0555] Based on the sentiment analysis results and the converted text information, the server uses a generative AI model to generate work operation procedures. These procedures incorporate dynamic adjustments based on the user's emotional state. The generated work instructions are stored in a remote storage device in the cloud.

[0556] For example, if the emotion analysis engine detects tension when a user provides voice input while giving instructions for meeting preparation, the generating AI model will use the prompt "Adjust the meeting procedure based on the user's instructions and emotion data" to generate a work procedure that includes helpful hints and reminders.

[0557] This process allows users to experience the flexible execution of tasks based on voice input and automated analysis. Work instructions that adjust according to emotional changes enable efficient and human-centered operations.

[0558] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0559] Step 1:

[0560] This is the stage where the terminal acquires voice input. The user speaks the necessary instructions for the task into the terminal's microphone. The terminal temporarily stores the acquired voice data and then prepares to send it to the server. In this step, voice data is received as input and the terminal enters a state of waiting to send to the server.

[0561] Step 2:

[0562] The terminal processes the acquired audio data and sends it to the server. The server receives the transmitted audio data and uses speech recognition software to convert it into natural language text. Here, the input is raw audio data, and the output is text information. A speech recognition algorithm is used for this conversion.

[0563] Step 3:

[0564] The server uses an emotion analysis engine to analyze the user's emotional state based on the converted text information. This analysis includes techniques to identify emotions such as stress and tension by evaluating word choice and voice characteristics. The input is the converted text information, and the output is the user's emotional state data.

[0565] Step 4:

[0566] The server uses emotional state data and textual information to input prompt statements into a generative AI model, which then generates business operation procedures. These prompt statements might be something like, "Create the optimal business procedure based on the user's instructions." The input for this step is emotional state data and textual information, and the output is business procedure data. The generative AI model then executes this process.

[0567] Step 5:

[0568] The server saves the generated work procedure data to remote storage. The saved work instructions are managed on the cloud and can be reused later. The input is work procedure data, and the output is saved work instructions.

[0569] Step 6:

[0570] For subsequent task executions, the server will automatically execute tasks based on user requests, referencing saved task instructions. A generative AI model is used here, enabling dynamic adjustments. The input is the saved task instructions, and the output is the task process to be executed. This allows users to enjoy flexible tasks that respond to their emotions.

[0571] (Application Example 2)

[0572] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0573] In today's work environment, there is a demand for both improved work efficiency and worker comfort. However, conventional systems have difficulty flexibly adjusting tasks while taking into account the emotional state of workers, which can result in increased worker stress and fatigue, potentially leading to decreased productivity.

[0574] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0575] In this invention, the server includes means for acquiring voice input and converting the voice into text data, means for recognizing the user's emotions and dynamically adjusting work procedures based on those emotions, and means for analyzing the work procedures and optimizing work instructions based on the converted text data and emotion data. This enables optimal work adjustment according to the worker's emotional state and an automated, efficient work process.

[0576] "Voice input" is the process of capturing the sound emitted by the user as a digital signal into the system.

[0577] "Text data" refers to digital information that represents voice input as a string of characters.

[0578] "Recognizing emotions" means analyzing and identifying the emotional state of a user from their voice or input information.

[0579] "Business procedures" refer to a series of steps or processes required to perform a specific task.

[0580] "Dynamic adjustment" means changing the work content and processes in real time in response to changes in circumstances and conditions.

[0581] "Emotional data" refers to information that indicates a user's emotional state, obtained during the process of recognizing emotions.

[0582] "Optimizing work instructions" means restructuring work procedures and processes to maximize work efficiency and effectiveness.

[0583] "Storage means" refers to a method or apparatus for storing data or information over a long period of time.

[0584] "Control means" refers to technologies and devices used to operate and manage equipment and systems.

[0585] A "generative AI model" is a model that automatically creates algorithms and information from data based on machine learning technology.

[0586] This invention is a system that combines voice input and emotion recognition, and the system for realizing this includes the following program.

[0587] The server receives voice data from the user's terminal over the network to acquire voice input. This voice data is converted into text data using the Google Cloud Speech-to-Text API. Simultaneously, the server recognizes the user's emotions from the voice using the Microsoft Azure Emotion API. This emotion data is an important element for the user's work, and the server uses it to analyze and optimize work procedures. Generative AI models are used for analysis and optimization.

[0588] The generated work instructions are stored in the cloud and made available for download to the user when needed. This allows workers to obtain the optimal work procedures tailored to their emotional state at the time. When the user performs the work again, the server refers to the stored data and automatically executes the tasks. This process includes dynamic adjustments in response to changes in emotions, enabling work to proceed in real time while considering the user's feelings.

[0589] As a concrete example, consider a scenario where a factory worker uses smart glasses and gives a voice command such as "Pick up the part for the next process." In this case, the server analyzes the emotional data from the voice command and adjusts the work pace if fatigue is detected. It can also display a message on the screen saying, "Please take a break."

[0590] The generative AI model generates instructions in real time that respond to the user's emotions each time it receives a voice command. An example of a prompt message is: "When the user says 'remove the part for the next step,' analyze their emotions and, if they are feeling fatigued, generate a message to inform the user of this."

[0591] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0592] Step 1:

[0593] The user wears smart glasses and initiates work by voice input. The voice data is captured by the terminal and transmitted to the server via the network. The input is the user's voice, and the output is the transmission of voice data to the server.

[0594] Step 2:

[0595] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. Audio data is input to the server, and the output is the converted text data. Data processing is performed to analyze the audio waveform information and represent it as string information.

[0596] Step 3:

[0597] The server uses Microsoft Azure's Emotion API to diagnose emotions from voice data. The input is voice data, and the output is data indicating the user's emotional state. It performs data calculations to analyze the tone and intonation of the voice and extract emotional parameters.

[0598] Step 4:

[0599] The server uses a generative AI model with converted text data and sentiment data to optimize work instructions. The input is text data and sentiment data, and the output is optimized work instructions. The generative AI model is given prompt text and executes an algorithm that generates the optimal work procedure.

[0600] Step 5:

[0601] Optimized work instructions are stored on a cloud server and managed so that users can refer to them in the future. The input is work instruction data, and the output is the stored work instruction data.

[0602] Step 6:

[0603] When the user performs the task again, the server automatically executes the work process based on the saved work instructions. The instructions are dynamically adjusted according to the user's real-time sentiment data. The input is real-time sentiment data, and the output is the adjusted work instructions.

[0604] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0605] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0606] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0607] [Fourth Embodiment]

[0608] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0609] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0610] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0611] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0612] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0613] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0614] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0615] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0616] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0617] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0618] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0619] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0620] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0621] This invention provides a system that automatically generates work manuals using voice and screen operation data to streamline operations, and automatically executes tasks based on those manuals. This system is provided on a cloud-based platform, and users can access and use it through their own devices.

[0622] For the system to be implemented, it is a prerequisite that the user terminal is equipped with hardware (microphone) and software (voice recognition function) that enables voice input. When performing a task, the user will explain the procedure for which they wish to create a manual by voice. This explanation will be captured by the user's terminal and sent to the server.

[0623] The server analyzes the audio data using a dedicated speech recognition engine and converts it into text data. This text data forms the basis for the subsequent process of generating work manuals. Additionally, the user terminal monitors the screen as the user performs their tasks and takes screenshots at critical operation steps. This image information is also sent to the server.

[0624] The server analyzes the received text and screen capture data to construct work procedures. A generative AI model is used for the analysis process, automatically recognizing the procedures based on the data and generating a work manual based on the extracted information. The generated manual includes detailed procedures and corresponding images, and is securely stored in the cloud.

[0625] For example, if a user explains the expense reimbursement process verbally and then performs the actual operations on a terminal, the server will generate a manual for the expense reimbursement process based on this verbal explanation and operation. This manual will clearly explain everything from how to start the software to the procedure for entering the necessary information and the verification process.

[0626] In subsequent instances, when a user performs a similar task, their terminal will request automated execution of the task. The server will refer to the task manual stored in the cloud, and the generated AI model will automatically complete the task according to the procedure. As a result, task standardization will progress, and more efficient work execution will be possible.

[0627] The introduction of this system will reduce the need for manual manual creation and work handover, thereby optimizing human resources and improving productivity. This technology will particularly contribute to the effective use of resources in small and medium-sized enterprises and provides a new solution for business operations.

[0628] The following describes the processing flow.

[0629] Step 1:

[0630] The user speaks voice instructions related to their work into the terminal. The terminal acquires this voice in real time and temporarily saves it as a digital audio file.

[0631] Step 2:

[0632] The terminal sends the acquired voice data to the server. The server uses a speech recognition engine to convert the voice data into text data. This converted text accurately reflects the work instructions given by the user.

[0633] Step 3:

[0634] When a user performs a task, the terminal monitors the user's screen operations. It then automatically captures a screen image when a specific operation trigger is detected. This capture contains important operational details necessary for creating procedure manuals.

[0635] Step 4:

[0636] The terminal sends the acquired screen capture data to the server. The server uses a generating AI model to analyze the text data and screen captures, and automatically extracts the components of the business procedure.

[0637] Step 5:

[0638] The server automatically generates an operational manual based on the extracted information. This manual combines screen captures corresponding to instructions with detailed operational procedures. The generated manual is securely stored in cloud storage.

[0639] Step 6:

[0640] From the next time onward, the user will request the same task to be performed automatically from the terminal. The server will read the task manual stored in the cloud, and the generated AI model will accurately reproduce the steps required for automated execution.

[0641] Step 7:

[0642] After the server completes its task, the terminal notifies the user of the results. If the process was successful, or if an error occurred, the user is informed of the details.

[0643] (Example 1)

[0644] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0645] Traditional methods of creating operational manuals and automating business processes were inefficient and required significant human resources. In particular, in small organizations, a lack of standardization and insufficient knowledge sharing led to decreased productivity. Furthermore, manual manual updates were infrequent, making it difficult to maintain consistency in work procedures.

[0646] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0647] In this invention, the server includes means equipped with an information processing device that acquires voice data and converts that data into text data; means equipped with a device that monitors the user's operation procedures and collects visual data for specific actions; and means equipped with a device that analyzes the business flow based on the converted text data and acquired visual data and automatically generates business guidelines. This enables the automatic generation of business manuals and the automatic execution of business processes.

[0648] "Audio data" refers to information expressed through the medium of sound, represented in digital format.

[0649] "Text data" refers to character information represented in digital format, and includes information converted from audio data.

[0650] An "information processing device" is a device that receives data and performs processing such as analysis, conversion, and storage.

[0651] "Operating procedure" refers to a sequence of actions or instructions that a user takes to achieve a specific objective.

[0652] "Visual data" refers to image information, such as screen captures, represented in digital format.

[0653] A "business process flow" is a diagram that shows the sequence of procedures and steps involved in carrying out a business task.

[0654] A "work guidelines" is a document that outlines the necessary procedures and points to note when carrying out a task.

[0655] "Equipment" refers to process equipment or devices that have a specific function and are designed to perform that function.

[0656] An "intelligent construction model" is a model created by applying artificial intelligence technology to automate various tasks, including data analysis and prediction.

[0657] A "remote database" is a data storage location that is physically distant but accessible via a network such as the internet.

[0658] "Storage" refers to the process of securely preserving data and making it accessible for later use.

[0659] This invention provides a system that automatically generates work guidelines using voice and visual data to streamline operations, and then automatically executes those guidelines. The system is cloud-based and can be accessed and used by users through their own devices.

[0660] The user's terminal is equipped with a microphone and speech recognition software for acquiring voice input. The user verbally explains the business workflow, and the audio is acquired and converted by the terminal. Examples of speech recognition software include available speech recognition APIs. The terminal also monitors the user's actions and takes screen captures at specific times. This visual data is later sent to a server and contributes to the formation of the business workflow.

[0661] The server uses a dedicated information processing device to convert audio data into text data. This text data is analyzed using a generative AI model and forms the basis for a process that recognizes and extracts business flows. The server also analyzes this data in conjunction with acquired visual data to automatically generate detailed business guidelines. These guidelines visually reproduce the business flow and assist in understanding the business processes. The generated business guidelines are securely stored in a remote database in the cloud.

[0662] For example, if a user explains the expense reimbursement process verbally and performs the corresponding operations on their device, the server will generate expense reimbursement guidelines based on this data. These guidelines will detail everything from how to launch the software to the procedures for entering necessary information and the verification process.

[0663] In subsequent uses, when a user performs a similar task, the terminal will refer to the task guidelines stored on the server and automatically proceed with the task using a generated AI model. An example of a prompt message that can be entered is, "Please explain and instruct me on the expense reimbursement process using voice." This invention promotes the standardization and efficiency of tasks, enabling users to perform tasks quickly and accurately.

[0664] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0665] Step 1:

[0666] The user explains the work procedure by voice using their device. During this process, voice data is input via the device's microphone. This voice data is passed to speech recognition software, which converts it into text data in real time. This converted text data is then prepared for transmission to the server.

[0667] Step 2:

[0668] The device monitors the user's screen activity in the background. When important operations or screen transitions occur, the device automatically takes a screen capture. This captured image is collected as the user's activity history and prepared as a dataset for transmission to the server.

[0669] Step 3:

[0670] The server receives audio data and screen captures sent from the terminal. The input audio data is analyzed again on the server using a high-precision speech recognition engine and converted into text data. The input screen captures are processed as information for visually analyzing the business flow. At this stage, the foundational data for the entire business process is complete.

[0671] Step 4:

[0672] The server applies a generating AI model using the converted text data and acquired screen captures. The model appropriately recognizes the business flow from the text and visual data and extracts detailed business guidelines. The output business guidelines include detailed descriptions of business procedures and corresponding images. These business guidelines are securely stored in a remote database in the cloud.

[0673] Step 5:

[0674] The next time the user performs a similar task, the user terminal will request the server to automatically execute the task via a prompt message. The server will refer to the task guidelines stored in the cloud, and the generated AI model will perform the task according to the automated instructions. This process allows the user to complete the task smoothly without any intervention.

[0675] (Application Example 1)

[0676] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0677] Improving work efficiency and standardization are crucial challenges, especially in manufacturing. Currently, work procedures are often performed manually based on manuals, which is time-consuming and labor-intensive, and can lead to variations in quality. Furthermore, processes that rely on worker experience increase the risk of errors and require training new employees. There is a need for solutions to these problems and ensure efficient and consistent work quality.

[0678] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0679] In this invention, the server includes means for acquiring voice input and converting the voice into text data, means for monitoring the user's information processing device and acquiring images at important moments, means for analyzing the operation procedure based on the converted text data and acquired images and automatically generating work instructions, and analysis means for analyzing the work procedure using an artificial structure and optimizing the work instructions. This enables the rapid automatic generation and optimization of work procedures.

[0680] "Voice input" refers to the process of acquiring a user's speech as a digital signal.

[0681] "Text data" refers to string information obtained by analyzing voice input.

[0682] An "information processing device" is an electronic device operated by a user to process and monitor data.

[0683] An "image" is a visual record that captures a specific scene on an information processing device.

[0684] An "operating procedure" is a set of steps necessary to complete a specific task or process.

[0685] A "work instruction sheet" is a guideline that specifically outlines the flow and procedures of a task.

[0686] An "artificial structure" is a digital architecture used for generation and analysis using AI models.

[0687] "Analysis means" refers to methods or techniques for analyzing data and processing it into meaningful information.

[0688] A "generative information processing model" is a framework of AI trained to process data and automate specific tasks.

[0689] To implement this invention, it is necessary to accurately acquire voice input and process that information as text data. Therefore, the server should utilize a speech recognition engine and use a service such as the Google Speech-to-Text API. Voice input is captured by the terminal via a microphone and converted into text data. In addition, the user's information processing device should use a camera to acquire images at important scenes in the work procedure and an image processing library such as OpenCV.

[0690] Text and image data acquired by the information processing device are sent to a server. The server then uses an AI model to analyze this data and automatically generates standardized work instructions. These work instructions contain detailed descriptions of specific operating procedures and can be used for the automated execution of subsequent tasks.

[0691] Once the work instructions are generated, the server saves them to a data storage device, making them available via the cloud when requested by the user in the future. This process improves work efficiency and standardization, particularly facilitating the automation of robotic tasks within factories.

[0692] As a concrete example, considering the assembly of parts in a factory, the user can give a voice command such as "Screw part A into part B," and a camera simultaneously records the operation. The server can then use this data to enable automated assembly in the future. An example of a prompt message would be, "Please tell me the process of assembling the toy according to the following steps."

[0693] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0694] Step 1:

[0695] The user inputs work instructions by voice into the device. The device receives this voice through its microphone and converts the voice data into text data using the Google Speech-to-Text API. The input is voice, and the output is text data.

[0696] Step 2:

[0697] The terminal monitors the user's actions and uses the camera to capture images at critical moments. Input is information about the user's actions, and output is the captured image. The image processing library OpenCV is used to capture the necessary scenes.

[0698] Step 3:

[0699] The terminal sends the acquired text data and captured images to the server. The input consists of text data and images, and the terminal performs the operation of transferring these together to the server.

[0700] Step 4:

[0701] The server uses a generative AI model to analyze the received text data and images. Through this analysis, it recognizes specific operating procedures and automatically generates work instructions. The input is text data and images, and the output is work instructions. This process involves structured analysis of the data.

[0702] Step 5:

[0703] The server saves the generated work instructions to an information storage device. The input is the work instruction data, and the output is the saved digital file. This process establishes the functionality necessary for future use.

[0704] Step 6:

[0705] The next time the user requests automated execution of a task, the server will refer to the saved work order and send data to the terminal to perform the requested operation. The inputs are the user's request and the saved work order, and the output is the automated operation. This supports the efficient execution of tasks.

[0706] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0707] This invention is a system that combines a voice input analysis function with an emotion engine that recognizes user emotions, in order to significantly improve work efficiency. In addition to automating work procedures, this system supports more flexible and effective work execution by understanding the user's emotional state in real time and adjusting work accordingly.

[0708] This system begins by acquiring voice input through a user terminal. Users input the details of their work as instructions via voice, and can also convey emotions and intentions included in those instructions. The terminal acquires this voice data in real time and sends it to the server.

[0709] The server uses a dedicated speech recognition engine to convert speech data into natural language text data. Simultaneously, an emotion engine analyzes the user's emotional state from the speech data. The results of this analysis function as a crucial element in generating business procedures and subsequent task execution.

[0710] The user's terminal monitors the user's screen operations and takes screen captures at specific stages of the operation. Based on these captures and audio data, the server uses a generated AI model to analyze the work procedure. The analyzed procedure will also include dynamic adjustments based on the user's emotions.

[0711] For example, when a user is preparing for an important meeting, if they include feelings of tension or stress while giving voice instructions, the emotion engine will detect that emotional state. This information will then be automatically incorporated into the work manual to provide additional assistance and tips to facilitate meeting preparation.

[0712] In the next task execution, automated execution will be possible in response to user requests. The server will use a generated AI model to automatically execute the procedures based on the stored task manual. During this process, the user's emotional state may change in real time, so the emotion engine will detect these changes and adjust the execution flow as needed.

[0713] This system allows users to enjoy optimal work processes tailored to their individual needs and emotions, contributing to improved work efficiency and a better work environment. As a result, it enables human-centered, flexible work design and smarter business operations.

[0714] The following describes the processing flow.

[0715] Step 1:

[0716] The user speaks work-related instructions and associated emotions into the device. The device captures this audio in real time and saves it as a digital audio file.

[0717] Step 2:

[0718] The device sends the acquired audio data to the server. The server uses a speech recognition engine to convert the audio data into natural language text data.

[0719] Step 3:

[0720] The server uses an emotion engine to analyze the user's emotional state from voice data and extracts emotional data. This emotional data is used when generating business procedures.

[0721] Step 4:

[0722] As the user performs their tasks, the terminal continuously monitors the user's screen operations and takes screen captures at critical operation steps.

[0723] Step 5:

[0724] The terminal sends the acquired screen capture data to the server. The server uses a generated AI model to analyze the work procedure based on the text data, screen captures, and sentiment data.

[0725] Step 6:

[0726] The server automatically generates dynamically adjusted operational manuals based on the analysis results. These manuals include emotionally responsive tips and assistance. The generated manuals are stored in the cloud.

[0727] Step 7:

[0728] From the next time onward, when a user requests automated task execution, the server will use the saved task manual as a basis, and the generated AI model will automatically execute the task according to the procedure.

[0729] Step 8:

[0730] During task execution, the server uses an emotion engine to monitor the user's real-time emotional changes and adjust the workflow as needed.

[0731] Step 9:

[0732] After completing a task, the terminal notifies the user of the results and provides relevant feedback if there have been any changes in emotions.

[0733] (Example 2)

[0734] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0735] Conventional business execution systems faced challenges in efficiently processing work instructions entered via user voice input and in dynamically adjusting tasks while considering the user's emotional state. This resulted in a lack of flexibility and efficiency in operations, and the need for additional confirmation by the user.

[0736] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0737] In this invention, the server includes means for converting speech into text information, means for analyzing the user's emotional state and dynamically adjusting the work operation procedure, and means for automatically executing tasks based on work instructions. This enables flexible and efficient work execution that takes the user's emotions into consideration.

[0738] "Voice input" refers to instructions or information that a user gives via voice, and is the data that the system acquires and processes from that voice.

[0739] "Textual information" refers to data in the form of a string of characters that has been converted from voice input through a speech recognition device.

[0740] "Display device operation" refers to a series of operations and controls performed by a user using a display device.

[0741] "Image information" refers to visual data obtained as screen captures of specific important scenes during a user's operation of their display device.

[0742] "Business operation procedures" refer to information that describes the steps and procedures necessary to perform a specific task.

[0743] A "work instruction sheet" is a document that is automatically generated based on work procedures and is used as a guide for carrying out work.

[0744] A "remote storage device" is a data storage device installed in an external facility such as the cloud, and is used to store and manage work instructions and related data.

[0745] "Emotional state" refers to the emotions a user experiences in relation to their work, as captured through analysis, and includes feelings such as joy, stress, and tension.

[0746] A "generative AI model" refers to an algorithm or process that automatically executes tasks based on work instructions, and is used to dynamically adjust tasks according to user requirements.

[0747] This invention is a voice input analysis system aimed at improving work efficiency, and it begins with the user giving work instructions via voice.

[0748] The user provides voice input through the device's microphone. This voice data is acquired by the device in real time. The device then sends the voice data to a server. The server uses speech recognition software to convert the voice data into natural language text. This process utilizes common speech recognition technologies.

[0749] The server uses an emotion analysis engine to analyze the user's emotional state from the converted text information. This is necessary to analyze emotions such as stress and joy. The emotion analysis engine determines the user's emotions by analyzing the tone of voice and word choice.

[0750] Based on the sentiment analysis results and the converted text information, the server uses a generative AI model to generate work operation procedures. These procedures incorporate dynamic adjustments based on the user's emotional state. The generated work instructions are stored in a remote storage device in the cloud.

[0751] For example, if the emotion analysis engine detects tension when a user provides voice input while giving instructions for meeting preparation, the generating AI model will use the prompt "Adjust the meeting procedure based on the user's instructions and emotion data" to generate a work procedure that includes helpful hints and reminders.

[0752] This process allows users to experience the flexible execution of tasks based on voice input and automated analysis. Work instructions that adjust according to emotional changes enable efficient and human-centered operations.

[0753] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0754] Step 1:

[0755] This is the stage where the terminal acquires voice input. The user speaks the necessary instructions for the task into the terminal's microphone. The terminal temporarily stores the acquired voice data and then prepares to send it to the server. In this step, voice data is received as input and the terminal enters a state of waiting to send to the server.

[0756] Step 2:

[0757] The terminal processes the acquired audio data and sends it to the server. The server receives the transmitted audio data and uses speech recognition software to convert it into natural language text. Here, the input is raw audio data, and the output is text information. A speech recognition algorithm is used for this conversion.

[0758] Step 3:

[0759] The server uses an emotion analysis engine to analyze the user's emotional state based on the converted text information. This analysis includes techniques to identify emotions such as stress and tension by evaluating word choice and voice characteristics. The input is the converted text information, and the output is the user's emotional state data.

[0760] Step 4:

[0761] The server uses emotional state data and textual information to input prompt statements into a generative AI model, which then generates business operation procedures. These prompt statements might be something like, "Create the optimal business procedure based on the user's instructions." The input for this step is emotional state data and textual information, and the output is business procedure data. The generative AI model then executes this process.

[0762] Step 5:

[0763] The server saves the generated work procedure data to remote storage. The saved work instructions are managed on the cloud and can be reused later. The input is work procedure data, and the output is saved work instructions.

[0764] Step 6:

[0765] For subsequent task executions, the server will automatically execute tasks based on user requests, referencing saved task instructions. A generative AI model is used here, enabling dynamic adjustments. The input is the saved task instructions, and the output is the task process to be executed. This allows users to enjoy flexible tasks that respond to their emotions.

[0766] (Application Example 2)

[0767] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0768] In today's work environment, there is a demand for both improved work efficiency and worker comfort. However, conventional systems have difficulty flexibly adjusting tasks while taking into account the emotional state of workers, which can result in increased worker stress and fatigue, potentially leading to decreased productivity.

[0769] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0770] In this invention, the server includes means for acquiring voice input and converting the voice into text data, means for recognizing the user's emotions and dynamically adjusting work procedures based on those emotions, and means for analyzing the work procedures and optimizing work instructions based on the converted text data and emotion data. This enables optimal work adjustment according to the worker's emotional state and an automated, efficient work process.

[0771] "Voice input" is the process of capturing the sound emitted by the user as a digital signal into the system.

[0772] "Text data" refers to digital information that represents voice input as a string of characters.

[0773] "Recognizing emotions" means analyzing and identifying the emotional state of a user from their voice or input information.

[0774] "Business procedures" refer to a series of steps or processes required to perform a specific task.

[0775] "Dynamic adjustment" means changing the work content and processes in real time in response to changes in circumstances and conditions.

[0776] "Emotional data" refers to information that indicates a user's emotional state, obtained during the process of recognizing emotions.

[0777] "Optimizing work instructions" means restructuring work procedures and processes to maximize work efficiency and effectiveness.

[0778] "Storage means" refers to a method or apparatus for storing data or information over a long period of time.

[0779] "Control means" refers to technologies and devices used to operate and manage equipment and systems.

[0780] A "generative AI model" is a model that automatically creates algorithms and information from data based on machine learning technology.

[0781] This invention is a system that combines voice input and emotion recognition, and the system for realizing this includes the following program.

[0782] The server receives voice data from the user's terminal over the network to acquire voice input. This voice data is converted into text data using the Google Cloud Speech-to-Text API. Simultaneously, the server recognizes the user's emotions from the voice using the Microsoft Azure Emotion API. This emotion data is an important element for the user's work, and the server uses it to analyze and optimize work procedures. Generative AI models are used for analysis and optimization.

[0783] The generated work instructions are stored in the cloud and made available for download to the user when needed. This allows workers to obtain the optimal work procedures tailored to their emotional state at the time. When the user performs the work again, the server refers to the stored data and automatically executes the tasks. This process includes dynamic adjustments in response to changes in emotions, enabling work to proceed in real time while considering the user's feelings.

[0784] As a concrete example, consider a scenario where a factory worker uses smart glasses and gives a voice command such as "Pick up the part for the next process." In this case, the server analyzes the emotional data from the voice command and adjusts the work pace if fatigue is detected. It can also display a message on the screen saying, "Please take a break."

[0785] The generative AI model generates instructions in real time that respond to the user's emotions each time it receives a voice command. An example of a prompt message is: "When the user says 'remove the part for the next step,' analyze their emotions and, if they are feeling fatigued, generate a message to inform the user of this."

[0786] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0787] Step 1:

[0788] The user wears smart glasses and initiates work by voice input. The voice data is captured by the terminal and transmitted to the server via the network. The input is the user's voice, and the output is the transmission of voice data to the server.

[0789] Step 2:

[0790] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. Audio data is input to the server, and the output is the converted text data. Data processing is performed to analyze the audio waveform information and represent it as string information.

[0791] Step 3:

[0792] The server uses Microsoft Azure's Emotion API to diagnose emotions from voice data. The input is voice data, and the output is data indicating the user's emotional state. It performs data calculations to analyze the tone and intonation of the voice and extract emotional parameters.

[0793] Step 4:

[0794] The server uses a generative AI model with converted text data and sentiment data to optimize work instructions. The input is text data and sentiment data, and the output is optimized work instructions. The generative AI model is given prompt text and executes an algorithm that generates the optimal work procedure.

[0795] Step 5:

[0796] Optimized work instructions are stored on a cloud server and managed so that users can refer to them in the future. The input is work instruction data, and the output is the stored work instruction data.

[0797] Step 6:

[0798] When the user performs the task again, the server automatically executes the work process based on the saved work instructions. The instructions are dynamically adjusted according to the user's real-time sentiment data. The input is real-time sentiment data, and the output is the adjusted work instructions.

[0799] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0800] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0801] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0802] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0803] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0804] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0805] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0806] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0807] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0808] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0809] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0810] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0811] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0812] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0813] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0814] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0815] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0816] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0817] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0818] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0819] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0820] The following is further disclosed regarding the embodiments described above.

[0821] (Claim 1)

[0822] A processing means for acquiring voice input and converting that voice into text data,

[0823] A processing means for monitoring user screen operations and acquiring screen captures at important moments,

[0824] A processing means that analyzes business procedures based on converted text data and acquired screen captures, and automatically generates a business manual.

[0825] A storage method for saving and managing the generated business manuals on the cloud,

[0826] From the next time onward, a control means will automatically execute tasks based on the saved business manual in response to user requests,

[0827] A system that includes this.

[0828] (Claim 2)

[0829] The system according to claim 1, which uses a speech recognition engine to convert speech data into natural language text.

[0830] (Claim 3)

[0831] The system according to claim 1, wherein the generated AI model automatically performs tasks by referring to saved work manuals.

[0832] "Example 1"

[0833] (Claim 1)

[0834] A means equipped with an information processing device that acquires audio data and converts that data into text data,

[0835] A means comprising a device that monitors the user's operating procedures and collects visual data for specific actions,

[0836] A means equipped with a device that analyzes business flows based on converted text data and acquired visual data, and automatically generates business guidelines,

[0837] A storage method for storing the generated business guidelines and managing them on a remote database,

[0838] A means equipped with a control device for automating operations based on stored operational guidelines in response to future user requests,

[0839] A system that includes this.

[0840] (Claim 2)

[0841] The system according to claim 1, which uses an information conversion engine to convert speech information into natural language text.

[0842] (Claim 3)

[0843] The system according to claim 1, wherein a generated intelligent construction model automatically performs tasks by referring to stored work guidelines.

[0844] "Application Example 1"

[0845] (Claim 1)

[0846] A processing means for acquiring voice input and converting that voice into text data,

[0847] A processing means for monitoring the user's information processing device and acquiring images at important moments,

[0848] A processing means that analyzes the operation procedure based on the converted text data and acquired images, and automatically generates a work instruction sheet,

[0849] A storage means for saving and managing the generated work instructions on an information storage device,

[0850] From the next time onward, a control means will automatically execute operations based on the saved work instructions in response to the user's request.

[0851] An analysis means that analyzes work procedures using artificial structures and optimizes work instructions,

[0852] A system that includes this.

[0853] (Claim 2)

[0854] The system according to claim 1, which uses a speech recognition system to convert speech data into linguistic information.

[0855] (Claim 3)

[0856] The system according to claim 1, wherein the generated information processing model automatically performs operations by referring to saved work instructions.

[0857] "Example 2 of combining an emotion engine"

[0858] (Claim 1)

[0859] A processing means for acquiring voice input and converting that voice into text information,

[0860] A processing means that monitors the user's operation of the display device and acquires image information at important moments,

[0861] A processing means that analyzes business operation procedures based on converted text information and acquired image information, and automatically generates work instruction documents,

[0862] A storage method for storing and managing generated work instructions on a remote storage device,

[0863] From the next time onward, a control means will automatically execute tasks based on stored work instructions in response to user requests,

[0864] An emotion analysis means that analyzes the user's emotional state and dynamically adjusts the work operation procedure,

[0865] A system that includes this.

[0866] (Claim 2)

[0867] The system according to claim 1, which uses a speech recognition device to convert speech data into natural language character information.

[0868] (Claim 3)

[0869] The system according to claim 1, wherein the generated AI model automatically executes tasks by referring to stored work instructions.

[0870] "Application example 2 when combining with an emotional engine"

[0871] (Claim 1)

[0872] A processing means for acquiring voice input and converting that voice into text data,

[0873] A processing means that recognizes the user's emotions and dynamically adjusts business procedures based on those emotions,

[0874] A processing means that analyzes work procedures and optimizes work instructions based on converted text data and sentiment data,

[0875] A storage method for saving generated work instructions and managing them on a shared resource,

[0876] A control means that automatically executes the work process based on saved work instructions in response to a subsequent request,

[0877] A system that includes this.

[0878] (Claim 2)

[0879] The system according to claim 1, comprising a speech recognition engine that converts speech data into natural language text and an emotion analysis engine that estimates the user's emotions.

[0880] (Claim 3)

[0881] The system according to claim 1, wherein the AI ​​model generates the work process automatically by referring to saved work instructions and adjusting the flow while taking into account the user's emotions. [Explanation of Symbols]

[0882] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A processing means for acquiring voice input and converting that voice into text data, A processing means for monitoring user screen operations and acquiring screen captures at important moments, A processing means that analyzes business procedures based on converted text data and acquired screen captures, and automatically generates a business manual. A storage method for saving and managing the generated business manuals on the cloud, From the next time onward, a control means will automatically execute tasks based on the saved business manual in response to user requests, A system that includes this.

2. The system according to claim 1, which uses a speech recognition engine to convert speech data into natural language text.

3. The system according to claim 1, wherein the generated AI model automatically performs tasks by referring to saved work manuals.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A