system
A system that analyzes audio and video data to automatically generate and update operation manuals addresses the inefficiencies in manual creation, improving knowledge sharing and operational efficiency by providing real-time support.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
The creation of operation manuals requires significant time and effort, and important operational knowledge is unevenly distributed among individuals, leading to inefficiencies and stagnation in updating manuals, hindering the effective operation of organizations.
A system that acquires audio and video data, analyzes them to automatically generate business procedures, and provides optimal information in response to user questions, enabling efficient knowledge sharing and updating.
Facilitates the efficient generation and updating of operation manuals, promotes knowledge sharing, and supports users with real-time information and guidance, enhancing organizational efficiency.
Smart Images

Figure 2026069094000001_ABST
Abstract
Description
Technical Field
[0004] , , ,
[0005] , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The creation of current operation manuals requires a great deal of time and effort, imposing an excessive burden on workers. Also, important knowledge related to operations is unevenly distributed among specific individuals, and the update of the manual may stagnate due to the absence or transfer of the updater. In this situation, the efficient operation of the entire organization is hindered.
Means for Solving the Problems
[0005] This invention provides a system that acquires audio and video data, analyzes them, and automatically generates business procedures. Using means to analyze audio data and convert it into text data, and means to analyze video data and recognize actions, the generated procedures are automatically compiled into manuals and stored. Furthermore, existing procedures are updatable, and the system includes a function to analyze user questions and provide optimal information, enabling efficient sharing and updating of business knowledge.
[0006] "Audio data" refers to information recorded in digital format from speech or other sounds.
[0007] "Video data" refers to information that records images or actions in digital format.
[0008] "Analysis" refers to the process of processing acquired data and extracting meaning and information from it.
[0009] "Text data" refers to information in the form of a string obtained by analyzing audio data.
[0010] "Recognizing motion" refers to the process of understanding specific actions or operations from video data.
[0011] A "procedure" refers to a series of steps or actions required to complete a specific task.
[0012] "Automatic generation" refers to the process by which a system creates procedures or information without human intervention.
[0013] "Storage" refers to the process of saving generated data and manuals as needed.
[0014] "User" refers to an individual or organization that uses the system.
[0015] "Analyzing a question" refers to the process of understanding a user's question and identifying and providing appropriate information.
[0016] "Providing information" refers to the act of displaying and notifying appropriate answers and procedures to the user based on the analysis results.
Brief Explanation of Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined. 7>
Embodiments for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the terminology used in the following description will be explained.
[0020] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] To implement this invention, it is first necessary to set up an environment for acquiring audio and video data. This requires a terminal equipped with a microphone for capturing audio and a camera for recording video. The terminal is responsible for recording the user's work in real time and transmitting that data to a server.
[0039] Step 1: Data Acquisition
[0040] User: Performs tasks while carrying a device. Detailed recording is possible, especially when performing new tasks or operating new equipment.
[0041] Terminal: Uses a microphone and camera to collect audio and video during work and transmits the data to the server in real time.
[0042] Step 2: Data Analysis
[0043] Server: The received audio data is analyzed by the speech recognition engine and converted into text data. This makes it clear what the user is saying.
[0044] Server: Video data is analyzed using computer vision algorithms to identify work procedures and characteristic movements.
[0045] Step 3: Automatic generation and storage of procedures
[0046] Server: Based on the obtained text data and video analysis information, work procedures are automatically generated. These procedures are documented in an easy-to-understand format and stored in a database.
[0047] Server: The generated manual is kept as the latest procedure at that time and is updated as needed.
[0048] Step 4: Information Provision
[0049] User: If any questions arise during the work process, users can ask them via their terminal.
[0050] Terminal: When a question is received via voice, it converts it to text using speech recognition and sends it to the server.
[0051] Server: Analyzes the question content and extracts appropriate information from manual databases and knowledge bases.
[0052] Terminal: The optimal answer is provided to the user in text or voice.
[0053] This embodiment allows users to naturally incorporate their knowledge into the system without any special operations, and provide it as useful information to future learners. This system effectively incorporates on-the-ground knowledge and promotes knowledge sharing throughout the organization.
[0054] The following describes the processing flow.
[0055] Step 1:
[0056] User: Start work and turn on the terminal. Perform normal tasks without making any special settings.
[0057] Terminal: Activates microphone and camera to record and videotape the user's work in real time. Recorded and videotaped data is sent to the server with security measures in place.
[0058] Step 2:
[0059] Server: The server analyzes the received audio data through a speech recognition engine and extracts the spoken content as text data. This process uses techniques to reduce audio noise and convert it into accurate text strings.
[0060] Server: Video data is analyzed by computer vision algorithms to recognize specific user actions and operations. For example, everyday actions such as pressing buttons on equipment or pulling levers are identified.
[0061] Step 3:
[0062] Server: Integrates the results of audio and video analysis to automatically generate work procedure manuals. The generated manuals are made easy to read using natural language generation technology and clearly explain information relevant to the work.
[0063] Server: Procedure manuals are stored in a database and made accessible within the organization. These manuals are continuously updated and used as the most up-to-date work procedures throughout the organization.
[0064] Step 4:
[0065] User: If questions or uncertainties arise during work, users can ask them via voice commands.
[0066] Terminal: Receives questions as voice input, performs speech recognition to convert them into text data, and sends them to the server.
[0067] Step 5:
[0068] Server: Analyzes questions submitted by users and extracts appropriate answers from relevant manuals and databases. Specifically, it aims to select the most relevant information to the question and provide it in a user-friendly format.
[0069] Terminal: Displays responses sent from the server to the user in audio or text format. This allows the user to quickly access solutions and work procedures.
[0070] (Example 1)
[0071] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0072] In recording and communicating work procedures on-site, traditional methods often relied on individual experience and memory, leading to challenges in the reproducibility and sharing of procedures. Furthermore, real-time information retrieval and question-and-answer sessions were not possible, preventing the provision of highly efficient support.
[0073] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0074] In this invention, the server includes a structure for acquiring audio and video information, a structure for automatically generating procedures based on the analyzed audio and video information, and a structure for answering questions to resolve information deficiencies that users may have during their work. This enables improved recording and reproducibility of on-site work procedures, real-time information provision, and efficient work support.
[0075] "Auditory information" refers to data obtained through changes in sound, and in particular, information that includes human voices.
[0076] "Visual information" refers to visual data acquired through visual devices such as cameras.
[0077] "Analysis" is the act of processing acquired data and extracting meaning and characteristics based on that data.
[0078] "Textual information" refers to audio data expressed as text, a format that allows for visual confirmation of the audio content.
[0079] A "structure that recognizes motion" is a structure that has the function of identifying and understanding the movements of people and objects contained in video data.
[0080] A "structure for automatically generating procedures" is a system that mechanically creates procedures and methods based on analyzed information.
[0081] A "storage and presentation structure" is a mechanism for saving generated information and displaying or providing it as needed.
[0082] A "question and answer structure" is a structure that receives questions from users, searches for answers to those questions, and provides responses.
[0083] To implement this invention, it is necessary to create an environment that can efficiently acquire and analyze audio and video information. Using a terminal equipped with a microphone and camera is appropriate for this purpose. The terminal is responsible for capturing the user's actions and voice in real time and transmitting that data to the server.
[0084] In terms of specific hardware, a high-resolution camera and a microphone capable of capturing clear audio are recommended. A stable network connection is also necessary for processing and transmitting data in real time.
[0085] Data transmitted from the terminal is processed on the server. The server uses speech recognition software to convert speech information into text. A natural language processing engine is utilized to accurately obtain text data from the speech. Furthermore, computer vision algorithms are used to identify important actions from the video information and utilize this information for procedure generation.
[0086] The generated procedures are stored in a database and provided to users as needed. Users can access this information in real time via their terminals and receive immediate support through the Q&A service if they have any questions during their work.
[0087] As a concrete example, when recording the steps required to operate a new machine, the user activates the terminal and provides instructions while operating it. This explanation and operation are recorded simultaneously, and after analysis on the server, the procedure is automatically generated. This procedure can then be referenced by other users.
[0088] Examples of prompts for a generative AI model include:
[0089] One example is, "Please explain the operating procedures for the new machine."
[0090] With the above configuration, it is possible to build a system that effectively records and shares on-site work knowledge and promotes the utilization of knowledge throughout the organization.
[0091] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0092] Step 1:
[0093] The device uses a microphone and camera to acquire user audio and video information. The input includes the user's real-time audio and video. This information is converted into digital data. Specifically, it records the user verbally explaining the procedures while operating the machine.
[0094] Step 2:
[0095] The terminal transmits acquired audio and video data to the server in real time. The input is digitized audio and video, and the output is data received on the server side. This transmission takes place over a stable network.
[0096] Step 3:
[0097] The server inputs the received audio data into a speech recognition engine and converts it into text information. Specifically, it analyzes the audio information using natural language processing technology and documents it. The output is a text representation of what the user said.
[0098] Step 4:
[0099] The server analyzes video data using computer vision algorithms. Using the video data as input, the server identifies user actions and key movements related to a procedure. The output is data representing what actions were performed. For example, a scene where the user presses a button is analyzed.
[0100] Step 5:
[0101] The server automatically generates procedures based on the analyzed audio and video information. Here, text and action data are integrated and formatted into easy-to-understand procedures. The output is a document of the generated business procedures.
[0102] Step 6:
[0103] The server stores the generated procedures in a database and prepares them for presentation to the user. Users can access this information via their terminal as needed. The saved procedures become an asset that other users can also refer to.
[0104] Step 7:
[0105] If a user has a question during work, they can ask it via their terminal. The input is a voice question. The terminal sends this question to a speech recognition engine, which then converts it into text and sends the text to the server.
[0106] Step 8:
[0107] The server analyzes the received question and searches for the most relevant information from its database and knowledge base. Based on the results, the server generates an answer. The output is the answer information to the question.
[0108] Step 9:
[0109] The terminal receives responses sent from the server and provides them to the user in text or audio format. Based on the information obtained by the user, tasks can be carried out smoothly.
[0110] (Application Example 1)
[0111] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0112] In manufacturing environments, there is a need to efficiently understand and improve work procedures and standard operations. This requires the introduction of new work procedures and efficiency methods, but traditional methods have made it difficult to generate and update procedures in real time. Furthermore, prompt and accurate responses to worker inquiries are required, but achieving this necessitates the management and provision of vast amounts of information. In response to these circumstances, there is a need for a system that provides efficient and up-to-date information, thereby improving work efficiency.
[0113] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0114] In this invention, the server includes a device for acquiring audio and video information, a process for analyzing the acquired audio information and converting it into text information, and a process for analyzing the acquired video information and recognizing the work. This makes it possible to automatically generate work procedures in real time, store and present the latest work procedures. It also records the actions of workers, supports work efficiency, and can effectively respond to user inquiries.
[0115] "Audio information" refers to voice and sound data acquired through acoustic sensors.
[0116] "Visual information" refers to visual data acquired using visual sensors.
[0117] A "device" is a piece of equipment used to acquire and process audio and video data.
[0118] "Analysis" is the process of converting acquired data into a meaningful format.
[0119] "Textual information" refers to audio data that has been represented as text.
[0120] "Work" refers to a series of processes or operations carried out in a factory or similar facility.
[0121] "Processing" refers to specific functions or calculations performed by a device or server.
[0122] A "procedure" is a series of steps required to perform a task or operation.
[0123] "Storage" refers to the act of saving generated information or data.
[0124] "Presentation" means providing stored information in a format that users can verify.
[0125] "Efficiency improvement" means aiming to perform tasks and processes more effectively with less time and effort.
[0126] A "server" is a computer system that performs data analysis and storage over a network.
[0127] "User" refers to a person who operates or uses the system.
[0128] A "question" refers to a question or problem that a user poses to the system.
[0129] A "response" is the information or solution that a system provides in response to a user's question.
[0130] This invention provides a system aimed at improving efficiency and standardizing work in factories and manufacturing sites. The server utilizes various hardware and software to acquire and analyze audio and video information. Specifically, the server uses a microphone to acquire audio information and a camera to acquire video information. This data is transmitted to the server in real time.
[0131] The server converts speech information into text using speech recognition software (e.g., Google® Cloud Speech-to-Text). For video information, computer vision algorithms (e.g., OpenCV, TENSORFLOW®) are used to analyze the work and recognize characteristic movements. This allows for the extraction of efficient work procedures and actions.
[0132] When a user performs a task, the optimal work procedure is presented based on the situation unfolding in real time via the terminal. This terminal uses a smart device to immediately provide the information the worker needs. Furthermore, when a user asks a question, a generative AI model is used to quickly generate and present the most appropriate answer.
[0133] As a concrete example, when a new assembly line is introduced, the system records and analyzes the workers' actions. It then presents optimized work procedures and shares them in a format that other workers can refer to. If a user prompts with "Please tell me the most efficient procedure for this task," the system can provide the most efficient procedure based on its accumulated data.
[0134] In this way, the collaboration between servers, terminals, and users promotes the sharing and standardization of knowledge within the factory, leading to improved work efficiency.
[0135] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0136] Step 1:
[0137] The terminal captures audio and video information at the work site. The input consists of the worker's actions and explanations, which are captured in real time by a microphone and camera and sent to the server as digital data. The output is the conversion of analog data into digital format, ready to be sent to the server.
[0138] Step 2:
[0139] The server converts received audio data into text using speech recognition software. The input is audio data sent from the terminal. A speech recognition engine (e.g., Google Cloud Speech-to-Text) is used to analyze the audio data and output it as text. As a result, the audio information is stored on the server in text format.
[0140] Step 3:
[0141] The server analyzes video data using computer vision algorithms to recognize work procedures and actions. The input is video data sent from the terminal. Using tools such as OpenCV and TensorFlow, motion analysis is performed frame by frame of the video. This identifies the characteristics of the actions and outputs them as procedure information.
[0142] Step 4:
[0143] The server automatically generates the optimal work procedure based on the analyzed text and video information. The input consists of the text and video procedure information obtained from the previous analysis. This information is integrated to generate an effective work procedure, which is then stored in a database. As output, the latest work procedure is documented and recorded as a manual.
[0144] Step 5:
[0145] Users can ask questions through their devices. The input consists of user voice or text queries. The device analyzes this input and sends it to the server. The server uses a generative AI model to extract information from its knowledge base and generate the optimal answer.
[0146] Step 6:
[0147] The terminal receives a response from the server and presents it to the user. The input is the response generated by the server, which is then provided to the user as text or audio. As a result, the user can obtain the necessary information in real time.
[0148] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0149] In implementing this invention, in addition to acquiring audio and video data, it is necessary to incorporate an emotion engine for recognizing the user's emotions. The emotion engine analyzes the user's emotional state from the tone and intonation of the voice, facial expressions and gestures in the video, etc. This emotion recognition technology enables the provision of adaptive information that takes into account the user's psychological state during learning and work.
[0150] Data acquisition
[0151] User: Performs daily tasks, and audio and video are collected through the device. Users do not need to perform complex operations and can focus on their work naturally.
[0152] Terminal: Inputs audio and video data in real time and simultaneously analyzes emotional data through an emotion engine.
[0153] Data and sentiment analysis
[0154] Server: Audio data is converted into text using a speech recognition algorithm, and video data is analyzed using an action recognition algorithm to identify the content of the work. Furthermore, the user's current emotional state is identified by an emotion engine. This enables data analysis linked to emotions.
[0155] Procedure generation and application
[0156] Server: Automatically generates business procedures based on analysis results and adjusts the content according to the user's emotions. For example, if it is determined that the user is confused, the approach will change to explain the procedure in more detail.
[0157] Server: The generated procedures are stored in a database and can be accessed by users as needed. Furthermore, information optimized by emotion is displayed to aid learners' understanding.
[0158] Information provision and feedback
[0159] User: If the user has any questions or uncertainties during the process, the emotion engine will provide answers that take into account their emotional state at the time. This allows the user to receive better support.
[0160] Terminal: Sends data to the server at the appropriate time and presents emotion-responsive feedback to the user in voice or text.
[0161] This system, which incorporates emotion recognition, provides users with detailed support in their work and learning, thereby enabling more effective knowledge acquisition and task execution.
[0162] The following describes the processing flow.
[0163] Step 1:
[0164] User: Start the terminal and begin normal work. Explore new procedures and tasks as needed.
[0165] Terminal: Uses a microphone and camera to record and videotape the user's voice and video in real time, and sends the data to the server.
[0166] Step 2:
[0167] Server: Receives audio data and converts it to text using a speech recognition engine. This process includes noise filtering to accurately analyze the spoken content.
[0168] Server: Simultaneously, it analyzes video data using computer vision algorithms to recognize user actions. It identifies the type and sequence of actions and stores the information in a database.
[0169] Step 3:
[0170] Server: Uses an emotion engine to analyze the user's emotional state from received audio and video data. Specifically, it identifies emotions based on the tone of voice and facial expressions and gestures in the video.
[0171] Server: Combines sentiment analysis results with other analysis data to understand the user's current situation and the necessary steps.
[0172] Step 4:
[0173] Server: Automatically generates business procedures based on analysis results. Adjusts the level of detail and explanation of the procedures according to the user's emotional state. For example, if the user is nervous, the procedures are presented more clearly and step-by-step.
[0174] Server: Stores the generated procedures in a database and updates them as needed. The procedures must always be up-to-date for daily operations.
[0175] Step 5:
[0176] User: If questions arise during the process or if additional information is needed, ask them at the terminal.
[0177] Step 6:
[0178] Terminal: Receives user questions via voice, performs speech recognition again to convert them to text for transmission to the server, and sends them to the server along with sentiment data.
[0179] Step 7:
[0180] Server: Analyzes the question content and submitted sentiment data, and searches the database for relevant information. In particular, it selects information in a way that takes the user's emotions into consideration and determines the steps to be provided.
[0181] Terminal: The terminal presents information selected by the server to the user via voice or text, and provides a relaxed learning environment by responding in a way that is sensitive to emotions.
[0182] In this way, the system aims to improve the quality of work and learning by utilizing emotion recognition capabilities to provide users with information in the most optimal format.
[0183] (Example 2)
[0184] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0185] In today's work and learning environments, users are required to efficiently handle diverse data. However, systems that understand users' psychological and emotional states and provide optimal information accordingly are still insufficient, resulting in insufficient improvements in user efficiency and satisfaction. Therefore, there is a need for means to enable flexible responses that take into account users' emotional states, thereby improving both work efficiency and user experience.
[0186] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0187] In this invention, the server includes means for recognizing the user's emotional state based on audio and video data, means for automatically generating and adjusting procedures based on the analyzed audio and video data and the user's emotional state, and means for analyzing the user's questions and emotional state and providing optimal information. This makes it possible to provide information adaptively according to the user's emotions and improve work efficiency and learning efficiency.
[0188] "Voice data" refers to information that records user speech and ambient sounds, and serves as the basis for speech recognition and emotion analysis.
[0189] "Video data" refers to visual information that records the user's movements and facial expressions, and is used for motion recognition and emotion analysis.
[0190] "Text data" refers to information obtained by analyzing audio data and converting it into written text, representing the content of the user's speech in written form.
[0191] "Motion recognition" is a technology that identifies a user's actions and gestures by analyzing video data.
[0192] "Emotional state" refers to the psychological state that indicates the user's emotions, and means the psychological response that can be inferred from audio and video.
[0193] A "procedure" is a set of steps that users need to complete their tasks or learning, and it is automatically generated and provided to the user.
[0194] "Automatic generation" refers to the process by which a system generates procedures and information based on data, without manual intervention by humans.
[0195] "Optimal information" refers to information that contains necessary and appropriate content, tailored to the user's situation and emotional state.
[0196] To implement this invention, a system is required for real-time acquisition, analysis, and provision of information to the user of audio and video data. The system is configured as follows.
[0197] Data acquisition
[0198] Terminal: As users perform tasks or studies, audio and video data are collected in real time by the camera and microphone built into the terminal. Users do not need to perform any special operations in a natural work environment.
[0199] Data and emotion recognition
[0200] Server: Audio data transmitted from the terminal is converted into text data using a speech recognition algorithm (e.g., Google Speech-to-Text). Video data can be analyzed to determine the user's actions by applying a motion recognition algorithm (e.g., OpenCV), and the user's emotional state can be estimated using an emotion engine (e.g., an emotion analysis tool). In this way, the server understands the user's psychological state.
[0201] Procedure generation and information provision
[0202] Server: Based on the analysis of the user's voice, video, and emotional state, it automatically generates optimized work procedures. For example, if it determines that the user is confused, it provides detailed and specific instructions. These generated procedures are stored in a database, allowing the user to access them as needed and continue their work.
[0203] Terminal: Receives instructions from the server and provides feedback tailored to the user's current emotional state. If voice guidance is appropriate, it uses speech synthesis technology (e.g., a speech synthesis engine) to provide voice guidance to the user. Alternatively, instructions can be displayed on the screen in text format.
[0204] Specific example
[0205] For example, when a user encounters a question on their device while learning how to use new software, the system detects confusion based on the user's tone of voice and video, and provides detailed step-by-step instructions. If a positive emotional state is detected, only an overview is provided, leaving room for the user to continue learning on their own.
[0206] Example of a prompt
[0207] "Analyze the user's current audio and video data, recognize their emotions, and generate instructions. Output a step-by-step guide if the user becomes confused."
[0208] This system allows users to receive individually optimized information, enabling them to efficiently advance their work and learning.
[0209] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0210] Step 1:
[0211] Data acquisition
[0212] Terminal: The terminal acquires audio and video data from the user in real time. Inputs include audio collected by the microphone and video recorded by the camera. These data are output as digital audio signals and video frames.
[0213] Step 2:
[0214] Data transfer and reception
[0215] Terminal: Transmits acquired audio and video data to the server via an internet connection.
[0216] Server: Receives data sent from the terminal and prepares the audio data for processing with a speech recognition algorithm. The output consists of two streams: audio data and video data.
[0217] Step 3:
[0218] Voice analysis
[0219] Server: Inputs audio data into a speech recognition algorithm (e.g., speech recognition engine) and converts it into text data. This is achieved by extracting linguistic characteristics from the waveform of the audio signal and outputting them as corresponding text strings.
[0220] Step 4:
[0221] Video analysis
[0222] Server: The received video data is input into the motion recognition algorithm. The algorithm analyzes the user's movements and facial expressions from the video frames and extracts corresponding motion data. The output is descriptor data related to the physical actions performed by the user.
[0223] Step 5:
[0224] Emotion analysis
[0225] Server: Uses an emotion engine to estimate the user's emotional state based on analysis of audio waveforms, text data, and video. By passing the input data through an emotion classification model, the user's state is output as an emotion tag (e.g., joy, confusion).
[0226] Step 6:
[0227] Automatic generation and adjustment of procedures
[0228] Server: Based on the analysis results, the server starts a process to automatically generate business procedures. By inputting data about the user's actions and emotions as prompts to the generating AI model, it outputs procedures that have been adjusted according to the emotions.
[0229] Step 7:
[0230] Information provision and feedback
[0231] Terminal: Receives instructions and feedback generated from the server and presents them to the user. Feedback is output as audio or visual messages depending on the user's emotional state. This allows the user to work smoothly based on optimized procedures.
[0232] (Application Example 2)
[0233] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0234] In on-site work, a challenge exists in that instructions and support that ignore the emotional state of workers make it difficult to effectively improve work efficiency and reduce psychological burden. In particular, when workers are feeling fatigued or confused, conventional systems may not provide appropriate support, potentially leading to decreased productivity and safety. Therefore, there is a need for a system that can analyze workers' emotions in real time and provide appropriate instructions and information according to their state.
[0235] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0236] In this invention, the server includes means for acquiring audio and video information, means for analyzing the user's emotional state, and means for adjusting and providing instructions based on the analyzed emotional state. This makes it possible to grasp the worker's emotions in real time and provide instructions and support appropriate to that emotional state.
[0237] "Audio information" refers to data consisting of human speech and ambient sounds, acquired using acoustic sensors such as microphones.
[0238] "Visual information" refers to visual data of objects and people acquired through cameras and other image sensors.
[0239] "Text data" refers to data in text format obtained by converting audio information into text.
[0240] "Means of recognizing motion" refers to technologies that analyze video information and detect specific patterns or actions related to the movement of people or objects.
[0241] "Means for automatically generating instructions" refers to a process that automatically generates appropriate work procedures and instructions based on the results of analysis of audio information, video information, emotional states, etc.
[0242] "Means of analyzing emotional states" refer to algorithms or engines that determine a person's emotions and psychological state from their tone of voice, facial expressions, and gestures.
[0243] "Means for adjusting and providing generated instructions" refers to a process that optimizes initially generated instructions according to the user's emotional state and presents them to the worker at the appropriate time.
[0244] This invention is a system that provides appropriate support tailored to the emotions of workers in the workplace. The system has the function of acquiring audio and video information, analyzing it, and recognizing actions and emotional states.
[0245] First, the device uses a microphone and camera to acquire audio and video information in real time. The audio information is converted into text data by a speech recognition algorithm (e.g., Google Cloud Speech-to-Text). The video information is used to identify the work being done using an action recognition algorithm.
[0246] Subsequently, the server analyzes the user's emotional state from voice tone and facial expressions obtained using an emotion engine (e.g., Microsoft® Azure® Emotion API). Based on the emotional state and behavioral information, the server automatically generates and adjusts instructions.
[0247] The analysis results and adjusted instructions are stored in a database and provided to the user in voice or text format at the appropriate time. For example, if a worker is confused, the system can provide specific instructions such as, "Next, use the F panel and attach it to the A base. When doing so, use the right side of the A base."
[0248] The following are examples of prompts to input into the generative AI model.
[0249] "Analyze the emotions captured from real-time audio and video data to identify at which stage of the task the user is experiencing difficulties. Then, propose work procedures or support tailored to their emotional state."
[0250] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0251] Step 1:
[0252] The terminal uses a microphone and camera to acquire audio and video information from the work environment in real time. The audio information includes the worker's speech and ambient sounds, while the video information is video captured by the camera. This data is transmitted directly to the server.
[0253] Step 2:
[0254] The server converts the acquired audio information into text data using a speech recognition algorithm (e.g., Google Cloud Speech-to-Text). The input is audio information, and the output is string data obtained by analyzing the audio. This process records the worker's speech in text.
[0255] Step 3:
[0256] The server analyzes the acquired video information using a motion recognition algorithm to identify the work content and actions. The input is video information, and the output is data related to the actions (e.g., specific actions performed by the worker). This allows for an accurate understanding of what actions the worker is performing.
[0257] Step 4:
[0258] The server analyzes the emotional state using an emotion engine (e.g., Microsoft Azure Emotion API) based on voice tone and facial expressions. Input is audio and video information, and output is emotional state data (e.g., joy, confusion, fatigue). This process identifies emotions using voice tone, speaking style, and facial expressions.
[0259] Step 5:
[0260] The server automatically generates appropriate instructions based on analyzed voice data, motion data, and emotional state data. The input is the data set obtained in the previous step, and the output is specific instructions to be presented to the user. These instructions are adjusted according to the worker's emotional state.
[0261] Step 6:
[0262] The user receives instructions from the server and performs the task. Instructions are provided in audio or text format and include specific guidance to help the user continue the task. Care is taken to ensure that the instructions are presented in a way that is easy for the user to understand.
[0263] Step 7:
[0264] The terminal continuously sends data to the server and updates it whenever processing is complete or new input is required. This enables real-time feedback that responds to changes in the worker's emotional state and work content.
[0265] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0266] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0267] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0268] [Second Embodiment]
[0269] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0270] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0271] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0272] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0273] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0274] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0275] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0276] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0277] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0278] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0279] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0280] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0281] To implement this invention, it is first necessary to set up an environment for acquiring audio and video data. This requires a terminal equipped with a microphone for capturing audio and a camera for recording video. The terminal is responsible for recording the user's work in real time and transmitting that data to a server.
[0282] Step 1: Data Acquisition
[0283] User: Performs tasks while carrying a device. Detailed recording is possible, especially when performing new tasks or operating new equipment.
[0284] Terminal: Using a microphone and a camera, it collects the voice and video during work and transmits the data to the server in real time.
[0285] Step 2: Data Analysis
[0286] Server: The received voice data is analyzed by a voice recognition engine and converted into text data. This clarifies what the user is saying.
[0287] Server: The video data is analyzed using a computer vision algorithm to identify the work procedures and characteristic actions.
[0288] Step 3: Automatic Generation and Storage of Procedures
[0289] Server: Based on the obtained text data and video analysis information, the business procedures are automatically generated. This procedure is manualized in an easy-to-understand format and stored in a database.This embodiment allows users to naturally incorporate their knowledge into the system without any special operations, and provide it as useful information to future learners. This system effectively incorporates on-the-ground knowledge and promotes knowledge sharing throughout the organization.
[0297] The following describes the processing flow.
[0298] Step 1:
[0299] User: Start work and turn on the terminal. Perform normal tasks without making any special settings.
[0300] Terminal: Activates microphone and camera to record and videotape the user's work in real time. Recorded and videotaped data is sent to the server with security measures in place.
[0301] Step 2:
[0302] Server: The server analyzes the received audio data through a speech recognition engine and extracts the spoken content as text data. This process uses techniques to reduce audio noise and convert it into accurate text strings.
[0303] Server: Video data is analyzed by computer vision algorithms to recognize specific user actions and operations. For example, everyday actions such as pressing buttons on equipment or pulling levers are identified.
[0304] Step 3:
[0305] Server: Integrates the results of audio and video analysis and automatically generates work procedure manuals. The generated manuals are made easy to read using natural language generation technology and clearly explain information relevant to the work.
[0306] Server: The procedure manual is stored in a database and can be referenced within the organization. This procedure manual can be continuously updated and is used as the latest working procedure throughout the organization.
[0307] Step 4:
[0308] User: When questions or uncertainties arise during work, the user can voice a question to the terminal.
[0309] Terminal: Receive the question as voice input, perform voice recognition to convert it into text data, and send it to the server.
[0310] Step 5:
[0311] Server: Analyze the question sent by the user and extract an appropriate answer from relevant procedure manuals or databases. Specifically, select the information most relevant to the question content and aim to provide it in a highly convenient form.
[0312] Terminal: Display the answer sent by the server to the user in voice or text. This allows the user to quickly access countermeasures and working procedures. <000 In this invention, the server includes a structure for acquiring audio and video information, a structure for automatically generating procedures based on the analyzed audio and video information, and a structure for answering questions to resolve information deficiencies that users may have during their work. This enables improved recording and reproducibility of on-site work procedures, real-time information provision, and efficient work support.
[0318] "Auditory information" refers to data obtained through changes in sound, and in particular, information that includes human voices.
[0319] "Visual information" refers to visual data acquired through visual devices such as cameras.
[0320] "Analysis" is the act of processing acquired data and extracting meaning and characteristics based on that data.
[0321] "Textual information" refers to audio data expressed as text, a format that allows for visual confirmation of the audio content.
[0322] A "structure that recognizes motion" is a structure that has the function of identifying and understanding the movements of people and objects contained in video data.
[0323] A "structure for automatically generating procedures" is a system that mechanically creates procedures and methods based on analyzed information.
[0324] A "storage and presentation structure" is a mechanism for saving generated information and displaying or providing it as needed.
[0325] A "question and answer structure" is a structure that receives questions from users, searches for answers to those questions, and provides responses.
[0326] To implement this invention, it is necessary to create an environment that can efficiently acquire and analyze audio and video information. Using a terminal equipped with a microphone and camera is appropriate for this purpose. The terminal is responsible for capturing the user's actions and voice in real time and transmitting that data to the server.
[0327] In terms of specific hardware, a high-resolution camera and a microphone capable of capturing clear audio are recommended. A stable network connection is also necessary for processing and transmitting data in real time.
[0328] Data transmitted from the terminal is processed on the server. The server uses speech recognition software to convert speech information into text. A natural language processing engine is utilized to accurately extract text data from the speech. Furthermore, computer vision algorithms are used to identify important actions from the video information and utilize this information for procedure generation.
[0329] The generated procedures are stored in a database and provided to users as needed. Users can access this information in real time via their terminals and receive immediate support through the Q&A service if they have any questions during their work.
[0330] As a concrete example, when recording the steps required to operate a new machine, the user activates the terminal and provides instructions while operating it. This explanation and operation are recorded simultaneously, and after analysis on the server, the procedure is automatically generated. This procedure can then be referenced by other users.
[0331] Examples of prompts for a generative AI model include:
[0332] One example is, "Please explain the operating procedures for the new machine."
[0333] With the above configuration, it is possible to build a system that effectively records and shares on-site work knowledge and promotes the utilization of knowledge throughout the organization.
[0334] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0335] Step 1:
[0336] The device uses a microphone and camera to acquire user audio and video information. The input is the user's real-time audio and video. This information is converted into digital data. Specifically, it records the user verbally explaining the procedures while operating the machine.
[0337] Step 2:
[0338] The terminal transmits acquired audio and video data to the server in real time. The input is digitized audio and video, and the output is data received on the server side. This transmission takes place over a stable network.
[0339] Step 3:
[0340] The server inputs the received audio data into a speech recognition engine and converts it into text information. Specifically, it analyzes the audio information using natural language processing technology and documents it. The output is a written representation of what the user said.
[0341] Step 4:
[0342] The server analyzes video data using computer vision algorithms. Using the video data as input, the server identifies user actions and key movements related to a procedure. The output is data representing what actions were performed. For example, a scene where the user presses a button is analyzed.
[0343] Step 5:
[0344] The server automatically generates procedures based on the analyzed audio and video information. Here, text and action data are integrated and formatted into easy-to-understand procedures. The output is a document of the generated business procedures.
[0345] Step 6:
[0346] The server stores the generated procedures in a database and prepares them for presentation to the user. Users can access this information via their terminal as needed. The saved procedures become an asset that other users can also refer to.
[0347] Step 7:
[0348] When a user has a question during work, they can ask it via their terminal. The input is a voice question. The terminal sends this question to a speech recognition engine, which then converts it into text and sends the text to the server.
[0349] Step 8:
[0350] The server analyzes the received question and searches for the most relevant information from its database and knowledge base. Based on the results, the server generates an answer. The output is the answer information to the question.
[0351] Step 9:
[0352] The terminal receives responses sent from the server and provides them to the user in text or audio format. Based on the information obtained by the user, tasks can be carried out smoothly.
[0353] (Application Example 1)
[0354] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0355] In manufacturing environments, there is a need to efficiently understand and improve work procedures and standard operations. This requires the introduction of new work procedures and efficiency methods, but traditional methods have made it difficult to generate and update procedures in real time. Furthermore, prompt and accurate responses to worker inquiries are required, but achieving this necessitates the management and provision of vast amounts of information. In response to these circumstances, there is a need for a system that provides efficient and up-to-date information, thereby improving work efficiency.
[0356] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0357] In this invention, the server includes a device for acquiring audio and video information, a process for analyzing the acquired audio information and converting it into text information, and a process for analyzing the acquired video information and recognizing the work. This makes it possible to automatically generate work procedures in real time, store and present the latest work procedures. It also records the actions of workers, supports work efficiency, and can effectively respond to user inquiries.
[0358] "Audio information" refers to voice and sound data acquired through acoustic sensors.
[0359] "Visual information" refers to visual data acquired using visual sensors.
[0360] A "device" is a piece of equipment used to acquire and process audio and video data.
[0361] "Analysis" is the process of converting acquired data into a meaningful format.
[0362] "Textual information" refers to audio data that has been represented as text.
[0363] "Work" refers to a series of processes or operations carried out in a factory or similar facility.
[0364] "Processing" refers to specific functions or calculations performed by a device or server.
[0365] A "procedure" is a series of steps required to perform a task or operation.
[0366] "Storage" refers to the act of saving generated information or data.
[0367] "Presentation" means providing stored information in a format that users can verify.
[0368] "Efficiency improvement" means aiming to perform tasks and processes more effectively with less time and effort.
[0369] A "server" is a computer system that performs data analysis and storage over a network.
[0370] "User" refers to a person who operates or uses the system.
[0371] A "question" refers to a question or problem that a user poses to the system.
[0372] A "response" is the information or solution that a system provides in response to a user's question.
[0373] This invention provides a system aimed at improving efficiency and standardizing work in factories and manufacturing sites. The server utilizes various hardware and software to acquire and analyze audio and video information. Specifically, the server uses a microphone to acquire audio information and a camera to acquire video information. This data is transmitted to the server in real time.
[0374] The server converts speech information into text using speech recognition software (e.g., Google Cloud Speech-to-Text). For video information, computer vision algorithms (e.g., OpenCV, TensorFlow) are used to analyze the work and recognize characteristic movements. This allows for the extraction of efficient work procedures and actions.
[0375] When a user performs a task, the optimal work procedure is presented based on the situation unfolding in real time via the terminal. This terminal uses a smart device to immediately provide the information the worker needs. Furthermore, when a user asks a question, a generative AI model is used to quickly generate and present the most appropriate answer.
[0376] As a concrete example, when a new assembly line is introduced, the system records and analyzes the workers' actions. It then presents optimized work procedures and shares them in a format that other workers can refer to. If a user prompts with "Please tell me the most efficient procedure for this task," the system can provide the most efficient procedure based on its accumulated data.
[0377] In this way, the collaboration between servers, terminals, and users promotes the sharing and standardization of knowledge within the factory, leading to improved work efficiency.
[0378] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0379] Step 1:
[0380] The terminal captures audio and video information at the work site. The input consists of the worker's actions and explanations, which are captured in real time by a microphone and camera and sent to the server as digital data. The output is the conversion of analog data into digital format, ready to be sent to the server.
[0381] Step 2:
[0382] The server converts received audio data into text using speech recognition software. The input is audio data sent from the terminal. A speech recognition engine (e.g., Google Cloud Speech-to-Text) is used to analyze the audio data and output it as text. As a result, the audio information is stored on the server in text format.
[0383] Step 3:
[0384] The server analyzes video data using computer vision algorithms to recognize work procedures and actions. The input is video data sent from the terminal. Using tools such as OpenCV and TensorFlow, it performs motion analysis frame by frame of the video. This identifies the characteristics of the actions and outputs them as procedure information.
[0385] Step 4:
[0386] The server automatically generates the optimal work procedure based on the analyzed text and video information. The input consists of the text and video procedure information obtained from the previous analysis. This information is integrated to generate an effective work procedure, which is then stored in a database. As output, the latest work procedure is documented and recorded as a manual.
[0387] Step 5:
[0388] Users can ask questions through their devices. The input consists of user voice or text queries. The device analyzes this input and sends it to the server. The server uses a generative AI model to extract information from its knowledge base and generate the optimal answer.
[0389] Step 6:
[0390] The terminal receives a response from the server and presents it to the user. The input is the response generated by the server, which is then provided to the user as text or audio. As a result, the user can obtain the necessary information in real time.
[0391] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0392] In implementing this invention, in addition to acquiring audio and video data, it is necessary to incorporate an emotion engine for recognizing the user's emotions. The emotion engine analyzes the user's emotional state from the tone and intonation of the voice, facial expressions and gestures in the video, etc. This emotion recognition technology enables the provision of adaptive information that takes into account the user's psychological state during learning and work.
[0393] Data acquisition
[0394] User: Performs daily tasks, and audio and video are collected through the device. Users do not need to perform complex operations and can focus on their work naturally.
[0395] Terminal: Inputs audio and video data in real time and simultaneously analyzes emotional data through an emotion engine.
[0396] Data and sentiment analysis
[0397] Server: Audio data is converted into text using a speech recognition algorithm, and video data is analyzed using an action recognition algorithm to identify the content of the work. Furthermore, the user's current emotional state is identified by an emotion engine. This enables data analysis linked to emotions.
[0398] Procedure generation and application
[0399] Server: Automatically generates business procedures based on analysis results and adjusts the content according to the user's emotions. For example, if it is determined that the user is confused, the approach will change to explain the procedure in more detail.
[0400] Server: The generated procedures are stored in a database and can be accessed by users as needed. Furthermore, information optimized by emotion is displayed to aid learners' understanding.
[0401] Information provision and feedback
[0402] User: If the user has any questions or uncertainties during the process, the emotion engine will provide answers that take into account their emotional state at the time. This allows the user to receive better support.
[0403] Terminal: Sends data to the server at the appropriate time and presents emotion-responsive feedback to the user in voice or text.
[0404] This system, which incorporates emotion recognition, provides users with detailed support in their work and learning, thereby enabling more effective knowledge acquisition and task execution.
[0405] The following describes the processing flow.
[0406] Step 1:
[0407] User: Start the terminal and begin normal work. Explore new procedures and tasks as needed.
[0408] Terminal: Uses a microphone and camera to record and video of the user in real time, and sends the data to the server.
[0409] Step 2:
[0410] Server: Receives audio data and converts it to text using a speech recognition engine. This process includes noise filtering to accurately analyze the spoken content.
[0411] Server: Simultaneously, it analyzes video data using computer vision algorithms to recognize user actions. It identifies the type and sequence of actions and stores the information in a database.
[0412] Step 3:
[0413] Server: Uses an emotion engine to analyze the user's emotional state from received audio and video data. Specifically, it identifies emotions based on the tone of voice and facial expressions and gestures in the video.
[0414] Server: Combines sentiment analysis results with other analysis data to understand the user's current situation and the necessary steps.
[0415] Step 4:
[0416] Server: Automatically generates business procedures based on analysis results. Adjusts the level of detail and explanation of the procedures according to the user's emotional state. For example, if the user is nervous, the procedures are presented more clearly and step-by-step.
[0417] Server: Stores the generated procedures in a database and updates them as needed. The procedures must always be up-to-date for daily operations.
[0418] Step 5:
[0419] User: If questions arise during the process or if additional information is needed, ask them at the terminal.
[0420] Step 6:
[0421] Terminal: Receives user questions via voice, performs speech recognition again to convert them to text, and sends them to the server along with sentiment data.
[0422] Step 7:
[0423] Server: Analyzes the question content and submitted sentiment data, and searches the database for relevant information. In particular, it selects information in a way that takes the user's emotions into consideration and determines the steps to be provided.
[0424] Terminal: The terminal presents information selected by the server to the user via voice or text, and provides a relaxed learning environment by responding in a way that is sensitive to emotions.
[0425] In this way, the system aims to improve the quality of work and learning by utilizing emotion recognition capabilities to provide users with information in the most optimal format.
[0426] (Example 2)
[0427] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0428] In today's work and learning environments, users are required to efficiently handle diverse data. However, systems that understand users' psychological and emotional states and provide optimal information accordingly are still insufficient, resulting in insufficient improvements in user efficiency and satisfaction. Therefore, there is a need for means to enable flexible responses that take into account users' emotional states, thereby improving both work efficiency and user experience.
[0429] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0430] In this invention, the server includes means for recognizing the user's emotional state based on audio and video data, means for automatically generating and adjusting procedures based on the analyzed audio and video data and the user's emotional state, and means for analyzing the user's questions and emotional state and providing optimal information. This makes it possible to provide information adaptively according to the user's emotions and improve work efficiency and learning efficiency.
[0431] "Voice data" refers to information that records user speech and ambient sounds, and serves as the basis for speech recognition and emotion analysis.
[0432] "Video data" refers to visual information that records the user's movements and facial expressions, and is used for motion recognition and emotion analysis.
[0433] "Text data" refers to information obtained by analyzing audio data and converting it into written text, representing the content of the user's speech in written form.
[0434] "Motion recognition" is a technology that identifies a user's actions and gestures by analyzing video data.
[0435] "Emotional state" refers to the psychological state that indicates the user's emotions, and means the psychological response that can be inferred from audio and video.
[0436] A "procedure" is a set of steps that users need to complete their tasks or learning, and it is automatically generated and provided to the user.
[0437] "Automatic generation" refers to the process by which a system generates procedures and information based on data, without manual intervention by humans.
[0438] "Optimal information" refers to information that contains necessary and appropriate content, tailored to the user's situation and emotional state.
[0439] To implement this invention, a system is required for real-time acquisition, analysis, and provision of information to the user of audio and video data. The system is configured as follows.
[0440] Data acquisition
[0441] Terminal: As users perform tasks or studies, audio and video data are collected in real time by the camera and microphone built into the terminal. Users do not need to perform any special operations in a natural work environment.
[0442] Data and emotion recognition
[0443] Server: Audio data transmitted from the terminal is converted into text data using a speech recognition algorithm (e.g., Google Speech-to-Text). Video data can be analyzed to capture user actions using a motion recognition algorithm (e.g., OpenCV), and the user's emotional state can be estimated using an emotion engine (e.g., an emotion analysis tool). This allows the server to understand the user's psychological state.
[0444] Procedure generation and information provision
[0445] Server: Based on the analysis of the user's voice, video, and emotional state, it automatically generates optimized work procedures. For example, if it determines that the user is confused, it provides detailed and specific instructions. These generated procedures are stored in a database, allowing the user to access them as needed and continue their work.
[0446] Terminal: Receives instructions from the server and provides feedback tailored to the user's current emotional state. If voice guidance is appropriate, it uses speech synthesis technology (e.g., a speech synthesis engine) to provide voice guidance to the user. Alternatively, instructions can be displayed on the screen in text format.
[0447] Specific example
[0448] For example, when a user encounters a question on their device while learning how to use new software, the system detects confusion based on the user's tone of voice and video, and provides detailed step-by-step instructions. If a positive emotional state is detected, only an overview is provided, leaving room for the user to continue learning on their own.
[0449] Example of a prompt
[0450] "Analyze the user's current audio and video data, recognize their emotions, and generate instructions. Output a step-by-step guide if the user becomes confused."
[0451] This system allows users to receive individually optimized information, enabling them to efficiently advance their work and learning.
[0452] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0453] Step 1:
[0454] Data acquisition
[0455] Terminal: The terminal acquires audio and video data from the user in real time. Inputs include audio collected by the microphone and video recorded by the camera. These data are output as digital audio signals and video frames.
[0456] Step 2:
[0457] Data transfer and reception
[0458] Terminal: Transmits acquired audio and video data to the server via an internet connection.
[0459] Server: Receives data sent from the terminal and prepares the audio data for processing with a speech recognition algorithm. The output consists of two streams: audio data and video data.
[0460] Step 3:
[0461] Voice analysis
[0462] Server: Inputs audio data into a speech recognition algorithm (e.g., speech recognition engine) and converts it into text data. This is achieved by extracting linguistic characteristics from the waveform of the audio signal and outputting them as corresponding text strings.
[0463] Step 4:
[0464] Video analysis
[0465] Server: The received video data is input into the motion recognition algorithm. The algorithm analyzes the user's movements and facial expressions from the video frames and extracts corresponding motion data. The output is descriptor data related to the physical actions performed by the user.
[0466] Step 5:
[0467] Emotion analysis
[0468] Server: Uses an emotion engine to estimate the user's emotional state based on analysis of audio waveforms, text data, and video. By passing the input data through an emotion classification model, the user's state is output as an emotion tag (e.g., joy, confusion).
[0469] Step 6:
[0470] Automatic generation and adjustment of procedures
[0471] Server: Based on the analysis results, the server starts a process to automatically generate business procedures. By inputting data about the user's actions and emotions as prompts to the generating AI model, it outputs procedures that have been adjusted according to the emotions.
[0472] Step 7:
[0473] Information provision and feedback
[0474] Terminal: Receives instructions and feedback generated from the server and presents them to the user. Feedback is output as audio or visual messages depending on the user's emotional state. This allows the user to work smoothly based on optimized procedures.
[0475] (Application Example 2)
[0476] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0477] In on-site work, providing instructions and support that ignore the emotional state of workers makes it difficult to effectively improve work efficiency and reduce psychological burden. In particular, when workers are feeling fatigued or confused, conventional systems may fail to provide appropriate support, potentially leading to decreased productivity and safety. Therefore, there is a need for a system that can analyze workers' emotions in real time and provide appropriate instructions and information tailored to their state.
[0478] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0479] In this invention, the server includes means for acquiring audio and video information, means for analyzing the user's emotional state, and means for adjusting and providing instructions based on the analyzed emotional state. This makes it possible to grasp the worker's emotions in real time and provide instructions and support appropriate to that emotional state.
[0480] "Audio information" refers to data consisting of human speech and ambient sounds, acquired using acoustic sensors such as microphones.
[0481] "Visual information" refers to visual data of objects and people acquired through cameras and other image sensors.
[0482] "Text data" refers to data in text format obtained by converting audio information into text.
[0483] "Means of recognizing motion" refers to technologies that analyze video information and detect specific patterns or actions related to the movement of people or objects.
[0484] "Means for automatically generating instructions" refers to a process that automatically generates appropriate work procedures and instructions based on the results of analysis of audio information, video information, emotional states, etc.
[0485] "Means of analyzing emotional states" refer to algorithms or engines that determine a person's emotions and psychological state from their tone of voice, facial expressions, and gestures.
[0486] "Means for adjusting and providing generated instructions" refers to a process that optimizes initially generated instructions according to the user's emotional state and presents them to the worker at the appropriate time.
[0487] This invention is a system that provides appropriate support tailored to the emotions of workers in the workplace. The system has the function of acquiring audio and video information, analyzing it, and recognizing actions and emotional states.
[0488] First, the device uses a microphone and camera to acquire audio and video information in real time. The audio information is converted into text data by a speech recognition algorithm (e.g., Google Cloud Speech-to-Text). The video information is used to identify the work being done using an action recognition algorithm.
[0489] Subsequently, the server analyzes the user's emotional state from voice tone and facial expressions obtained using an emotion engine (e.g., Microsoft Azure Emotion API). Based on the emotional state and behavioral information, the server automatically generates and adjusts instructions.
[0490] The analysis results and adjusted instructions are stored in a database and provided to the user in voice or text format at the appropriate time. For example, if a worker is confused, the system can provide specific instructions such as, "Next, use the F panel and attach it to the A base. When doing so, use the right side of the A base."
[0491] The following are examples of prompts to input into the generative AI model.
[0492] "Analyze the emotions captured from real-time audio and video data to identify at which stage of the task the user is experiencing difficulties. Then, propose work procedures or support tailored to their emotional state."
[0493] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0494] Step 1:
[0495] The terminal uses a microphone and camera to acquire audio and video information from the work environment in real time. The audio information includes the worker's speech and ambient sounds, while the video information is video captured by the camera. This data is transmitted directly to the server.
[0496] Step 2:
[0497] The server converts the acquired audio information into text data using a speech recognition algorithm (e.g., Google Cloud Speech-to-Text). The input is audio information, and the output is string data obtained by analyzing the audio. This process records the worker's speech in text.
[0498] Step 3:
[0499] The server analyzes the acquired video information using a motion recognition algorithm to identify the work content and actions. The input is video information, and the output is data related to the actions (e.g., specific actions performed by the worker). This allows for an accurate understanding of what actions the worker is performing.
[0500] Step 4:
[0501] The server analyzes the emotional state using an emotion engine (e.g., Microsoft Azure Emotion API) based on voice tone and facial expressions. Input is audio and video information, and output is emotional state data (e.g., joy, confusion, fatigue). This process identifies emotions using voice tone, speaking style, and facial expressions.
[0502] Step 5:
[0503] The server automatically generates appropriate instructions based on analyzed voice data, motion data, and emotional state data. The input is the data set obtained in the previous step, and the output is specific instructions to be presented to the user. These instructions are adjusted according to the worker's emotional state.
[0504] Step 6:
[0505] The user receives instructions from the server and performs the task. Instructions are provided in audio or text format and include specific guidance to help the user continue the task. Care is taken to ensure that the instructions are presented in a way that is easy for the user to understand.
[0506] Step 7:
[0507] The terminal continuously sends data to the server and updates it whenever processing is complete or new input is required. This enables real-time feedback that responds to changes in the worker's emotional state and work content.
[0508] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0509] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0510] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0511] [Third Embodiment]
[0512] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0513] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0514] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0515] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0516] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0517] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0518] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0519] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0520] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0521] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0522] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0523] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0524] To implement this invention, it is first necessary to set up an environment for acquiring audio and video data. This requires a terminal equipped with a microphone for capturing audio and a camera for recording video. The terminal is responsible for recording the user's work in real time and transmitting that data to a server.
[0525] Step 1: Data Acquisition
[0526] User: Performs tasks while carrying a device. Detailed recording is possible, especially when performing new tasks or operating new equipment.
[0527] Terminal: Uses a microphone and camera to collect audio and video during work and transmits the data to the server in real time.
[0528] Step 2: Data Analysis
[0529] Server: The received audio data is analyzed by the speech recognition engine and converted into text data. This makes it clear what the user is saying.
[0530] Server: Video data is analyzed using computer vision algorithms to identify work procedures and characteristic movements.
[0531] Step 3: Automatic generation and storage of procedures
[0532] Server: Based on the obtained text data and video analysis information, work procedures are automatically generated. These procedures are documented in an easy-to-understand format and stored in a database.
[0533] Server: The generated manual is kept as the latest procedure at that time and is updated as needed.
[0534] Step 4: Information Provision
[0535] User: If any questions arise during the work process, users can ask them via their terminal.
[0536] Terminal: When a question is received via voice, it converts it to text using speech recognition and sends it to the server.
[0537] Server: Analyzes the question content and extracts appropriate information from manual databases and knowledge bases.
[0538] Terminal: The optimal answer is provided to the user in text or voice.
[0539] This embodiment allows users to naturally incorporate their knowledge into the system without any special operations, and provide it as useful information to future learners. This system effectively incorporates on-the-ground knowledge and promotes knowledge sharing throughout the organization.
[0540] The following describes the processing flow.
[0541] Step 1:
[0542] User: Start work and turn on the terminal. Perform normal tasks without making any special settings.
[0543] Terminal: Activates microphone and camera to record and videotape the user's work in real time. Recorded and videotaped data is sent to the server with security measures in place.
[0544] Step 2:
[0545] Server: The server analyzes the received audio data through a speech recognition engine and extracts the spoken content as text data. This process uses techniques to reduce audio noise and convert it into accurate text strings.
[0546] Server: Video data is analyzed by computer vision algorithms to recognize specific user actions and operations. For example, everyday actions such as pressing buttons on equipment or pulling levers are identified.
[0547] Step 3:
[0548] Server: Integrates the results of audio and video analysis and automatically generates work procedure manuals. The generated manuals are made easy to read using natural language generation technology and clearly explain information relevant to the work.
[0549] Server: Procedure manuals are stored in a database and made accessible within the organization. These manuals are continuously updated and used as the most up-to-date work procedures throughout the organization.
[0550] Step 4:
[0551] User: If questions or uncertainties arise during work, users can ask them via voice commands.
[0552] Terminal: Receives questions as voice input, performs speech recognition to convert them into text data, and sends them to the server.
[0553] Step 5:
[0554] Server: Analyzes questions submitted by users and extracts appropriate answers from relevant manuals and databases. Specifically, it aims to select the most relevant information to the question and provide it in a user-friendly format.
[0555] Terminal: Displays responses sent from the server to the user in audio or text format. This allows the user to quickly access solutions and work procedures.
[0556] (Example 1)
[0557] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0558] In recording and communicating work procedures on-site, traditional methods often relied on individual experience and memory, leading to challenges in the reproducibility and sharing of procedures. Furthermore, real-time information retrieval and question-and-answer sessions were not possible, preventing the provision of highly efficient support.
[0559] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0560] In this invention, the server includes a structure for acquiring audio and video information, a structure for automatically generating procedures based on the analyzed audio and video information, and a structure for answering questions to resolve information deficiencies that users may have during their work. This enables improved recording and reproducibility of on-site work procedures, real-time information provision, and efficient work support.
[0561] "Auditory information" refers to data obtained through changes in sound, and in particular, information that includes human voices.
[0562] "Visual information" refers to visual data acquired through visual devices such as cameras.
[0563] "Analysis" is the act of processing acquired data and extracting meaning and characteristics based on that data.
[0564] "Textual information" refers to audio data expressed as text, a format that allows for visual confirmation of the audio content.
[0565] A "structure that recognizes motion" is a structure that has the function of identifying and understanding the movements of people and objects contained in video data.
[0566] A "structure for automatically generating procedures" is a system that mechanically creates procedures and methods based on analyzed information.
[0567] A "storage and presentation structure" is a mechanism for saving generated information and displaying or providing it as needed.
[0568] A "question and answer structure" is a structure that receives questions from users, searches for answers to those questions, and provides responses.
[0569] To implement the present invention, it is necessary to create an environment that can efficiently acquire and analyze audio and video information. Using a terminal equipped with a microphone and camera is appropriate for this purpose. The terminal is responsible for capturing the user's actions and voice in real time and transmitting that data to the server.
[0570] In terms of specific hardware, a high-resolution camera and a microphone capable of capturing clear audio are recommended. A stable network connection is also necessary for processing and transmitting data in real time.
[0571] Data transmitted from the terminal is processed on the server. The server uses speech recognition software to convert speech information into text. A natural language processing engine is utilized to accurately extract text data from the speech. Furthermore, computer vision algorithms are used to identify important actions from the video information and utilize this information for procedure generation.
[0572] The generated procedures are stored in a database and provided to users as needed. Users can access this information in real time via their terminals and receive immediate support through the Q&A service if they have any questions during their work.
[0573] As a concrete example, when recording the steps required to operate a new machine, the user activates the terminal and provides instructions while operating it. This explanation and operation are recorded simultaneously, and after analysis on the server, the procedure is automatically generated. This procedure can then be referenced by other users.
[0574] Examples of prompts for a generative AI model include:
[0575] One example is, "Please explain the operating procedures for the new machine."
[0576] With the above configuration, it is possible to build a system that effectively records and shares on-site work knowledge and promotes the utilization of knowledge throughout the organization.
[0577] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0578] Step 1:
[0579] The device uses a microphone and camera to acquire user audio and video information. The input is the user's real-time audio and video. This information is converted into digital data. Specifically, it records the user verbally explaining the procedures while operating the machine.
[0580] Step 2:
[0581] The terminal transmits acquired audio and video data to the server in real time. The input is digitized audio and video, and the output is data received on the server side. This transmission takes place over a stable network.
[0582] Step 3:
[0583] The server inputs the received audio data into a speech recognition engine and converts it into text information. Specifically, it analyzes the audio information using natural language processing technology and documents it. The output is a written representation of what the user said.
[0584] Step 4:
[0585] The server analyzes video data using computer vision algorithms. Using the video data as input, the server identifies user actions and key movements related to a procedure. The output is data representing what actions were performed. For example, a scene where the user presses a button is analyzed.
[0586] Step 5:
[0587] The server automatically generates procedures based on the analyzed audio and video information. Here, text and action data are integrated and formatted into easy-to-understand procedures. The output is a document of the generated business procedures.
[0588] Step 6:
[0589] The server stores the generated procedures in a database and prepares them for presentation to the user. Users can access this information via their terminal as needed. The saved procedures become an asset that other users can also refer to.
[0590] Step 7:
[0591] When a user has a question during work, they can ask it via their terminal. The input is a voice question. The terminal sends this question to a speech recognition engine, which then converts it into text and sends the text to the server.
[0592] Step 8:
[0593] The server analyzes the received question and searches for the most relevant information from its database and knowledge base. Based on the results, the server generates an answer. The output is the answer information to the question.
[0594] Step 9:
[0595] The terminal receives responses sent from the server and provides them to the user in text or audio format. Based on the information obtained by the user, tasks can be carried out smoothly.
[0596] (Application Example 1)
[0597] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0598] In manufacturing environments, there is a need to efficiently understand and improve work procedures and standard operations. This requires the introduction of new work procedures and efficiency methods, but traditional methods have made it difficult to generate and update procedures in real time. Furthermore, prompt and accurate responses to worker inquiries are required, but achieving this necessitates the management and provision of vast amounts of information. In response to these circumstances, there is a need for a system that provides efficient and up-to-date information, thereby improving work efficiency.
[0599] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0600] In this invention, the server includes a device for acquiring audio and video information, a process for analyzing the acquired audio information and converting it into text information, and a process for analyzing the acquired video information and recognizing the work. This makes it possible to automatically generate work procedures in real time, store and present the latest work procedures. It also records the actions of workers, supports work efficiency, and can effectively respond to user inquiries.
[0601] "Audio information" refers to voice and sound data acquired through acoustic sensors.
[0602] "Visual information" refers to visual data acquired using visual sensors.
[0603] A "device" is a piece of equipment used to acquire and process audio and video data.
[0604] "Analysis" is the process of converting acquired data into a meaningful format.
[0605] "Textual information" refers to audio data that has been represented as text.
[0606] "Work" refers to a series of processes or operations carried out in a factory or similar facility.
[0607] "Processing" refers to specific functions or calculations performed by a device or server.
[0608] A "procedure" is a series of steps required to perform a task or operation.
[0609] "Storage" refers to the act of saving generated information or data.
[0610] "Presentation" means providing stored information in a format that users can verify.
[0611] "Efficiency improvement" means aiming to perform tasks and processes more effectively with less time and effort.
[0612] A "server" is a computer system that performs data analysis and storage over a network.
[0613] "User" refers to a person who operates or uses the system.
[0614] A "question" refers to a question or problem that a user poses to the system.
[0615] A "response" is the information or solution that a system provides in response to a user's question.
[0616] This invention provides a system aimed at improving efficiency and standardizing work in factories and manufacturing sites. The server utilizes various hardware and software to acquire and analyze audio and video information. Specifically, the server uses a microphone to acquire audio information and a camera to acquire video information. This data is transmitted to the server in real time.
[0617] The server converts speech information into text using speech recognition software (e.g., Google Cloud Speech-to-Text). For video information, computer vision algorithms (e.g., OpenCV, TensorFlow) are used to analyze the work and recognize characteristic movements. This allows for the extraction of efficient work procedures and actions.
[0618] When a user performs a task, the optimal work procedure is presented based on the situation unfolding in real time via the terminal. This terminal uses a smart device to immediately provide the information the worker needs. Furthermore, when a user asks a question, a generative AI model is used to quickly generate and present the most appropriate answer.
[0619] As a concrete example, when a new assembly line is introduced, the system records and analyzes the workers' actions. It then presents optimized work procedures and shares them in a format that other workers can refer to. If a user prompts with "Please tell me the most efficient procedure for this task," the system can provide the most efficient procedure based on its accumulated data.
[0620] In this way, the collaboration between servers, terminals, and users promotes the sharing and standardization of knowledge within the factory, leading to improved work efficiency.
[0621] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0622] Step 1:
[0623] The terminal captures audio and video information at the work site. The input consists of the worker's actions and explanations, which are captured in real time by a microphone and camera and sent to the server as digital data. The output is the conversion of analog data into digital format, ready to be sent to the server.
[0624] Step 2:
[0625] The server converts received audio data into text using speech recognition software. The input is audio data sent from the terminal. A speech recognition engine (e.g., Google Cloud Speech-to-Text) is used to analyze the audio data and output it as text. As a result, the audio information is stored on the server in text format.
[0626] Step 3:
[0627] The server analyzes video data using computer vision algorithms to recognize work procedures and actions. The input is video data sent from the terminal. Using tools such as OpenCV and TensorFlow, it performs motion analysis frame by frame of the video. This identifies the characteristics of the actions and outputs them as procedure information.
[0628] Step 4:
[0629] The server automatically generates the optimal work procedure based on the analyzed text and video information. The input consists of the text and video procedure information obtained from the previous analysis. This information is integrated to generate an effective work procedure, which is then stored in a database. As output, the latest work procedure is documented and recorded as a manual.
[0630] Step 5:
[0631] Users can ask questions through their devices. The input consists of user voice or text queries. The device analyzes this input and sends it to the server. The server uses a generative AI model to extract information from its knowledge base and generate the optimal answer.
[0632] Step 6:
[0633] The terminal receives a response from the server and presents it to the user. The input is the response generated by the server, which is then provided to the user as text or audio. As a result, the user can obtain the necessary information in real time.
[0634] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0635] In implementing this invention, in addition to acquiring audio and video data, it is necessary to incorporate an emotion engine for recognizing the user's emotions. The emotion engine analyzes the user's emotional state from the tone and intonation of the voice, facial expressions and gestures in the video, etc. This emotion recognition technology enables the provision of adaptive information that takes into account the user's psychological state during learning and work.
[0636] Data acquisition
[0637] User: Performs daily tasks, and audio and video are collected through the device. Users do not need to perform complex operations and can focus on their work naturally.
[0638] Terminal: Inputs audio and video data in real time and simultaneously analyzes emotional data through an emotion engine.
[0639] Data and sentiment analysis
[0640] Server: Audio data is converted into text using a speech recognition algorithm, and video data is analyzed using an action recognition algorithm to identify the content of the work. Furthermore, the user's current emotional state is identified by an emotion engine. This enables data analysis linked to emotions.
[0641] Procedure generation and application
[0642] Server: Automatically generates business procedures based on analysis results and adjusts the content according to the user's emotions. For example, if it is determined that the user is confused, the approach will change to explain the procedure in more detail.
[0643] Server: The generated procedures are stored in a database and can be accessed by users as needed. Furthermore, information optimized by emotion is displayed to aid learners' understanding.
[0644] Information provision and feedback
[0645] User: If the user has any questions or uncertainties during the process, the emotion engine will provide answers that take into account their emotional state at the time. This allows the user to receive better support.
[0646] Terminal: Sends data to the server at the appropriate time and presents emotion-responsive feedback to the user in voice or text.
[0647] This system, which incorporates emotion recognition, provides users with detailed support in their work and learning, thereby enabling more effective knowledge acquisition and task execution.
[0648] The following describes the processing flow.
[0649] Step 1:
[0650] User: Start the terminal and begin normal work. Explore new procedures and tasks as needed.
[0651] Terminal: Uses a microphone and camera to record and video of the user in real time, and sends the data to the server.
[0652] Step 2:
[0653] Server: Receives audio data and converts it to text using a speech recognition engine. This process includes noise filtering to accurately analyze the spoken content.
[0654] Server: Simultaneously, it analyzes video data using computer vision algorithms to recognize user actions. It identifies the type and sequence of actions and stores the information in a database.
[0655] Step 3:
[0656] Server: Uses an emotion engine to analyze the user's emotional state from received audio and video data. Specifically, it identifies emotions based on the tone of voice and facial expressions and gestures in the video.
[0657] Server: Combines sentiment analysis results with other analysis data to understand the user's current situation and the necessary steps.
[0658] Step 4:
[0659] Server: Automatically generates business procedures based on analysis results. Adjusts the level of detail and explanation of the procedures according to the user's emotional state. For example, if the user is nervous, the procedures are presented more clearly and step-by-step.
[0660] Server: Stores the generated procedures in a database and updates them as needed. The procedures must always be up-to-date for daily operations.
[0661] Step 5:
[0662] User: If questions arise during the process or if additional information is needed, ask them at the terminal.
[0663] Step 6:
[0664] Terminal: Receives user questions via voice, performs speech recognition again to convert them to text, and sends them to the server along with sentiment data.
[0665] Step 7:
[0666] Server: Analyzes the question content and submitted sentiment data, and searches the database for relevant information. In particular, it selects information in a way that takes the user's emotions into consideration and determines the steps to be provided.
[0667] Terminal: The terminal presents information selected by the server to the user via voice or text, and provides a relaxed learning environment by responding in a way that is sensitive to emotions.
[0668] In this way, the system aims to improve the quality of work and learning by utilizing emotion recognition capabilities to provide users with information in the most optimal format.
[0669] (Example 2)
[0670] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0671] In today's work and learning environments, users are required to efficiently handle diverse data. However, systems that understand users' psychological and emotional states and provide optimal information accordingly are still insufficient, resulting in insufficient improvements in user efficiency and satisfaction. Therefore, there is a need for means to enable flexible responses that take into account users' emotional states, thereby improving both work efficiency and user experience.
[0672] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0673] In this invention, the server includes means for recognizing the user's emotional state based on audio and video data, means for automatically generating and adjusting procedures based on the analyzed audio and video data and the user's emotional state, and means for analyzing the user's questions and emotional state and providing optimal information. This makes it possible to provide information adaptively according to the user's emotions and improve work efficiency and learning efficiency.
[0674] "Voice data" refers to information that records user speech and ambient sounds, and serves as the basis for speech recognition and emotion analysis.
[0675] "Video data" refers to visual information that records the user's movements and facial expressions, and is used for motion recognition and emotion analysis.
[0676] "Text data" refers to information obtained by analyzing audio data and converting it into written text, representing the content of the user's speech in written form.
[0677] "Motion recognition" is a technology that identifies a user's actions and gestures by analyzing video data.
[0678] "Emotional state" refers to the psychological state that indicates the user's emotions, and means the psychological response that can be inferred from audio and video.
[0679] A "procedure" is a set of steps that users need to complete their tasks or learning, and it is automatically generated and provided to the user.
[0680] "Automatic generation" refers to the process by which a system generates procedures and information based on data, without manual intervention by humans.
[0681] "Optimal information" refers to information that contains necessary and appropriate content, tailored to the user's situation and emotional state.
[0682] To implement this invention, a system is required for real-time acquisition, analysis, and provision of information to the user of audio and video data. The system is configured as follows.
[0683] Data acquisition
[0684] Terminal: As users perform tasks or studies, audio and video data are collected in real time by the camera and microphone built into the terminal. Users do not need to perform any special operations in a natural work environment.
[0685] Data and emotion recognition
[0686] Server: Audio data transmitted from the terminal is converted into text data using a speech recognition algorithm (e.g., Google Speech-to-Text). Video data can be analyzed to capture user actions using a motion recognition algorithm (e.g., OpenCV), and the user's emotional state can be estimated using an emotion engine (e.g., an emotion analysis tool). This allows the server to understand the user's psychological state.
[0687] Procedure generation and information provision
[0688] Server: Based on the analysis of the user's voice, video, and emotional state, it automatically generates optimized work procedures. For example, if it determines that the user is confused, it provides detailed and specific instructions. These generated procedures are stored in a database, allowing the user to access them as needed and continue their work.
[0689] Terminal: Receives instructions from the server and provides feedback tailored to the user's current emotional state. If voice guidance is appropriate, it uses speech synthesis technology (e.g., a speech synthesis engine) to provide voice guidance to the user. Alternatively, instructions can be displayed on the screen in text format.
[0690] Specific example
[0691] For example, when a user encounters a question on their device while learning how to use new software, the system detects confusion based on the user's tone of voice and video, and provides detailed step-by-step instructions. If a positive emotional state is detected, only an overview is provided, leaving room for the user to continue learning on their own.
[0692] Example of a prompt
[0693] "Analyze the user's current audio and video data, recognize their emotions, and generate instructions. Output a step-by-step guide if the user becomes confused."
[0694] This system allows users to receive individually optimized information, enabling them to efficiently advance their work and learning.
[0695] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0696] Step 1:
[0697] Data acquisition
[0698] Terminal: The terminal acquires audio and video data from the user in real time. Inputs include audio collected by the microphone and video recorded by the camera. These data are output as digital audio signals and video frames.
[0699] Step 2:
[0700] Data transfer and reception
[0701] Terminal: Transmits acquired audio and video data to the server via an internet connection.
[0702] Server: Receives data sent from the terminal and prepares the audio data for processing with a speech recognition algorithm. The output consists of two streams: audio data and video data.
[0703] Step 3:
[0704] Voice analysis
[0705] Server: Inputs audio data into a speech recognition algorithm (e.g., speech recognition engine) and converts it into text data. This is achieved by extracting linguistic characteristics from the waveform of the audio signal and outputting them as corresponding text strings.
[0706] Step 4:
[0707] Video analysis
[0708] Server: The received video data is input into the motion recognition algorithm. The algorithm analyzes the user's movements and facial expressions from the video frames and extracts corresponding motion data. The output is descriptor data related to the physical actions performed by the user.
[0709] Step 5:
[0710] Emotion analysis
[0711] Server: Uses an emotion engine to estimate the user's emotional state based on analysis of audio waveforms, text data, and video. By passing the input data through an emotion classification model, the user's state is output as an emotion tag (e.g., joy, confusion).
[0712] Step 6:
[0713] Automatic generation and adjustment of procedures
[0714] Server: Based on the analysis results, the server starts a process to automatically generate business procedures. By inputting data about the user's actions and emotions as prompts to the generating AI model, it outputs procedures that have been adjusted according to the emotions.
[0715] Step 7:
[0716] Information provision and feedback
[0717] Terminal: Receives instructions and feedback generated from the server and presents them to the user. Feedback is output as audio or visual messages depending on the user's emotional state. This allows the user to work smoothly based on optimized procedures.
[0718] (Application Example 2)
[0719] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0720] In on-site work, providing instructions and support that ignore the emotional state of workers makes it difficult to effectively improve work efficiency and reduce psychological burden. In particular, when workers are feeling fatigued or confused, conventional systems may fail to provide appropriate support, potentially leading to decreased productivity and safety. Therefore, there is a need for a system that can analyze workers' emotions in real time and provide appropriate instructions and information tailored to their state.
[0721] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0722] In this invention, the server includes means for acquiring audio and video information, means for analyzing the user's emotional state, and means for adjusting and providing instructions based on the analyzed emotional state. This makes it possible to grasp the worker's emotions in real time and provide instructions and support appropriate to that emotional state.
[0723] "Audio information" refers to data consisting of human speech and ambient sounds, acquired using acoustic sensors such as microphones.
[0724] "Visual information" refers to visual data of objects and people acquired through cameras and other image sensors.
[0725] "Text data" refers to data in text format obtained by converting audio information into text.
[0726] "Means of recognizing motion" refers to technologies that analyze video information and detect specific patterns or actions related to the movement of people or objects.
[0727] "Means for automatically generating instructions" refers to a process that automatically generates appropriate work procedures and instructions based on the results of analysis of audio information, video information, emotional states, etc.
[0728] "Means of analyzing emotional states" refer to algorithms or engines that determine a person's emotions and psychological state from their tone of voice, facial expressions, and gestures.
[0729] "Means for adjusting and providing generated instructions" refers to a process that optimizes initially generated instructions according to the user's emotional state and presents them to the worker at the appropriate time.
[0730] This invention is a system that provides appropriate support tailored to the emotions of workers in the workplace. The system has the function of acquiring audio and video information, analyzing it, and recognizing actions and emotional states.
[0731] First, the device uses a microphone and camera to acquire audio and video information in real time. The audio information is converted into text data by a speech recognition algorithm (e.g., Google Cloud Speech-to-Text). The video information is used to identify the work being done using an action recognition algorithm.
[0732] Subsequently, the server analyzes the user's emotional state from voice tone and facial expressions obtained using an emotion engine (e.g., Microsoft Azure Emotion API). Based on the emotional state and behavioral information, the server automatically generates and adjusts instructions.
[0733] The analysis results and adjusted instructions are stored in a database and provided to the user in voice or text format at the appropriate time. For example, if a worker is confused, the system can provide specific instructions such as, "Next, use the F panel and attach it to the A base. When doing so, use the right side of the A base."
[0734] The following are examples of prompts to input into the generative AI model.
[0735] "Analyze the emotions captured from real-time audio and video data to identify at which stage of the task the user is experiencing difficulties. Then, propose work procedures or support tailored to their emotional state."
[0736] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0737] Step 1:
[0738] The terminal uses a microphone and camera to acquire audio and video information from the work environment in real time. The audio information includes the worker's speech and ambient sounds, while the video information is video captured by the camera. This data is transmitted directly to the server.
[0739] Step 2:
[0740] The server converts the acquired audio information into text data using a speech recognition algorithm (e.g., Google Cloud Speech-to-Text). The input is audio information, and the output is string data obtained by analyzing the audio. This process records the worker's speech in text.
[0741] Step 3:
[0742] The server analyzes the acquired video information using a motion recognition algorithm to identify the work content and actions. The input is video information, and the output is data related to the actions (e.g., specific actions performed by the worker). This allows for an accurate understanding of what actions the worker is performing.
[0743] Step 4:
[0744] The server analyzes the emotional state using an emotion engine (e.g., Microsoft Azure Emotion API) based on voice tone and facial expressions. Input is audio and video information, and output is emotional state data (e.g., joy, confusion, fatigue). This process identifies emotions using voice tone, speaking style, and facial expressions.
[0745] Step 5:
[0746] The server automatically generates appropriate instructions based on analyzed voice data, motion data, and emotional state data. The input is the data set obtained in the previous step, and the output is specific instructions to be presented to the user. These instructions are adjusted according to the worker's emotional state.
[0747] Step 6:
[0748] The user receives instructions from the server and performs the task. Instructions are provided in audio or text format and include specific guidance to help the user continue the task. Care is taken to ensure that the instructions are presented in a way that is easy for the user to understand.
[0749] Step 7:
[0750] The terminal continuously sends data to the server and updates it whenever processing is complete or new input is required. This enables real-time feedback that responds to changes in the worker's emotional state and work content.
[0751] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0752] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0753] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0754] [Fourth Embodiment]
[0755] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0756] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0757] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0758] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0759] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0760] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0761] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0762] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0763] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0764] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0765] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0766] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0767] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0768] To implement this invention, it is first necessary to set up an environment for acquiring audio and video data. This requires a terminal equipped with a microphone for capturing audio and a camera for recording video. The terminal is responsible for recording the user's work in real time and transmitting that data to a server.
[0769] Step 1: Data Acquisition
[0770] User: Performs tasks while carrying a device. Detailed recording is possible, especially when performing new tasks or operating new equipment.
[0771] Terminal: Uses a microphone and camera to collect audio and video during work and transmits the data to the server in real time.
[0772] Step 2: Data Analysis
[0773] Server: The received audio data is analyzed by the speech recognition engine and converted into text data. This makes it clear what the user is saying.
[0774] Server: Video data is analyzed using computer vision algorithms to identify work procedures and characteristic movements.
[0775] Step 3: Automatic generation and storage of procedures
[0776] Server: Based on the obtained text data and video analysis information, work procedures are automatically generated. These procedures are documented in an easy-to-understand format and stored in a database.
[0777] Server: The generated manual is kept as the latest procedure at that time and is updated as needed.
[0778] Step 4: Information Provision
[0779] User: If any questions arise during the work process, users can ask them via their terminal.
[0780] Terminal: When a question is received via voice, it converts it to text using speech recognition and sends it to the server.
[0781] Server: Analyzes the question content and extracts appropriate information from manual databases and knowledge bases.
[0782] Terminal: The optimal answer is provided to the user in text or voice.
[0783] This embodiment allows users to naturally incorporate their knowledge into the system without any special operations, and provide it as useful information to future learners. This system effectively incorporates on-the-ground knowledge and promotes knowledge sharing throughout the organization.
[0784] The following describes the processing flow.
[0785] Step 1:
[0786] User: Start work and turn on the terminal. Perform normal tasks without making any special settings.
[0787] Terminal: Activates microphone and camera to record and videotape the user's work in real time. Recorded and videotaped data is sent to the server with security measures in place.
[0788] Step 2:
[0789] Server: The server analyzes the received audio data through a speech recognition engine and extracts the spoken content as text data. This process uses techniques to reduce audio noise and convert it into accurate text strings.
[0790] Server: Video data is analyzed by computer vision algorithms to recognize specific user actions and operations. For example, everyday actions such as pressing buttons on equipment or pulling levers are identified.
[0791] Step 3:
[0792] Server: Integrates the results of audio and video analysis and automatically generates work procedure manuals. The generated manuals are made easy to read using natural language generation technology and clearly explain information relevant to the work.
[0793] Server: Procedure manuals are stored in a database and made accessible within the organization. These manuals are continuously updated and used as the most up-to-date work procedures throughout the organization.
[0794] Step 4:
[0795] User: If questions or uncertainties arise during work, users can ask them via voice commands.
[0796] Terminal: Receives questions as voice input, performs speech recognition to convert them into text data, and sends them to the server.
[0797] Step 5:
[0798] Server: Analyzes questions submitted by users and extracts appropriate answers from relevant manuals and databases. Specifically, it aims to select the most relevant information to the question and provide it in a user-friendly format.
[0799] Terminal: Displays responses sent from the server to the user in audio or text format. This allows the user to quickly access solutions and work procedures.
[0800] (Example 1)
[0801] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0802] In recording and communicating work procedures on-site, traditional methods often relied on individual experience and memory, leading to challenges in the reproducibility and sharing of procedures. Furthermore, real-time information retrieval and question-and-answer sessions were not possible, preventing the provision of highly efficient support.
[0803] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0804] In this invention, the server includes a structure for acquiring audio and video information, a structure for automatically generating procedures based on the analyzed audio and video information, and a structure for answering questions to resolve information deficiencies that users may have during their work. This enables improved recording and reproducibility of on-site work procedures, real-time information provision, and efficient work support.
[0805] "Auditory information" refers to data obtained through changes in sound, and in particular, information that includes human voices.
[0806] "Visual information" refers to visual data acquired through visual devices such as cameras.
[0807] "Analysis" is the act of processing acquired data and extracting meaning and characteristics based on that data.
[0808] "Textual information" refers to audio data expressed as text, a format that allows for visual confirmation of the audio content.
[0809] A "structure that recognizes motion" is a structure that has the function of identifying and understanding the movements of people and objects contained in video data.
[0810] A "structure for automatically generating procedures" is a system that mechanically creates procedures and methods based on analyzed information.
[0811] A "storage and presentation structure" is a mechanism for saving generated information and displaying or providing it as needed.
[0812] A "question and answer structure" is a structure that receives questions from users, searches for answers to those questions, and provides responses.
[0813] To implement the present invention, it is necessary to create an environment that can efficiently acquire and analyze audio and video information. Using a terminal equipped with a microphone and camera is appropriate for this purpose. The terminal is responsible for capturing the user's actions and voice in real time and transmitting that data to the server.
[0814] In terms of specific hardware, a high-resolution camera and a microphone capable of capturing clear audio are recommended. A stable network connection is also necessary for processing and transmitting data in real time.
[0815] Data transmitted from the terminal is processed on the server. The server uses speech recognition software to convert speech information into text. A natural language processing engine is utilized to accurately extract text data from the speech. Furthermore, computer vision algorithms are used to identify important actions from the video information and utilize this information for procedure generation.
[0816] The generated procedures are stored in a database and provided to users as needed. Users can access this information in real time via their terminals and receive immediate support through the Q&A service if they have any questions during their work.
[0817] As a concrete example, when recording the steps required to operate a new machine, the user activates the terminal and provides instructions while operating it. This explanation and operation are recorded simultaneously, and after analysis on the server, the procedure is automatically generated. This procedure can then be referenced by other users.
[0818] Examples of prompts for a generative AI model include:
[0819] One example is, "Please explain the operating procedures for the new machine."
[0820] With the above configuration, it is possible to build a system that effectively records and shares on-site work knowledge and promotes the utilization of knowledge throughout the organization.
[0821] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0822] Step 1:
[0823] The device uses a microphone and camera to acquire user audio and video information. The input is the user's real-time audio and video. This information is converted into digital data. Specifically, it records the user verbally explaining the procedures while operating the machine.
[0824] Step 2:
[0825] The terminal transmits acquired audio and video data to the server in real time. The input is digitized audio and video, and the output is data received on the server side. This transmission takes place over a stable network.
[0826] Step 3:
[0827] The server inputs the received audio data into a speech recognition engine and converts it into text information. Specifically, it analyzes the audio information using natural language processing technology and documents it. The output is a written representation of what the user said.
[0828] Step 4:
[0829] The server analyzes video data using computer vision algorithms. Using the video data as input, the server identifies user actions and key movements related to a procedure. The output is data representing what actions were performed. For example, a scene where the user presses a button is analyzed.
[0830] Step 5:
[0831] The server automatically generates procedures based on the analyzed audio and video information. Here, text and action data are integrated and formatted into easy-to-understand procedures. The output is a document of the generated business procedures.
[0832] Step 6:
[0833] The server stores the generated procedures in a database and prepares them for presentation to the user. Users can access this information via their terminal as needed. The saved procedures become an asset that other users can also refer to.
[0834] Step 7:
[0835] When a user has a question during work, they can ask it via their terminal. The input is a voice question. The terminal sends this question to a speech recognition engine, which then converts it into text and sends the text to the server.
[0836] Step 8:
[0837] The server analyzes the received question and searches for the most relevant information from its database and knowledge base. Based on the results, the server generates an answer. The output is the answer information to the question.
[0838] Step 9:
[0839] The terminal receives responses sent from the server and provides them to the user in text or audio format. Based on the information obtained by the user, tasks can be carried out smoothly.
[0840] (Application Example 1)
[0841] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0842] In manufacturing environments, there is a need to efficiently understand and improve work procedures and standard operations. This requires the introduction of new work procedures and efficiency methods, but traditional methods have made it difficult to generate and update procedures in real time. Furthermore, prompt and accurate responses to worker inquiries are required, but achieving this necessitates the management and provision of vast amounts of information. In response to these circumstances, there is a need for a system that provides efficient and up-to-date information, thereby improving work efficiency.
[0843] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0844] In this invention, the server includes a device for acquiring audio and video information, a process for analyzing the acquired audio information and converting it into text information, and a process for analyzing the acquired video information and recognizing the work. This makes it possible to automatically generate work procedures in real time, store and present the latest work procedures. It also records the actions of workers, supports work efficiency, and can effectively respond to user inquiries.
[0845] "Audio information" refers to voice and sound data acquired through acoustic sensors.
[0846] "Visual information" refers to visual data acquired using visual sensors.
[0847] A "device" is a piece of equipment used to acquire and process audio and video data.
[0848] "Analysis" is the process of converting acquired data into a meaningful format.
[0849] "Textual information" refers to audio data that has been represented as text.
[0850] "Work" refers to a series of processes or operations carried out in a factory or similar facility.
[0851] "Processing" refers to specific functions or calculations performed by a device or server.
[0852] A "procedure" is a series of steps required to perform a task or operation.
[0853] "Storage" refers to the act of saving generated information or data.
[0854] "Presentation" means providing stored information in a format that users can verify.
[0855] "Efficiency improvement" means aiming to perform tasks and processes more effectively with less time and effort.
[0856] A "server" is a computer system that performs data analysis and storage over a network.
[0857] "User" refers to a person who operates or uses the system.
[0858] A "question" refers to a question or problem that a user poses to the system.
[0859] A "response" is the information or solution that a system provides in response to a user's question.
[0860] This invention provides a system aimed at improving efficiency and standardizing work in factories and manufacturing sites. The server utilizes various hardware and software to acquire and analyze audio and video information. Specifically, the server uses a microphone to acquire audio information and a camera to acquire video information. This data is transmitted to the server in real time.
[0861] The server converts speech information into text using speech recognition software (e.g., Google Cloud Speech-to-Text). For video information, computer vision algorithms (e.g., OpenCV, TensorFlow) are used to analyze the work and recognize characteristic movements. This allows for the extraction of efficient work procedures and actions.
[0862] When a user performs a task, the optimal work procedure is presented based on the situation unfolding in real time via the terminal. This terminal uses a smart device to immediately provide the information the worker needs. Furthermore, when a user asks a question, a generative AI model is used to quickly generate and present the most appropriate answer.
[0863] As a concrete example, when a new assembly line is introduced, the system records and analyzes the workers' actions. It then presents optimized work procedures and shares them in a format that other workers can refer to. If a user prompts with "Please tell me the most efficient procedure for this task," the system can provide the most efficient procedure based on its accumulated data.
[0864] In this way, the collaboration between servers, terminals, and users promotes the sharing and standardization of knowledge within the factory, leading to improved work efficiency.
[0865] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0866] Step 1:
[0867] The terminal captures audio and video information at the work site. The input consists of the worker's actions and explanations, which are captured in real time by a microphone and camera and sent to the server as digital data. The output is the conversion of analog data into digital format, ready to be sent to the server.
[0868] Step 2:
[0869] The server converts received audio data into text using speech recognition software. The input is audio data sent from the terminal. A speech recognition engine (e.g., Google Cloud Speech-to-Text) is used to analyze the audio data and output it as text. As a result, the audio information is stored on the server in text format.
[0870] Step 3:
[0871] The server analyzes video data using computer vision algorithms to recognize work procedures and actions. The input is video data sent from the terminal. Using tools such as OpenCV and TensorFlow, it performs motion analysis frame by frame of the video. This identifies the characteristics of the actions and outputs them as procedure information.
[0872] Step 4:
[0873] The server automatically generates the optimal work procedure based on the analyzed text and video information. The input consists of the text and video procedure information obtained from the previous analysis. This information is integrated to generate an effective work procedure, which is then stored in a database. As output, the latest work procedure is documented and recorded as a manual.
[0874] Step 5:
[0875] Users can ask questions through their devices. The input consists of user voice or text queries. The device analyzes this input and sends it to the server. The server uses a generative AI model to extract information from its knowledge base and generate the optimal answer.
[0876] Step 6:
[0877] The terminal receives a response from the server and presents it to the user. The input is the response generated by the server, which is then provided to the user as text or audio. As a result, the user can obtain the necessary information in real time.
[0878] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0879] In implementing this invention, in addition to acquiring audio and video data, it is necessary to incorporate an emotion engine for recognizing the user's emotions. The emotion engine analyzes the user's emotional state from the tone and intonation of the voice, facial expressions and gestures in the video, etc. This emotion recognition technology enables the provision of adaptive information that takes into account the user's psychological state during learning and work.
[0880] Data acquisition
[0881] User: Performs daily tasks, and audio and video are collected through the device. Users do not need to perform complex operations and can focus on their work naturally.
[0882] Terminal: Inputs audio and video data in real time and simultaneously analyzes emotional data through an emotion engine.
[0883] Data and sentiment analysis
[0884] Server: Audio data is converted into text using a speech recognition algorithm, and video data is analyzed using an action recognition algorithm to identify the content of the work. Furthermore, the user's current emotional state is identified by an emotion engine. This enables data analysis linked to emotions.
[0885] Procedure generation and application
[0886] Server: Automatically generates business procedures based on analysis results and adjusts the content according to the user's emotions. For example, if it is determined that the user is confused, the approach will change to explain the procedure in more detail.
[0887] Server: The generated procedures are stored in a database and can be accessed by users as needed. Furthermore, information optimized by emotion is displayed to aid learners' understanding.
[0888] Information provision and feedback
[0889] User: If the user has any questions or uncertainties during the process, the emotion engine will provide answers that take into account their emotional state at the time. This allows the user to receive better support.
[0890] Terminal: Sends data to the server at the appropriate time and presents emotion-responsive feedback to the user in voice or text.
[0891] This system, which incorporates emotion recognition, provides users with detailed support in their work and learning, thereby enabling more effective knowledge acquisition and task execution.
[0892] The following describes the processing flow.
[0893] Step 1:
[0894] User: Start the terminal and begin normal work. Explore new procedures and tasks as needed.
[0895] Terminal: Uses a microphone and camera to record and video of the user in real time, and sends the data to the server.
[0896] Step 2:
[0897] Server: Receives audio data and converts it to text using a speech recognition engine. This process includes noise filtering to accurately analyze the spoken content.
[0898] Server: Simultaneously, it analyzes video data using computer vision algorithms to recognize user actions. It identifies the type and sequence of actions and stores the information in a database.
[0899] Step 3:
[0900] Server: Uses an emotion engine to analyze the user's emotional state from received audio and video data. Specifically, it identifies emotions based on the tone of voice and facial expressions and gestures in the video.
[0901] Server: Combines sentiment analysis results with other analysis data to understand the user's current situation and the necessary steps.
[0902] Step 4:
[0903] Server: Automatically generates business procedures based on analysis results. Adjusts the level of detail and explanation of the procedures according to the user's emotional state. For example, if the user is nervous, the procedures are presented more clearly and step-by-step.
[0904] Server: Stores the generated procedures in a database and updates them as needed. The procedures must always be up-to-date for daily operations.
[0905] Step 5:
[0906] User: If questions arise during the process or if additional information is needed, ask them at the terminal.
[0907] Step 6:
[0908] Terminal: Receives user questions via voice, performs speech recognition again to convert them to text, and sends them to the server along with sentiment data.
[0909] Step 7:
[0910] Server: Analyzes the question content and submitted sentiment data, and searches the database for relevant information. In particular, it selects information in a way that takes the user's emotions into consideration and determines the steps to be provided.
[0911] Terminal: The terminal presents information selected by the server to the user via voice or text, and provides a relaxed learning environment by responding in a way that is sensitive to emotions.
[0912] In this way, the system aims to improve the quality of work and learning by utilizing emotion recognition capabilities to provide users with information in the most optimal format.
[0913] (Example 2)
[0914] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0915] In today's work and learning environments, users are required to efficiently handle diverse data. However, systems that understand users' psychological and emotional states and provide optimal information accordingly are still insufficient, resulting in insufficient improvements in user efficiency and satisfaction. Therefore, there is a need for means to enable flexible responses that take into account users' emotional states, thereby improving both work efficiency and user experience.
[0916] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0917] In this invention, the server includes means for recognizing the user's emotional state based on audio and video data, means for automatically generating and adjusting procedures based on the analyzed audio and video data and the user's emotional state, and means for analyzing the user's questions and emotional state and providing optimal information. This makes it possible to provide information adaptively according to the user's emotions and improve work efficiency and learning efficiency.
[0918] "Voice data" refers to information that records user speech and ambient sounds, and serves as the basis for speech recognition and emotion analysis.
[0919] "Video data" refers to visual information that records the user's movements and facial expressions, and is used for motion recognition and emotion analysis.
[0920] "Text data" refers to information obtained by analyzing audio data and converting it into written text, representing the content of the user's speech in written form.
[0921] "Motion recognition" is a technology that identifies a user's actions and gestures by analyzing video data.
[0922] "Emotional state" refers to the psychological state that indicates the user's emotions, and means the psychological response that can be inferred from audio and video.
[0923] A "procedure" is a set of steps that users need to complete their tasks or learning, and it is automatically generated and provided to the user.
[0924] "Automatic generation" refers to the process by which a system generates procedures and information based on data, without manual intervention by humans.
[0925] "Optimal information" refers to information that contains necessary and appropriate content, tailored to the user's situation and emotional state.
[0926] To implement this invention, a system is required for real-time acquisition, analysis, and provision of information to the user of audio and video data. The system is configured as follows.
[0927] Data acquisition
[0928] Terminal: As users perform tasks or studies, audio and video data are collected in real time by the camera and microphone built into the terminal. Users do not need to perform any special operations in a natural work environment.
[0929] Data and emotion recognition
[0930] Server: Audio data transmitted from the terminal is converted into text data using a speech recognition algorithm (e.g., Google Speech-to-Text). Video data can be analyzed to capture user actions using a motion recognition algorithm (e.g., OpenCV), and the user's emotional state can be estimated using an emotion engine (e.g., an emotion analysis tool). This allows the server to understand the user's psychological state.
[0931] Procedure generation and information provision
[0932] Server: Based on the analysis of the user's voice, video, and emotional state, it automatically generates optimized work procedures. For example, if it determines that the user is confused, it provides detailed and specific instructions. These generated procedures are stored in a database, allowing the user to access them as needed and continue their work.
[0933] Terminal: Receives instructions from the server and provides feedback tailored to the user's current emotional state. If voice guidance is appropriate, it uses speech synthesis technology (e.g., a speech synthesis engine) to provide voice guidance to the user. Alternatively, instructions can be displayed on the screen in text format.
[0934] Specific example
[0935] For example, when a user encounters a question on their device while learning how to use new software, the system detects confusion based on the user's tone of voice and video, and provides detailed step-by-step instructions. If a positive emotional state is detected, only an overview is provided, leaving room for the user to continue learning on their own.
[0936] Example of a prompt
[0937] "Analyze the user's current audio and video data, recognize their emotions, and generate instructions. Output a step-by-step guide if the user becomes confused."
[0938] This system allows users to receive individually optimized information, enabling them to efficiently advance their work and learning.
[0939] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0940] Step 1:
[0941] Data acquisition
[0942] Terminal: The terminal acquires audio and video data from the user in real time. Inputs include audio collected by the microphone and video recorded by the camera. These data are output as digital audio signals and video frames.
[0943] Step 2:
[0944] Data transfer and reception
[0945] Terminal: Transmits acquired audio and video data to the server via an internet connection.
[0946] Server: Receives data sent from the terminal and prepares the audio data for processing with a speech recognition algorithm. The output consists of two streams: audio data and video data.
[0947] Step 3:
[0948] Voice analysis
[0949] Server: Inputs audio data into a speech recognition algorithm (e.g., speech recognition engine) and converts it into text data. This is achieved by extracting linguistic characteristics from the waveform of the audio signal and outputting them as corresponding text strings.
[0950] Step 4:
[0951] Video analysis
[0952] Server: The received video data is input into the motion recognition algorithm. The algorithm analyzes the user's movements and facial expressions from the video frames and extracts corresponding motion data. The output is descriptor data related to the physical actions performed by the user.
[0953] Step 5:
[0954] Emotion analysis
[0955] Server: Uses an emotion engine to estimate the user's emotional state based on analysis of audio waveforms, text data, and video. By passing the input data through an emotion classification model, the user's state is output as an emotion tag (e.g., joy, confusion).
[0956] Step 6:
[0957] Automatic generation and adjustment of procedures
[0958] Server: Based on the analysis results, the server starts a process to automatically generate business procedures. By inputting data about the user's actions and emotions as prompts to the generating AI model, it outputs procedures that have been adjusted according to the emotions.
[0959] Step 7:
[0960] Information provision and feedback
[0961] Terminal: Receives instructions and feedback generated from the server and presents them to the user. Feedback is output as audio or visual messages depending on the user's emotional state. This allows the user to work smoothly based on optimized procedures.
[0962] (Application Example 2)
[0963] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0964] In on-site work, providing instructions and support that ignore the emotional state of workers makes it difficult to effectively improve work efficiency and reduce psychological burden. In particular, when workers are feeling fatigued or confused, conventional systems may fail to provide appropriate support, potentially leading to decreased productivity and safety. Therefore, there is a need for a system that can analyze workers' emotions in real time and provide appropriate instructions and information tailored to their state.
[0965] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0966] In this invention, the server includes means for acquiring audio and video information, means for analyzing the user's emotional state, and means for adjusting and providing instructions based on the analyzed emotional state. This makes it possible to grasp the worker's emotions in real time and provide instructions and support appropriate to that emotional state.
[0967] "Audio information" refers to data consisting of human speech and ambient sounds, acquired using acoustic sensors such as microphones.
[0968] "Visual information" refers to visual data of objects and people acquired through cameras and other image sensors.
[0969] "Text data" refers to data in text format obtained by converting audio information into text.
[0970] "Means of recognizing motion" refers to technologies that analyze video information and detect specific patterns or actions related to the movement of people or objects.
[0971] "Means for automatically generating instructions" refers to a process that automatically generates appropriate work procedures and instructions based on the results of analysis of audio information, video information, emotional states, etc.
[0972] "Means of analyzing emotional states" refer to algorithms or engines that determine a person's emotions and psychological state from their tone of voice, facial expressions, and gestures.
[0973] "Means for adjusting and providing generated instructions" refers to a process that optimizes initially generated instructions according to the user's emotional state and presents them to the worker at the appropriate time.
[0974] This invention is a system that provides appropriate support tailored to the emotions of workers in the workplace. The system has the function of acquiring audio and video information, analyzing it, and recognizing actions and emotional states.
[0975] First, the device uses a microphone and camera to acquire audio and video information in real time. The audio information is converted into text data by a speech recognition algorithm (e.g., Google Cloud Speech-to-Text). The video information is used to identify the work being done using an action recognition algorithm.
[0976] Subsequently, the server analyzes the user's emotional state from voice tone and facial expressions obtained using an emotion engine (e.g., Microsoft Azure Emotion API). Based on the emotional state and behavioral information, the server automatically generates and adjusts instructions.
[0977] The analysis results and adjusted instructions are stored in a database and provided to the user in voice or text format at the appropriate time. For example, if a worker is confused, the system can provide specific instructions such as, "Next, use the F panel and attach it to the A base. When doing so, use the right side of the A base."
[0978] The following are examples of prompts to input into the generative AI model.
[0979] "Analyze the emotions captured from real-time audio and video data to identify at which stage of the task the user is experiencing difficulties. Then, propose work procedures or support tailored to their emotional state."
[0980] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0981] Step 1:
[0982] The terminal uses a microphone and camera to acquire audio and video information from the work environment in real time. The audio information includes the worker's speech and ambient sounds, while the video information is video captured by the camera. This data is transmitted directly to the server.
[0983] Step 2:
[0984] The server converts the acquired audio information into text data using a speech recognition algorithm (e.g., Google Cloud Speech-to-Text). The input is audio information, and the output is string data obtained by analyzing the audio. This process records the worker's speech in text.
[0985] Step 3:
[0986] The server analyzes the acquired video information using a motion recognition algorithm to identify the work content and actions. The input is video information, and the output is data related to the actions (e.g., specific actions performed by the worker). This allows for an accurate understanding of what actions the worker is performing.
[0987] Step 4:
[0988] The server analyzes the emotional state using an emotion engine (e.g., Microsoft Azure Emotion API) based on voice tone and facial expressions. Input is audio and video information, and output is emotional state data (e.g., joy, confusion, fatigue). This process identifies emotions using voice tone, speaking style, and facial expressions.
[0989] Step 5:
[0990] The server automatically generates appropriate instructions based on analyzed voice data, motion data, and emotional state data. The input is the data set obtained in the previous step, and the output is specific instructions to be presented to the user. These instructions are adjusted according to the worker's emotional state.
[0991] Step 6:
[0992] The user receives instructions from the server and performs the task. Instructions are provided in audio or text format and include specific guidance to help the user continue the task. Care is taken to ensure that the instructions are presented in a way that is easy for the user to understand.
[0993] Step 7:
[0994] The terminal continuously sends data to the server and updates it whenever processing is complete or new input is required. This enables real-time feedback that responds to changes in the worker's emotional state and work content.
[0995] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0996] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0997] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0998] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0999] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1000] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1001] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1002] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1003] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1004] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1005] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1006] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1007] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1008] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1009] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1010] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1011] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1012] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1013] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1014] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1015] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1016] The following is further disclosed regarding the embodiments described above.
[1017] (Claim 1)
[1018] Means for acquiring audio and video data,
[1019] A means for analyzing acquired audio data and converting it into text data,
[1020] A means for analyzing acquired video data and recognizing motion,
[1021] A means for automatically generating procedures based on analyzed audio and video data,
[1022] A means for storing and displaying the generated procedure,
[1023] A system that includes this.
[1024] (Claim 2)
[1025] The system according to claim 1, comprising means for updating existing procedures based on acquired data.
[1026] (Claim 3)
[1027] The system according to claim 1, comprising means for analyzing user questions and providing optimal information.
[1028] "Example 1"
[1029] (Claim 1)
[1030] A structure for acquiring audio and video information,
[1031] A structure that analyzes acquired audio information and converts it into text information,
[1032] A structure that analyzes acquired video information and recognizes movement,
[1033] A structure that automatically generates procedures based on analyzed audio and video information,
[1034] A structure for storing and presenting the generated procedure,
[1035] A structure for answering questions to resolve information gaps that users may have during their work,
[1036] A system that includes this.
[1037] (Claim 2)
[1038] The system according to claim 1, comprising a structure that updates stored procedural information based on acquired information.
[1039] (Claim 3)
[1040] The system according to claim 1, comprising a structure that analyzes user questions and provides relevant or new information.
[1041] "Application Example 1"
[1042] (Claim 1)
[1043] A device that acquires audio and video information,
[1044] The process involves analyzing the acquired audio information and converting it into text information,
[1045] The process involves analyzing the acquired video information to recognize the task,
[1046] A process that automatically generates work procedures based on analyzed audio and video information,
[1047] The process of storing and presenting the generated work procedures,
[1048] A device that records the actions of workers and provides a method to support the efficiency of work,
[1049] A system that includes this.
[1050] (Claim 2)
[1051] The system according to claim 1, including a process for updating existing work procedures based on acquired information.
[1052] (Claim 3)
[1053] The system according to claim 1, comprising processing to respond to user inquiries and provide the most appropriate information.
[1054] "Example 2 of combining an emotion engine"
[1055] (Claim 1)
[1056] Means for acquiring audio and video data,
[1057] A means for analyzing acquired audio data and converting it into text data,
[1058] A means for analyzing acquired video data and recognizing motion,
[1059] A means for recognizing the user's emotional state based on acquired audio and video data,
[1060] A means for automatically generating and adjusting procedures based on analyzed audio data, video data, and the user's emotional state,
[1061] A means for storing and displaying the generated procedure,
[1062] A system that includes this.
[1063] (Claim 2)
[1064] The system according to claim 1, comprising means for updating existing procedures based on acquired data and the emotional state of the user.
[1065] (Claim 3)
[1066] The system according to claim 1, comprising means for analyzing the user's questions and emotional state and providing optimal information.
[1067] "Application example 2 when combining with an emotional engine"
[1068] (Claim 1)
[1069] Means for acquiring audio and video information,
[1070] A means for analyzing acquired audio information and converting it into text data,
[1071] A means for analyzing acquired video information and recognizing motion,
[1072] A means for automatically generating instructions based on analyzed audio and video information,
[1073] A means for storing and displaying the generated instructions,
[1074] A means of analyzing the emotional state of users,
[1075] A means of adjusting and providing generated instructions based on the analyzed emotional state,
[1076] A system that includes this.
[1077] (Claim 2)
[1078] The system according to claim 1, comprising means for updating existing instructions based on acquired emotional information.
[1079] (Claim 3)
[1080] The system according to claim 1, comprising means for analyzing the user's questions and emotional state and providing optimal information. [Explanation of symbols]
[1081] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for acquiring audio and video data, A means for analyzing acquired audio data and converting it into text data, A means for analyzing acquired video data and recognizing motion, A means for automatically generating procedures based on analyzed audio and video data, A means for storing and displaying the generated procedure, A system that includes this.
2. The system according to claim 1, comprising means for updating existing procedures based on acquired data.
3. The system according to claim 1, comprising means for analyzing user questions and providing optimal information.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A