System

The system addresses the shortage of trainers and know-how sharing in manufacturing by analyzing video, audio, and text data to provide real-time, intuitive advice, enhancing training efficiency and reducing costs.

JP2026030588APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024133572
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

The manufacturing industry faces a shortage of trainers for machine operation training, difficulty in sharing know-how, and high labor costs, exacerbated by noisy environments that hinder verbal communication for questions and explanations.

Method used

A system that captures video and audio, analyzes this data, integrates it with text input, and provides advice based on analysis results, enabling efficient and intuitive training through a server and user terminal interface.

Benefits of technology

Enables intuitive and efficient training by analyzing multimodal data to provide real-time, appropriate advice, reducing the need for human trainers and lowering training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026030588000001_ABST
    Figure 2026030588000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for capturing video; means for capturing audio; means for accepting text input; means for analyzing the captured video; means for analyzing the captured audio; means for analyzing the input text; and means for providing advice to a user based on the analysis.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] ---

[0005] In the manufacturing industry, a shortage of trainers for machine operation training and the difficulty of sharing know-how are serious problems. Furthermore, training new personnel requires a great deal of time and money, and rising labor costs are putting pressure on corporate management. Furthermore, the noisy environment on-site often makes it difficult to ask questions or give explanations verbally. The present invention aims to solve these problems and provide a system for providing efficient and intuitive training. [Means for solving the problem]

[0006] The present invention provides a system including a means for capturing video, a means for capturing audio, a means for accepting text input, a means for analyzing the captured video, a means for analyzing the captured audio, a means for analyzing the input text, and a means for providing advice to a user based on the analysis results. The system further includes a means for transmitting the captured video and audio to a server and a means for displaying the advice transmitted from the server on a user terminal. The system also provides a means for integrating and analyzing the user's video data, audio data, and text data, enabling intuitive and efficient training.

[0007] ---

[0008] ---

[0009] A "means for capturing video" is a device or function that records video of the machine being operated by the user.

[0010] A "means for capturing audio" is a device or function that records audio information such as a user's questions or instructions.

[0011] A "means for accepting text input" is a device or function that allows a user to input text information using a keyboard or touch screen.

[0012] "Means for analyzing captured video" refers to an algorithm or device that analyzes captured video data and identifies machine parts and operating locations.

[0013] A "means for analyzing captured audio" is an algorithm or device that converts recorded audio data into text and understands its content.

[0014] The "means for analyzing input text" refers to an algorithm or device for analyzing the text information input by the user and understanding its content.

[0015] "Means for providing advice to users based on analysis results" refers to a function that generates and provides specific advice and instructions to users based on the analysis results of video, audio, and text data.

[0016] The "means for transmitting captured video and audio to a server" refers to a device or function that uploads data captured by a user terminal to a server via a network.

[0017] The "means for displaying advice sent from the server on the user terminal" is a function for receiving advice information generated by the server and displaying it on the user terminal.

[0018] The "means for integrating and analyzing video data, audio data, and text data" refers to an algorithm or device for integrating and centrally analyzing data in multiple different formats.

[0019] ---

[0020] The above are definitions of important words included in the scope of patent claims. It is important to use these definitions to create a more detailed specification. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0029] [First embodiment]

[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0042] ---

[0043] The purpose of this invention is to provide efficient training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. Specifically, it analyzes multimodal data from video, audio, and text input, and provides intuitive and appropriate advice to users.

[0044] System Configuration

[0045] The system of the present invention includes the following major components:

[0046] 1. Terminal: A mobile terminal such as a smartphone that provides an interface for users to operate.

[0047] Video capture method

[0048] Audio capture method

[0049] Text input method

[0050] Data transmission method

[0051] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[0052] Video analysis methods

[0053] Voice analysis methods

[0054] Text Analysis Methods

[0055] Advice Generation Method

[0056] Data Receiving Method

[0057] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[0058] Specific actions

[0059] Video, audio and text capture

[0060] 1. The user takes a photo of the machine operation using a device such as a smartphone.

[0061] 2. Users can ask questions by voice about points they don't understand or points they're unsure about. In some cases, they can input text in a noisy environment.

[0062] 3. The device temporarily stores the video data, audio data, and text data.

[0063] Data upload and analysis

[0064] 1. The device sends the captured data to the server.

[0065] 2. The server analyzes the received video data and detects specific actions and parts in each frame.

[0066] 3. The server converts the voice data into text and analyzes the user's question.

[0067] 4. The server analyzes the text input data and understands the question.

[0068] 5. The server integrates and analyzes the video, audio, and text data to generate optimal advice.

[0069] Advice generation and display

[0070] 1. The server generates advice and sends it to the device as data.

[0071] 2. The device notifies the user of the received advice and displays it to them. This advice is provided as text, audio, or in some cases as a guidance video.

[0072] Specific examples

[0073] Example 1: Machine part replacement

[0074] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[0075] 2. The device sends video and audio data to the server.

[0076] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[0077] 4. The server generates advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool."

[0078] 5. The device displays the received advice to the user and also plays it back aloud.

[0079] 6. The user replaces the part following the appropriate procedure.

[0080] Example 2: Checking for abnormal sounds

[0081] 1. A user records an abnormal sound made by a machine on their smartphone and asks by text, "Is this sound normal?"

[0082] 2. The device sends video and audio data to the server.

[0083] 3. The server uses voice analysis to detect abnormal sounds and analyzes the text input to understand the question.

[0084] 4. The server generates the advice, "This noise is caused by worn bearings. We recommend replacing the bearings."

[0085] 5. The device displays the received advice to the user as text and plays it back aloud.

[0086] 6. The user follows the advice, checks the cause of the abnormal sound, and takes action.

[0087] This concludes the description of the "Multimodal Factory Coach" system, which enables efficient and intuitive training in manufacturing workplaces.

[0088] The processing flow will be explained below.

[0089] ---

[0090] In case of replacing machine parts

[0091] Step 1:

[0092] The user takes a video of the machine operation using a smartphone, and uses the device's camera function to record the specific operation points.

[0093] Step 2:

[0094] The user can ask questions by voice about things they don't understand, such as, "Please tell me how to remove this bolt."

[0095] Step 3:

[0096] Temporarily saves video and audio data captured by the device. Stores data in internal storage.

[0097] Step 4:

[0098] Your device establishes an internet connection, using Wi-Fi or your mobile data network, to send the stored data to the server.

[0099] Step 5:

[0100] The device uploads video and audio data to the server, which then calls the appropriate API to package and transmit the data.

[0101] Step 6:

[0102] The server analyzes the received video data and uses an image analysis algorithm to identify the location of the bolt in each video frame.

[0103] Step 7:

[0104] The server converts the voice data into text. Using voice recognition technology, the server converts the user's speech into a string of characters.

[0105] Step 8:

[0106] The server analyzes the converted text and understands the question, using natural language processing technology to analyze the intent of the question.

[0107] Step 9:

[0108] The server integrates and analyzes the video and text data, and uses a multimodal AI model to generate optimal advice in response to the user's question.

[0109] Step 10:

[0110] Formatting server-generated advice into text and speech, e.g. "Turn this bolt clockwise to remove it. Use the special tool."

[0111] Step 11:

[0112] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[0113] Step 12:

[0114] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[0115] Step 13:

[0116] The user follows the displayed advice and continues the part replacement operation.

[0117] ---

[0118] Checking for abnormal machine noise

[0119] Step 1:

[0120] The user records abnormal sounds from the machine using their smartphone. The abnormal sounds are recorded using the device's microphone function.

[0121] Step 2:

[0122] The user texts in a question: "Is this sound normal?"

[0123] Step 3:

[0124] The device temporarily stores captured voice data and entered text. Stores data in internal storage.

[0125] Step 4:

[0126] Your device establishes an internet connection, using Wi-Fi or your mobile data network, to send the stored data to the server.

[0127] Step 5:

[0128] The device uploads the voice and text data to the server, which calls the appropriate API to package and transmit the data.

[0129] Step 6:

[0130] The server analyzes the received audio data and uses an audio analysis algorithm to detect abnormal sound patterns.

[0131] Step 7:

[0132] The server analyzes the input text data and understands the question. It uses natural language processing technology to analyze the intent of the question.

[0133] Step 8:

[0134] The server integrates and analyzes the voice and text data, and uses a multimodal AI model to generate optimal advice for the user's question.

[0135] Step 9:

[0136] Formatting server-generated advice into text and speech, e.g. "This sound is caused by worn bearings. We recommend replacing the bearings."

[0137] Step 10:

[0138] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[0139] Step 11:

[0140] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[0141] Step 12:

[0142] The user follows the displayed advice to identify the cause of the abnormal sound and take action.

[0143] ---

[0144] The above are the specific processing steps of the "Multimodal Factory Coach" system.

[0145] Example 1

[0146] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0147] To resolve the shortage of trainers and lack of know-how sharing in the manufacturing industry, provide efficient and intuitive training, and enable users to receive appropriate advice in real time. There is a demand for time-saving and effortless responses to complex machine operations and abnormality detection on manufacturing sites. Conventional methods lack the mechanisms for integrating and analyzing multimodal data such as video, audio, and text to provide appropriate advice.

[0148] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0149] In this invention, the server includes means for acquiring video, means for acquiring audio, means for accepting text, means for analyzing the acquired video, means for analyzing the acquired audio, means for analyzing input text, means for providing advice to the user based on the analysis results, and means for integrating and analyzing the acquired data, thereby making it possible to perform an integrated analysis of multimodal data and provide appropriate advice to the user in real time.

[0150] The "means for acquiring images" is a function for capturing visual information of the surroundings using a sensor device such as a camera mounted on the equipment operated by the user.

[0151] The "means for acquiring audio" is a function for collecting surrounding audio information using a microphone or audio capture device installed in the device used by the user.

[0152] "Means for accepting text" refers to a function that receives and records text information entered by a user using a keyboard or touch screen.

[0153] "Means for analyzing the captured video" refers to software algorithms or machine learning models that process the captured video data and recognize specific actions or objects.

[0154] The "means for analyzing the captured voice" refers to a voice recognition system or natural language processing technology that converts the captured voice data into text and further understands its content.

[0155] The "means for analyzing input text" is a natural language processing algorithm that processes the character data entered by the user and understands its content.

[0156] The "means for providing advice to the user based on the analysis results" is a system that integrates the analyzed video, audio, and text data, generates appropriate advice, and notifies or displays that advice to the user.

[0157] "Means for integrating and analyzing acquired data" refers to data fusion technology and analysis algorithms that process video, audio, and text data in an integrated manner and make comprehensive judgments.

[0158] "Computer system" refers collectively to hardware and software for receiving, processing, analyzing data, and generating advice.

[0159] The "user interface" refers to an interface such as a display device or audio output device for directly providing the generated advice to the user.

[0160] The "Multimodal Factory Coach" system is a system that provides efficient training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. This system analyzes multimodal data from video, audio, and text input, and provides intuitive and appropriate advice to users.

[0161] System Configuration

[0162] The system of the present invention includes the following major components:

[0163] 1. Terminal: A mobile terminal that provides an interface for users to operate. Specifically, a smartphone or tablet is used.

[0164] Means of acquiring images (camera)

[0165] A means of capturing audio (microphone)

[0166] A means of accepting text (keyboard or touchscreen)

[0167] Data transmission method (Internet data upload function)

[0168] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[0169] A method for analyzing video (video analysis algorithm)

[0170] A means of analyzing the voice (voice recognition technology, e.g., Google Cloud Speech-to-Text API)

[0171] Means of analyzing text (natural language processing technology)

[0172] A means of generating advice (generative AI model)

[0173] Data reception method (data download function via the Internet)

[0174] 3. Communication network: The infrastructure for sending and receiving data between devices and servers, usually using an internet connection.

[0175] Specific actions

[0176] Data Capture

[0177] A user uses a device such as a smartphone to record video of a machine operation. For example, a user may record a video of a machine part replacement and ask a question by voice, such as, "Please tell me how to remove this bolt." In a noisy environment, text input is also possible, and the user may input, "Is this sound normal?" The device temporarily stores this video, audio, and text data.

[0178] Data transmission

[0179] The device sends the captured data to a server, using an internet connection and uploading the data securely.

[0180] Data reception and analysis

[0181] The server receives and analyzes the data sent from the terminal. It uses video analysis means to detect specific actions and components in each frame. It uses audio analysis means to convert the audio data into text and analyze the question. It also uses input text analysis means to understand the question.

[0182] Advice Generation

[0183] The server integrates the analyzed data and generates optimal advice. For example, it generates advice on specific steps and tool selection, such as "Turn this bolt clockwise to remove it. Use the appropriate tool." Using a generative AI model, it is possible to generate the optimal advice that the user desires.

[0184] Providing advice

[0185] The server sends the generated advice to the terminal, which receives it and notifies and displays it to the user. The advice is provided as text, audio, or in some cases, guidance video. The user can follow the instructions and perform the appropriate task.

[0186] Specific examples

[0187] Example 1: Machine part replacement

[0188] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[0189] 2. The device sends video and audio data to the server.

[0190] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[0191] 4. The server generates advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool."

[0192] 5. The device displays the received advice to the user and also plays it back aloud.

[0193] 6. The user replaces the part following the appropriate procedure.

[0194] Example 2: Checking for abnormal sounds

[0195] 1. A user records an abnormal sound made by a machine on their smartphone and asks by text, "Is this sound normal?"

[0196] 2. The device sends the voice data and text data to the server.

[0197] 3. The server uses voice analysis to detect abnormal sounds and analyzes the text input to understand the question.

[0198] 4. The server generates the advice, "This noise is caused by worn bearings. We recommend replacing the bearings."

[0199] 5. The device displays the received advice to the user as text and plays it back aloud.

[0200] 6. The user follows the advice, checks the cause of the abnormal sound, and takes action.

[0201] In this way, the system enables real-time support for users, supporting rapid response and efficient work on-site.

[0202] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0203] Step 1: Data Capture

[0204] A user uses a smartphone to record video of a machine being operated and ask questions by voice. The input is video data and audio data, and the output is data temporarily stored on the device. For example, if a user records a machine part replacement and asks, "Please tell me how to remove this bolt," this data is stored on the device.

[0205] Step 2: Send data

[0206] The device sends the stored video and audio data to the server. The input is the data stored on the device, and the output is the data uploaded to the server. The data is sent via the Internet and received by the server.

[0207] Step 3: Receiving data

[0208] The server receives the data sent from the device. The input is the video and audio data sent from the device, and the output is the data saved on the server. After receiving, the data is stored in an appropriate folder for analysis.

[0209] Step 4: Video analysis

[0210] The server analyzes the received video data. The input is the video data, and the output is analysis data that identifies specific actions and objects. For example, a video analysis algorithm is used to analyze the position of bolts and the user's hand movements in each frame.

[0211] Step 5: Audio analysis

[0212] The server converts the received voice data into text and analyzes the question. The input is voice data, and the output is text data and the analysis results. Using a voice recognition system, for example, the voice data is converted into text, such as "Please tell me how to remove this bolt," and the intent is understood.

[0213] Step 6: Text Analysis

[0214] The server analyzes the text data and understands the question. The input is text data, and the output is an analysis result that indicates the intent of the question. Natural language processing technology is used to determine what information the input text is seeking.

[0215] Step 7: Data Integration

[0216] The server integrates the results of video analysis, audio analysis, and text analysis. The input is the results of each analysis, and the output is the integrated analysis data. For example, the server can combine and analyze the location information of a bolt and the content of a voice question to determine how to remove the bolt.

[0217] Step 8: Advice Generation

[0218] The server generates advice based on the integrated analysis data. The input is the integrated analysis data, and the output is specific advice. Using a generative AI model, advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool" is generated.

[0219] Step 9: Send Advice

[0220] The server sends the generated advice to the terminal. The input is the advice content, and the output is the data sent to the terminal. The advice is sent to the terminal via the Internet.

[0221] Step 10: Providing advice

[0222] The device notifies and displays the received advice to the user. The input is advice data sent from the server, and the output is text, audio, and guidance video displayed to the user. For example, it may play audio such as "Turn this bolt clockwise to remove it," encouraging the user to take specific action.

[0223] The above is the specific processing flow of the "Multimodal Factory Coach" system program.

[0224] (Application example 1)

[0225] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0226] In factory operations, there is a need to provide appropriate training and advice in real time when operating or maintaining robots. However, conventional systems lack human trainers, making it difficult to efficiently share know-how. This results in problems that cannot be dealt with quickly, leading to a decline in production efficiency. In particular, there is a need for technology that can provide effective support when an immediate response is required, such as when complex operations or abnormal sounds are generated.

[0227] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0228] In this invention, the server includes a means for analyzing video, a means for analyzing audio, and a means for analyzing text. This allows the server to analyze the video, audio, and text data captured by the user and provide advice based on the analysis results in real time using a generative AI model. Specifically, the server identifies the work content from the video data, understands the user's questions and problems from the audio and text data, and generates appropriate advice. By performing data analysis on the server side using prompt statements, the server can provide quick and accurate support to on-site workers and quickly respond to problems when they occur.

[0229] "Means for capturing images" refers to the function of acquiring visual information as digital data using a camera or sensor.

[0230] "Means for capturing audio" refers to a function for obtaining audio information as digital data using a microphone.

[0231] "Means for accepting text input" refers to the ability to input text information using a keyboard or touch screen.

[0232] "Means for analyzing captured video" refers to the function of processing and analyzing acquired video data using algorithms and AI.

[0233] The "means for analyzing captured audio" is a function for converting acquired audio data into text and analyzing the content.

[0234] "Means for analyzing input text" refers to a function that understands the character information input by the user and analyzes the context and meaning.

[0235] "Means for providing advice to the user based on the analysis results" is a function that provides the user with appropriate instructions and advice based on the analysis results of video, audio, and text.

[0236] "Means for notifying the user of the analysis results in real time" is a function that enables a quick response by immediately transmitting the analysis results to the user.

[0237] A "server" is a computer system that performs analysis and advice generation, and is often located in a cloud environment.

[0238] A "generative AI model" is an artificial intelligence model that generatively provides specific advice and answers in response to a user's questions.

[0239] A "prompt statement" is an instruction statement that gives AI instructions for analysis or generation.

[0240] This invention provides a system that provides an application called "Smart Robo Coach" that is installed on robots in factories. This system captures video, audio, and text, analyzes the data in real time, and provides appropriate advice to users.

[0241] The system has the following main components:

[0242] 1. Terminal

[0243] Camera: Built into the robot to acquire visual information.

[0244] Microphone: Captures audio information.

[0245] Text input interface: A user inputs text information by operating a keyboard or touch screen.

[0246] 2. Server

[0247] Video analysis method: The acquired video data is processed and analyzed using algorithms and AI.

[0248] Voice analysis means: Converts acquired voice data into text and analyzes the content.

[0249] Text analysis means: Understands the text information entered by the user and analyzes the context and meaning.

[0250] Advice generation method: Uses a generative AI model to generate appropriate advice based on the analysis results.

[0251] Data transmission and reception means: Data is transmitted from the terminal to the server, and the server notifies the terminal of the analysis results.

[0252] 3. Communication Network

[0253] It is the infrastructure for sending and receiving data between terminals and servers.

[0254] The specific operation is as follows.

[0255] Video, audio and text capture

[0256] 1. The user uses the robot's built-in camera to capture video of the work being done.

[0257] 2. Users can ask questions by voice about any points they do not understand or are unsure about. In noisy environments, a text input interface can also be used.

[0258] 3. The device temporarily stores the captured video data, audio data, and text data.

[0259] Data upload and analysis

[0260] 1. The device sends the captured data to the server in real time.

[0261] 2. The server analyzes the received video data and detects specific actions and parts in real time.

[0262] 3. The server converts the voice data into text and analyzes the user's question.

[0263] 4. The server analyzes the text input data and understands the question.

[0264] 5. The server integrates and analyzes the video, audio, and text data, and uses a generative AI model to generate optimal advice.

[0265] Advice generation and display

[0266] 1. The server generates advice and sends it to the device as data.

[0267] 2. The device notifies the user of the received advice and displays it to them. This advice is provided as text, audio, or in some cases as a guidance video.

[0268] The hardware used is a camera, microphone, and communication network built into the robot, and the software is video analysis, audio analysis, and generative AI models (such as GPT-3) implemented in Python and other languages.

[0269] As a specific example, consider a robot trying to remove a machine part. When an operator asks, "Please tell me how to remove this bolt," the robot's built-in camera captures an image of the bolt and the question is recorded by a microphone. This data is sent to a server, which analyzes it and generates advice such as, "Turn the bolt clockwise to remove it."

[0270] An example of a prompt for the generative AI model is as follows:

[0271] "Your task is to analyze video and audio data captured by robots in a factory and provide the best advice for a given task. For example, how to remove a bolt or the cause of an unusual noise."

[0272] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0273] Step 1:

[0274] The user captures video of the robot working using the built-in camera and asks questions by voice using the microphone. The input is video data and audio data. The output is the captured video file and audio file. Specifically, the camera captures video at a specified frame rate, and the microphone records audio.

[0275] Step 2:

[0276] The device temporarily stores captured video and audio data. The input is the captured video and audio files. The output is the video and audio files stored in the device's storage. Specifically, the video and audio files are stored in a specific directory.

[0277] Step 3:

[0278] The device sends stored video data, audio data, and text input data to the server in real time. The input is the video file, audio file, and text data. The output is the data sent to the server. Specifically, the data is uploaded to the server using the HTTPS protocol.

[0279] Step 4:

[0280] The server analyzes the received video data and detects specific actions and parts in real time. The input is a video file. The output is the analyzed video data, such as the location information of specific parts. Specifically, it uses OpenCV and machine learning models to analyze the video data frame by frame.

[0281] Step 5:

[0282] The server converts the received voice data into text and analyzes the question content. The input is an audio file. The output is text data and the analyzed question content. Specifically, the voice data is transcribed using the SpeechRecognition library and the content is analyzed using natural language processing (NLP) algorithms.

[0283] Step 6:

[0284] The server analyzes the input text data and understands the question. The input is text data. The output is the analyzed question. Specifically, it uses NLP algorithms to analyze the context and meaning of the text.

[0285] Step 7:

[0286] The server integrates and analyzes video data, audio data, and text data, and generates optimal advice using a generative AI model. The input is video data, audio data, and text data. The output is text or audio data as advice. Specifically, the server integrates the results of video analysis and text analysis, and generates advice using a generative AI model (e.g., GPT-3).

[0287] Step 8:

[0288] The server sends the generated advice to the terminal as data. The input is text or audio data as advice. The output is the advice data sent to the terminal. Specifically, the data is sent to the terminal using the HTTPS protocol.

[0289] Step 9:

[0290] The terminal notifies and displays the received advice to the user. This advice is provided as text or audio, or in some cases as a guidance video. The input is text or audio data as advice. The output is the advice notified to the user. Specifically, the text is displayed on the display and audio is played from the speaker.

[0291] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0292] ---

[0293] This invention is a system that provides efficient and intuitive training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. Furthermore, it aims to improve the user experience and training effectiveness by adding an emotion engine that recognizes the user's emotions.

[0294] System Configuration

[0295] The system of the present invention includes the following major components:

[0296] 1. Terminal: A mobile terminal such as a smartphone that provides an interface for users to operate.

[0297] Video capture method

[0298] Audio capture method

[0299] Text input method

[0300] Data transmission method

[0301] Emotion Recognition Module

[0302] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[0303] Video analysis methods

[0304] Voice analysis methods

[0305] Text Analysis Methods

[0306] Advice Generation Method

[0307] Data Receiving Method

[0308] Emotion Engine

[0309] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[0310] Specific actions

[0311] Video, audio, text and emotional capture

[0312] 1. The user takes a video of the machine operation using a device such as a smartphone.

[0313] 2. Users can ask questions by voice if they have any questions or concerns. In noisy environments, they can input text.

[0314] 3. The device temporarily stores the video data, audio data, and text data.

[0315] 4. The emotion recognition module analyzes the user's emotions in real time from the video and audio data.

[0316] Data upload and analysis

[0317] 1. The device sends the captured data to the server.

[0318] 2. The server analyzes the received video data and detects specific actions and parts in each frame.

[0319] 3. The server converts the voice data into text and analyzes the user's question.

[0320] 4. The server analyzes the text input data and understands the question.

[0321] 5. The emotion engine analyzes the received emotion data and detects the user's stress level and emotional state.

[0322] 6. The server integrates and analyzes the video, audio, text, and emotional data to generate optimal advice.

[0323] Advice generation and display

[0324] 1. The server formats the generated advice into text and speech. An emotion engine adjusts the advice content to the user's emotional state.

[0325] For example, advice like "Turn this bolt clockwise to remove it. Use the appropriate tool" becomes "Turn this bolt clockwise slowly. If you have trouble, use the appropriate tool" if the user is feeling stressed.

[0326] 2. The server sends the generated advice data to the terminal, securing an appropriate network path to transfer the data.

[0327] 3. The device displays the advice data received from the server by displaying a text message on the screen and playing the advice via audio output.

[0328] Specific examples

[0329] Example 1: Machine part replacement

[0330] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[0331] 2. The device sends video and audio data to the server.

[0332] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[0333] 4. The emotion engine analyzes the user's emotional state from video and audio.

[0334] 5. The server combines video, audio, text, and emotion data to generate advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[0335] 6. The device displays the received advice to the user and also plays it back aloud.

[0336] 7. The user replaces the part according to the advice.

[0337] Example 2: Checking for abnormal sounds

[0338] 1. A user records an abnormal machine sound on their smartphone and asks via text, "Is this sound normal?"

[0339] 2. The device sends the voice and text data to the server.

[0340] 3. The server uses voice analysis to detect abnormal sound patterns and uses text analysis to understand the content of the question.

[0341] 4. The emotion engine analyzes the user's emotional state.

[0342] 5. The server generates advice on the cause of the abnormal sound and how to deal with it, for example, by telling a nervous user, "This sound is caused by worn bearings. It would be a good idea to replace the bearings."

[0343] 6. The device displays the received advice to the user and plays it aloud.

[0344] 7. The user follows the advice to identify the cause of the abnormal sound and take action.

[0345] This concludes the description of the "Multimodal Factory Coach" system, which enables efficient and intuitive training in manufacturing workplaces and provides personalized support based on the user's emotions.

[0346] The processing flow will be explained below.

[0347] ---

[0348] In case of replacing machine parts

[0349] Step 1:

[0350] The user takes a video of the machine operation using a smartphone, and uses the device's camera function to record the specific operation points.

[0351] Step 2:

[0352] The user can ask questions by voice about things they don't understand, such as, "Please tell me how to remove this bolt."

[0353] Step 3:

[0354] Temporarily saves video and audio data captured by the device. Stores data in internal storage.

[0355] Step 4:

[0356] The device's emotion recognition module analyzes the audio and video data to determine the user's emotional state, for example, determining whether the user is nervous based on the tone of their voice or facial expression.

[0357] Step 5:

[0358] The device establishes an internet connection using Wi-Fi or a mobile data network to send video data, audio data, text data, and emotion data to the server.

[0359] Step 6:

[0360] The device uploads all data to the server, which calls the appropriate API to package and transmit the data.

[0361] Step 7:

[0362] The server analyzes the received video data and uses an image analysis algorithm to identify the location of the bolt in each video frame.

[0363] Step 8:

[0364] The server converts the voice data into text. Using voice recognition technology, the server converts the user's speech into a string of characters.

[0365] Step 9:

[0366] The server analyzes the converted text and understands the question, using natural language processing technology to analyze the intent of the question.

[0367] Step 10:

[0368] The server's emotion engine analyzes the emotion data to understand the user's stress level and emotional state.

[0369] Step 11:

[0370] The server integrates and analyzes video data, audio data, text data, and emotional data, and uses a multimodal AI model to generate optimal advice in response to the user's question.

[0371] Step 12:

[0372] The server adjusts the advice based on the user's emotional data. For example, if the user is nervous, the advice will be adjusted to something like, "Turn this bolt clockwise slowly. If you have trouble, use the appropriate tool."

[0373] Step 13:

[0374] Formatting server-generated advice into text and speech, e.g. "Turn this bolt clockwise to remove it. Use the special tool."

[0375] Step 14:

[0376] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[0377] Step 15:

[0378] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[0379] Step 16:

[0380] The user follows the displayed advice and continues the part replacement operation.

[0381] ---

[0382] Checking for abnormal machine noise

[0383] Step 1:

[0384] The user records abnormal sounds from the machine using their smartphone. The abnormal sounds are recorded using the device's microphone function.

[0385] Step 2:

[0386] The user texts in a question: "Is this sound normal?"

[0387] Step 3:

[0388] The device temporarily stores captured voice data and entered text. Stores data in internal storage.

[0389] Step 4:

[0390] The device's emotion recognition module analyzes the video and audio data to determine the user's emotional state.

[0391] Step 5:

[0392] The device establishes an internet connection using Wi-Fi or a mobile data network to send voice, text, and emotion data to the server.

[0393] Step 6:

[0394] The device uploads all data to the server, which calls the appropriate API to package and transmit the data.

[0395] Step 7:

[0396] The server analyzes the received audio data and uses an audio analysis algorithm to detect abnormal sound patterns.

[0397] Step 8:

[0398] The server analyzes the input text data and understands the question. It uses natural language processing technology to analyze the intent of the question.

[0399] Step 9:

[0400] The server's emotion engine analyzes the emotion data to understand the user's stress level and emotional state.

[0401] Step 10:

[0402] The server integrates and analyzes voice data, text data, and emotion data, and uses a multimodal AI model to generate optimal advice for the user's question.

[0403] Step 11:

[0404] The server adjusts the advice based on the user's emotional data. For example, if the user is nervous, the advice will be adjusted to say, "This noise is caused by worn bearings. It would be a good idea to replace the bearings."

[0405] Step 12:

[0406] Formatting server-generated advice into text and speech, e.g. "This sound is caused by worn bearings. We recommend replacing the bearings."

[0407] Step 13:

[0408] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[0409] Step 14:

[0410] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[0411] Step 15:

[0412] The user follows the displayed advice to identify the cause of the abnormal sound and take action.

[0413] ---

[0414] The above are the specific processing steps when combining the "Multimodal Factory Coach" system with an emotion engine.

[0415] Example 2

[0416] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0417] In modern manufacturing, a lack of trainers and a lack of sharing of know-how are serious problems. This makes it difficult for new or inexperienced employees to receive training quickly and effectively, resulting in reduced work efficiency. Furthermore, simple advice that does not take users' emotions into consideration makes it difficult to address the stress and tension they feel, preventing optimal performance. A system that can solve these problems and provide efficient and intuitive training is needed.

[0418] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0419] In this invention, the server

[0420] a means for capturing video;

[0421] a means for capturing audio;

[0422] means for accepting text input;

[0423] a means for analyzing emotions;

[0424] means for analyzing the captured video;

[0425] means for analyzing the captured audio;

[0426] means for parsing input text;

[0427] means for analyzing user emotion data;

[0428] and means for providing advice to the user based on the analysis results.

[0429] This makes it possible to provide effective and intuitive training that takes into account the stress and tension felt by the user. By analyzing the user's emotional state in real time and generating and providing appropriate advice based on that state, the user's understanding and work efficiency are improved.

[0430] "Means for capturing video" refers to devices or functions that capture the equipment operated by the user and the work environment in real time, and record and save the video data.

[0431] "Audio capturing means" refers to a device or function that collects the user's voice using a microphone or the like and records it as audio data.

[0432] "Means for accepting text input" refers to a function that allows a user to input text data using an interface such as a keyboard or touch screen.

[0433] "Means for analyzing emotions" refers to software or algorithms that use video and audio data of users to recognize and evaluate their emotional state, such as stress and tension, in real time.

[0434] "Means for analyzing captured video" refers to a function that analyzes received video data and detects and identifies specific actions, positions, parts, etc.

[0435] The "means for analyzing captured audio" is a function that converts collected audio data into text data and further analyzes the content of that text data.

[0436] "Means for analyzing input text" refers to software or algorithms that understand the content of the text data entered by the user and extract the necessary information.

[0437] The "means for analyzing the user's emotional data" is a function that analyzes the user's emotional state from the user's video data and audio data, and evaluates the stress level and tension.

[0438] The "means for providing advice to the user based on the analysis results" is a function for generating and providing optimal instructions and advice to the user based on the analysis results of video, audio, text, and emotional data.

[0439] This invention is a system for providing efficient and intuitive training, resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. The system aims to improve the effectiveness of training by recognizing the user's emotions and providing appropriate advice.

[0440] System Configuration

[0441] The system includes the following major components:

[0442] 1. Terminal: A mobile terminal operated by the user, equipped with the following functions:

[0443] Video capture method

[0444] Audio capture method

[0445] Text input method

[0446] Data transmission method

[0447] Emotion Recognition Module

[0448] 2. Server: A computer system for analysis and advice generation located in a cloud environment.

[0449] Video analysis methods

[0450] Voice analysis methods

[0451] Text Analysis Methods

[0452] Advice Generation Method

[0453] Data Receiving Method

[0454] Emotion Engine

[0455] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[0456] Specific actions

[0457] Video, audio, text and emotional capture

[0458] 1. The user uses a device such as a smartphone to record video of the machine operation. For example, while recording the operation, the user can ask a question by voice, such as "Please tell me how to remove this bolt."

[0459] 2. If the user is in a noisy environment, he or she inputs text. For example, the user inputs text such as "Is this sound normal?"

[0460] 3. The device temporarily stores the video data, audio data, and text data.

[0461] 4. The emotion recognition module analyzes the user's emotions in real time from their video and audio data, for example, determining whether they are nervous.

[0462] Sending data

[0463] 1. The device sends all captured data to the server.

[0464] 2. Use a high-speed and stable communication network for data transmission.

[0465] Data analysis on the server

[0466] 1. The server analyzes the video data received by the server using video analysis means to detect specific actions and parts in each frame. For example, it identifies the position of a bolt.

[0467] 2. The server converts the voice data into text and uses speech analysis to understand the question. For example, it analyzes the question, "Please tell me how to remove this bolt."

[0468] 3. The server analyzes the text input data and understands the question.

[0469] 4. The emotion engine analyzes the received emotion data and detects the user's stress level and emotional state. For example, it recognizes that the user is nervous.

[0470] Generating Advice

[0471] 1. The server integrates and analyzes video, audio, text, and emotion data to generate optimal advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[0472] 2. The emotion engine adjusts advice according to the user's emotional state.

[0473] Providing advice and feedback

[0474] 1. The server sends the generated advice to the device.

[0475] 2. The device displays the received advice data to the user and also plays it aloud. For example, the device might display the message "Turn this bolt clockwise slowly. Use a tool if necessary" on the screen and simultaneously play it aloud.

[0476] 3. The user follows the advice and performs the task. During this time, the device monitors the user's emotional state and sends feedback to the server as needed.

[0477] Specific examples

[0478] Example 1: Replacing machine parts

[0479] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[0480] 2. The device sends video and audio data to the server.

[0481] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[0482] 4. The emotion engine analyzes the user's emotional state from video and audio.

[0483] 5. The server combines video, audio, text, and emotion data to generate advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[0484] 6. The device displays the received advice to the user and also plays it back aloud.

[0485] 7. The user replaces the part according to the advice.

[0486] Example 2: Checking for abnormal sounds

[0487] 1. A user records an abnormal machine sound on their smartphone and asks via text, "Is this sound normal?"

[0488] 2. The device sends the voice and text data to the server.

[0489] 3. The server uses voice analysis to detect abnormal sound patterns and uses text analysis to understand the content of the question.

[0490] 4. The emotion engine analyzes the user's emotional state.

[0491] 5. The server generates advice on the cause of the abnormal sound and how to deal with it, for example, by telling a nervous user, "This sound is caused by worn bearings. It would be a good idea to replace the bearings."

[0492] 6. The device displays the received advice to the user and plays it aloud.

[0493] 7. The user follows the advice to identify the cause of the abnormal sound and take action.

[0494] Prompt Sentence Examples

[0495] "Please tell me how to replace the machine parts. I'm a little nervous."

[0496] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0497] Step 1:

[0498] The user captures video using a device. For example, the user may record video of a machine operation or part replacement using a smartphone, and then ask a question by voice, such as "Please tell me how to remove this bolt." Specifically, the user's operation of the equipment is recorded, and the audio is collected using the device's microphone.

[0499] Input: Video and audio data captured by a smartphone.

[0500] Output: The captured video and audio data is saved on the device.

[0501] Step 2:

[0502] When a user is in a noisy environment, he or she uses the device to input text. For example, he or she inputs text such as "Is this sound normal?". The input is performed using the device's keyboard or touch screen.

[0503] Input: Text data that a user types into a terminal.

[0504] Output: The entered text data is saved to the terminal.

[0505] Step 3:

[0506] The device temporarily stores video, audio, and text data and prepares it for later transmission to the server. The data is temporarily stored in the device's memory.

[0507] Input: Captured video, audio, and text data.

[0508] Output: Video, audio, and text data temporarily stored in the device's memory.

[0509] Step 4:

[0510] The emotion recognition module analyzes the user's emotions in real time from their video and audio data. For example, it analyzes the user's facial expressions and tone of voice to detect tension or stress. This emotion analysis is performed in real time and is executed on the device.

[0511] Input: Captured video and audio data.

[0512] Output: Parsed user emotion data.

[0513] Step 5:

[0514] All data captured by the device (video data, audio data, text data, and emotion data) is sent to the server. The data is transferred using a high-speed, stable communication network.

[0515] Input: Temporarily stored video, audio, and text data, as well as analyzed emotion data.

[0516] Output: Video, audio, text data and emotion data sent to the server.

[0517] Step 6:

[0518] The server analyzes the received video data. Using video analysis means, it detects specific actions and parts in each frame. For example, analysis is performed to identify the location and operation method of a bolt in the video.

[0519] Input: Video data sent to the server.

[0520] Output: Analyzed video data.

[0521] Step 7:

[0522] The server converts the received voice data into text and uses voice analysis means to understand the content of the user's question. For example, it analyzes voice data such as "Please tell me how to remove this bolt."

[0523] Input: The audio data sent to the server.

[0524] Output: The audio data converted to text and the parsed questions.

[0525] Step 8:

[0526] The server uses text analysis means to understand the content of the input text data, for example, analyzing the content of a question such as "Is this sound normal?"

[0527] Input: The text data sent to the server.

[0528] Output: Parsed text data and question content.

[0529] Step 9:

[0530] The emotion engine analyzes the received emotion data to detect the user's stress level and emotional state, for example, assessing whether the user is nervous.

[0531] Input: Emotion data sent to the server.

[0532] Output: The analyzed emotional state of the user.

[0533] Step 10:

[0534] The server combines video, audio, text, and emotion data to generate optimal advice. For example, if the user is nervous, the server generates advice such as, "Turn this bolt clockwise slowly. Use a tool if necessary."

[0535] Input: Analyzed video data, audio data, text data, and emotion data.

[0536] Output: The generated advice data.

[0537] Step 11:

[0538] The server sends the generated advice data to the terminal, securing an appropriate network path for rapid transfer.

[0539] Input: Generated advice data.

[0540] Output: Advice data sent to the terminal.

[0541] Step 12:

[0542] The device displays the received advice data to the user. It displays a text message on the screen and plays the advice aloud. For example, the device might display "Turn this bolt slowly clockwise. Use a tool if necessary" and play the advice aloud.

[0543] Input: Advice data sent by the server.

[0544] Output: The advice that was displayed and played back to the user.

[0545] The above is a flow of the specific processing steps of the system. It includes details of the data processing and calculations performed at each step, which allows for effective and intuitive training for users.

[0546] (Application example 2)

[0547] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0548] In the manufacturing industry, there are problems with efficient and intuitive training due to a shortage of trainers and insufficient sharing of know-how. Furthermore, training and advice that do not take into account the user's emotional state can reduce its effectiveness in the field.

[0549] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video, means for capturing audio, means for accepting text input, means for analyzing the user's emotional state, means for providing advice to the user based on the analysis results and the user's emotional state, means for transmitting the captured video and audio to the server, means for displaying the advice transmitted from the server on the user terminal, means for analyzing prompt sentences using a generative AI model and generating advice corresponding to the emotion, and means for integrating and analyzing the user's video data, audio data, and text data, inputting the prompt sentences into the generative AI model, and generating advice based on the analysis results. This enables efficient and intuitive training and advice provision while taking the user's emotional state into consideration.

[0550] A "video capture means" is a device or method for obtaining images or video of a user or the work environment.

[0551] An "audio capturing means" is a device or method for recording the sounds of a user or the work environment.

[0552] The "means for accepting text input" refers to a device or method that allows a user to input characters using a keyboard, an on-screen keyboard, or the like.

[0553] A "means for analyzing a user's emotional state" is a device or method for inferring a user's emotional state using audio, video, and other data.

[0554] The "means for providing advice to a user based on the analysis results and the emotional state of the user" is a device or method for combining the results of data analysis with the emotional state to provide the user with optimal advice.

[0555] The "means for transmitting captured video and audio to a server" refers to a device or method for transferring captured video and audio data to a server via a network.

[0556] The "means for displaying advice sent from the server on the user terminal" refers to a device or method for displaying instructions or advice received from the server in a form that can be confirmed by the user.

[0557] "Means for analyzing prompt sentences using a generative AI model and generating advice based on emotions" refers to a device or method that utilizes AI technology to analyze text data entered by a user and generate appropriate advice based on that data and the user's emotional state.

[0558] "Means for integrating and analyzing a user's video data, audio data, and text data, inputting prompt sentences into a generative AI model, and generating advice based on the analysis results" refers to a device or method for integrating and analyzing multiple data sources, analyzing prompt sentences using AI based on the results, and providing appropriate advice to the user.

[0559] The present invention is a system for a factory robot that provides specific advice while taking into account the emotional state of the user. The system includes the following main components:

[0560] System Configuration

[0561] 1. Terminal: This applies to robots that work in factories. Robots use cameras and microphones to capture video and audio of their work.

[0562] Video capture method

[0563] Audio capture method

[0564] Text input method

[0565] Emotion Recognition Module

[0566] Data transmission method

[0567] 2. Server: A system deployed in a cloud environment that performs analysis and generates advice. The server includes the following means:

[0568] Video analysis methods

[0569] Voice analysis methods

[0570] Text Analysis Methods

[0571] Emotion Recognition Module

[0572] Advice Generation Method

[0573] Data Receiving Method

[0574] Generative AI Models

[0575] 3. Communication network: Infrastructure for sending and receiving data between terminals and servers.

[0576] Operation overview

[0577] Data Capture and Transmission

[0578] As the user performs a task, the device (robot) captures video and audio using a camera and microphone. The device also includes a means for the user to input text using a keyboard or voice input. This data is sent to a server in real time.

[0579] Data analysis and emotion recognition

[0580] The server integrates and analyzes the received video, audio, and text data. The emotion recognition module analyzes the user's emotional state, and the generative AI model analyzes the prompt sentences to generate appropriate advice.

[0581] Advice generation and delivery

[0582] The generated advice is sent from the server to the terminal in text and audio format, and the terminal displays the received advice to the user and plays it back aloud.

[0583] Hardware and software used

[0584] Hardware: Camera, microphone, user input devices (keyboard, touch screen, etc.), robot body

[0585] Software: Video analysis software (OpenCV, etc.), audio analysis software (SpeechRecognition, etc.), emotion recognition software (EmotionRecognizer, etc.), generative AI models (cloud-based AI services, etc.)

[0586] Specific examples

[0587] Machine part replacement

[0588] If a user is unsure how to remove a bolt, they can enter "Please tell me how to remove this bolt" in text format, and the robot will send video and audio to the server. The server will analyze the data and use an emotion recognition module to confirm that the user is nervous. The advice generator will then generate advice such as "Turn this bolt slowly clockwise. Use a tool if necessary," which will then be displayed and played aloud on the device.

[0589] Check for abnormal sounds

[0590] If a user wants to check for abnormal machine sounds, they record the sound on their smartphone and then type "Is this sound normal?" into the robot while playing it back. The robot sends this to the server, which then analyzes the sound pattern using a voice analysis module and detects tension using an emotion recognition module. The server then generates advice such as "This sound is caused by worn bearings. It would be a good idea to replace the bearings," which the robot displays and plays back in audio.

[0591] Prompt Sentence Examples

[0592] Can you tell me how to remove this bolt?

[0593] "Is this sound normal?"

[0594] Through these specific operations, user support is realized through cooperation between the terminal and the server.

[0595] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0596] Step 1:

[0597] The device uses a camera and microphone to capture video and audio of the user working. As input, video data is obtained from the camera and audio data is obtained from the microphone. Specifically, the device records the user's work scene in real time and temporarily stores the data.

[0598] Step 2:

[0599] The user enters a question requesting clarification through text input. As input, the text data entered by the user via a keyboard or touch screen is acquired. Specifically, the device provides an interface for accepting the user's input and saves the entered text.

[0600] Step 3:

[0601] The device sends the captured video, audio, and text data to the server. The device uses the saved video, audio, and text data as input and sends this data to the server as output. Specifically, the device uploads the data to the cloud server using the network.

[0602] Step 4:

[0603] The server analyzes the video data and identifies the user's work. It uses the video data sent from the device as input and obtains the analysis results as output. Specifically, the server uses video analysis software to recognize specific actions and parts in each frame and compares them with a database.

[0604] Step 5:

[0605] The server analyzes the voice data and understands the user's question. It uses the voice data sent from the device as input and outputs the converted text string and the analysis results. Specifically, it uses voice recognition software to convert the voice into text and analyzes the content of the question.

[0606] Step 6:

[0607] The server analyzes the user's emotional state. It uses video and audio data as input and obtains the emotion analysis results as output. Specifically, it uses an emotion recognition module to estimate the user's emotional state from facial expressions, tone of voice, etc.

[0608] Step 7:

[0609] The server integrates the analysis results and the emotional state and generates advice using a generative AI model. The analysis results, emotion analysis results, and prompt sentences are used as input, and advice for the user is obtained as output. Specifically, the generative AI model analyzes the prompt sentence and generates appropriate advice based on the data.

[0610] Step 8:

[0611] The server sends the generated advice to the terminal. The generated advice data is used as input to obtain data to be sent to the terminal as output. In concrete terms, the server uploads the advice data to the terminal via the network.

[0612] Step 9:

[0613] The device displays the advice received from the server to the user and plays it back aloud. It uses the advice data sent from the server as input and provides visual and auditory feedback to the user as output. Specifically, the device displays a text message on the screen and plays back an audio message from the speaker.

[0614] Prompt Sentence Examples

[0615] Can you tell me how to remove this bolt?

[0616] "Is this sound normal?"

[0617] The above are the specific processing steps of this system.

[0618] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0619] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0620] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0621] [Second embodiment]

[0622] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0623] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0624] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0625] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0626] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0627] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0628] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0629] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0630] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0631] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0632] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0633] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0634] ---

[0635] The purpose of this invention is to provide efficient training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. Specifically, it analyzes multimodal data from video, audio, and text input, and provides intuitive and appropriate advice to users.

[0636] System Configuration

[0637] The system of the present invention includes the following major components:

[0638] 1. Terminal: A mobile terminal such as a smartphone that provides an interface for users to operate.

[0639] Video capture method

[0640] Audio capture method

[0641] Text input method

[0642] Data transmission method

[0643] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[0644] Video analysis methods

[0645] Voice analysis methods

[0646] Text Analysis Methods

[0647] Advice Generation Method

[0648] Data Receiving Method

[0649] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[0650] Specific actions

[0651] Video, audio and text capture

[0652] 1. The user takes a photo of the machine operation using a device such as a smartphone.

[0653] 2. Users can ask questions by voice about points they don't understand or points they're unsure about. In some cases, they can input text in a noisy environment.

[0654] 3. The device temporarily stores the video data, audio data, and text data.

[0655] Data upload and analysis

[0656] 1. The device sends the captured data to the server.

[0657] 2. The server analyzes the received video data and detects specific actions and parts in each frame.

[0658] 3. The server converts the voice data into text and analyzes the user's question.

[0659] 4. The server analyzes the text input data and understands the question.

[0660] 5. The server integrates and analyzes the video, audio, and text data to generate optimal advice.

[0661] Advice generation and display

[0662] 1. The server generates advice and sends it to the device as data.

[0663] 2. The device notifies the user of the received advice and displays it to them. This advice is provided as text, audio, or in some cases as a guidance video.

[0664] Specific examples

[0665] Example 1: Machine part replacement

[0666] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[0667] 2. The device sends video and audio data to the server.

[0668] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[0669] 4. The server generates advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool."

[0670] 5. The device displays the received advice to the user and also plays it back aloud.

[0671] 6. The user replaces the part following the appropriate procedure.

[0672] Example 2: Checking for abnormal sounds

[0673] 1. A user records an abnormal sound made by a machine on their smartphone and asks by text, "Is this sound normal?"

[0674] 2. The device sends video and audio data to the server.

[0675] 3. The server uses voice analysis to detect abnormal sounds and analyzes the text input to understand the question.

[0676] 4. The server generates the advice, "This noise is caused by worn bearings. We recommend replacing the bearings."

[0677] 5. The device displays the received advice to the user as text and plays it back aloud.

[0678] 6. The user follows the advice, checks the cause of the abnormal sound, and takes action.

[0679] This concludes the description of the "Multimodal Factory Coach" system, which enables efficient and intuitive training in manufacturing workplaces.

[0680] The processing flow will be explained below.

[0681] ---

[0682] In case of replacing machine parts

[0683] Step 1:

[0684] The user takes a video of the machine operation using a smartphone, and uses the device's camera function to record the specific operation points.

[0685] Step 2:

[0686] The user can ask questions by voice about things they don't understand, such as, "Please tell me how to remove this bolt."

[0687] Step 3:

[0688] Temporarily saves video and audio data captured by the device. Stores data in internal storage.

[0689] Step 4:

[0690] Your device establishes an internet connection, using Wi-Fi or your mobile data network, to send the stored data to the server.

[0691] Step 5:

[0692] The device uploads video and audio data to the server, which then calls the appropriate API to package and transmit the data.

[0693] Step 6:

[0694] The server analyzes the received video data and uses an image analysis algorithm to identify the location of the bolt in each video frame.

[0695] Step 7:

[0696] The server converts the voice data into text. Using voice recognition technology, the server converts the user's speech into a string of characters.

[0697] Step 8:

[0698] The server analyzes the converted text and understands the question, using natural language processing technology to analyze the intent of the question.

[0699] Step 9:

[0700] The server integrates and analyzes the video and text data, and uses a multimodal AI model to generate optimal advice in response to the user's question.

[0701] Step 10:

[0702] Formatting server-generated advice into text and speech, e.g. "Turn this bolt clockwise to remove it. Use the special tool."

[0703] Step 11:

[0704] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[0705] Step 12:

[0706] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[0707] Step 13:

[0708] The user follows the displayed advice and continues the part replacement operation.

[0709] ---

[0710] Checking for abnormal machine noise

[0711] Step 1:

[0712] The user records abnormal sounds from the machine using their smartphone. The abnormal sounds are recorded using the device's microphone function.

[0713] Step 2:

[0714] The user texts in a question: "Is this sound normal?"

[0715] Step 3:

[0716] The device temporarily stores captured voice data and entered text. Stores data in internal storage.

[0717] Step 4:

[0718] Your device establishes an internet connection, using Wi-Fi or your mobile data network, to send the stored data to the server.

[0719] Step 5:

[0720] The device uploads the voice and text data to the server, which calls the appropriate API to package and transmit the data.

[0721] Step 6:

[0722] The server analyzes the received audio data and uses an audio analysis algorithm to detect abnormal sound patterns.

[0723] Step 7:

[0724] The server analyzes the input text data and understands the question. It uses natural language processing technology to analyze the intent of the question.

[0725] Step 8:

[0726] The server integrates and analyzes the voice and text data, and uses a multimodal AI model to generate optimal advice for the user's question.

[0727] Step 9:

[0728] Formatting server-generated advice into text and speech, e.g. "This sound is caused by worn bearings. We recommend replacing the bearings."

[0729] Step 10:

[0730] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[0731] Step 11:

[0732] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[0733] Step 12:

[0734] The user follows the displayed advice to identify the cause of the abnormal sound and take action.

[0735] ---

[0736] The above are the specific processing steps of the "Multimodal Factory Coach" system.

[0737] Example 1

[0738] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0739] To resolve the shortage of trainers and lack of know-how sharing in the manufacturing industry, provide efficient and intuitive training, and enable users to receive appropriate advice in real time. There is a demand for time-saving and effortless responses to complex machine operations and abnormality detection on manufacturing sites. Conventional methods lack the mechanisms for integrating and analyzing multimodal data such as video, audio, and text to provide appropriate advice.

[0740] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0741] In this invention, the server includes means for acquiring video, means for acquiring audio, means for accepting text, means for analyzing the acquired video, means for analyzing the acquired audio, means for analyzing input text, means for providing advice to the user based on the analysis results, and means for integrating and analyzing the acquired data, thereby making it possible to perform an integrated analysis of multimodal data and provide appropriate advice to the user in real time.

[0742] The "means for acquiring images" is a function for capturing visual information of the surroundings using a sensor device such as a camera mounted on the equipment operated by the user.

[0743] The "means for acquiring audio" is a function for collecting surrounding audio information using a microphone or audio capture device installed in the device used by the user.

[0744] "Means for accepting text" refers to a function that receives and records text information entered by a user using a keyboard or touch screen.

[0745] "Means for analyzing the captured video" refers to software algorithms or machine learning models that process the captured video data and recognize specific actions or objects.

[0746] The "means for analyzing the captured voice" refers to a voice recognition system or natural language processing technology that converts the captured voice data into text and further understands its content.

[0747] The "means for analyzing input text" is a natural language processing algorithm that processes the character data entered by the user and understands its content.

[0748] The "means for providing advice to the user based on the analysis results" is a system that integrates the analyzed video, audio, and text data, generates appropriate advice, and notifies or displays that advice to the user.

[0749] "Means for integrating and analyzing acquired data" refers to data fusion technology and analysis algorithms that process video, audio, and text data in an integrated manner and make comprehensive judgments.

[0750] "Computer system" refers collectively to hardware and software for receiving, processing, analyzing data, and generating advice.

[0751] The "user interface" refers to an interface such as a display device or audio output device for directly providing the generated advice to the user.

[0752] The "Multimodal Factory Coach" system is a system that provides efficient training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. This system analyzes multimodal data from video, audio, and text input, and provides intuitive and appropriate advice to users.

[0753] System Configuration

[0754] The system of the present invention includes the following major components:

[0755] 1. Terminal: A mobile terminal that provides an interface for users to operate. Specifically, a smartphone or tablet is used.

[0756] Means of acquiring images (camera)

[0757] A means of capturing audio (microphone)

[0758] A means of accepting text (keyboard or touchscreen)

[0759] Data transmission method (Internet data upload function)

[0760] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[0761] A method for analyzing video (video analysis algorithm)

[0762] A means of analyzing the voice (voice recognition technology, e.g., Google Cloud Speech-to-Text API)

[0763] Means of analyzing text (natural language processing technology)

[0764] A means of generating advice (generative AI model)

[0765] Data reception method (data download function via the Internet)

[0766] 3. Communication network: The infrastructure for sending and receiving data between devices and servers, usually using an internet connection.

[0767] Specific actions

[0768] Data Capture

[0769] A user uses a device such as a smartphone to record video of a machine operation. For example, a user may record a video of a machine part replacement and ask a question by voice, such as, "Please tell me how to remove this bolt." In a noisy environment, text input is also possible, and the user may input, "Is this sound normal?" The device temporarily stores this video, audio, and text data.

[0770] Data transmission

[0771] The device sends the captured data to a server, using an internet connection and uploading the data securely.

[0772] Data reception and analysis

[0773] The server receives and analyzes the data sent from the terminal. It uses video analysis means to detect specific actions and components in each frame. It uses audio analysis means to convert the audio data into text and analyze the question. It also uses input text analysis means to understand the question.

[0774] Advice Generation

[0775] The server integrates the analyzed data and generates optimal advice. For example, it generates advice on specific steps and tool selection, such as "Turn this bolt clockwise to remove it. Use the appropriate tool." Using a generative AI model, it is possible to generate the optimal advice that the user desires.

[0776] Providing advice

[0777] The server sends the generated advice to the terminal, which receives it and notifies and displays it to the user. The advice is provided as text, audio, or in some cases, guidance video. The user can follow the instructions and perform the appropriate task.

[0778] Specific examples

[0779] Example 1: Machine part replacement

[0780] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[0781] 2. The device sends video and audio data to the server.

[0782] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[0783] 4. The server generates advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool."

[0784] 5. The device displays the received advice to the user and also plays it back aloud.

[0785] 6. The user replaces the part following the appropriate procedure.

[0786] Example 2: Checking for abnormal sounds

[0787] 1. A user records an abnormal sound made by a machine on their smartphone and asks by text, "Is this sound normal?"

[0788] 2. The device sends the voice data and text data to the server.

[0789] 3. The server uses voice analysis to detect abnormal sounds and analyzes the text input to understand the question.

[0790] 4. The server generates the advice, "This noise is caused by worn bearings. We recommend replacing the bearings."

[0791] 5. The device displays the received advice to the user as text and plays it back aloud.

[0792] 6. The user follows the advice, checks the cause of the abnormal sound, and takes action.

[0793] In this way, the system enables real-time support for users, supporting rapid response and efficient work on-site.

[0794] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0795] Step 1: Data Capture

[0796] A user uses a smartphone to record video of a machine being operated and ask questions by voice. The input is video data and audio data, and the output is data temporarily stored on the device. For example, if a user records a machine part replacement and asks, "Please tell me how to remove this bolt," this data is stored on the device.

[0797] Step 2: Send data

[0798] The device sends the stored video and audio data to the server. The input is the data stored on the device, and the output is the data uploaded to the server. The data is sent via the Internet and received by the server.

[0799] Step 3: Receiving data

[0800] The server receives the data sent from the device. The input is the video and audio data sent from the device, and the output is the data saved on the server. After receiving, the data is stored in an appropriate folder for analysis.

[0801] Step 4: Video analysis

[0802] The server analyzes the received video data. The input is the video data, and the output is analysis data that identifies specific actions and objects. For example, a video analysis algorithm is used to analyze the position of bolts and the user's hand movements in each frame.

[0803] Step 5: Audio analysis

[0804] The server converts the received voice data into text and analyzes the question. The input is voice data, and the output is text data and the analysis results. Using a voice recognition system, for example, the voice data is converted into text, such as "Please tell me how to remove this bolt," and the intent is understood.

[0805] Step 6: Text Analysis

[0806] The server analyzes the text data and understands the question. The input is text data, and the output is an analysis result that indicates the intent of the question. Natural language processing technology is used to determine what information the input text is seeking.

[0807] Step 7: Data Integration

[0808] The server integrates the results of video analysis, audio analysis, and text analysis. The input is the results of each analysis, and the output is the integrated analysis data. For example, the server can combine and analyze the location information of a bolt and the content of a voice question to determine how to remove the bolt.

[0809] Step 8: Advice Generation

[0810] The server generates advice based on the integrated analysis data. The input is the integrated analysis data, and the output is specific advice. Using a generative AI model, advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool" is generated.

[0811] Step 9: Send Advice

[0812] The server sends the generated advice to the terminal. The input is the advice content, and the output is the data sent to the terminal. The advice is sent to the terminal via the Internet.

[0813] Step 10: Providing advice

[0814] The device notifies and displays the received advice to the user. The input is advice data sent from the server, and the output is text, audio, and guidance video displayed to the user. For example, it may play audio such as "Turn this bolt clockwise to remove it," encouraging the user to take specific action.

[0815] The above is the specific processing flow of the "Multimodal Factory Coach" system program.

[0816] (Application example 1)

[0817] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0818] In factory operations, there is a need to provide appropriate training and advice in real time when operating or maintaining robots. However, conventional systems lack human trainers, making it difficult to efficiently share know-how. This results in problems that cannot be dealt with quickly, leading to a decline in production efficiency. In particular, there is a need for technology that can provide effective support when an immediate response is required, such as when complex operations or abnormal sounds are generated.

[0819] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0820] In this invention, the server includes a means for analyzing video, a means for analyzing audio, and a means for analyzing text. This allows the server to analyze the video, audio, and text data captured by the user and provide advice based on the analysis results in real time using a generative AI model. Specifically, the server identifies the work content from the video data, understands the user's questions and problems from the audio and text data, and generates appropriate advice. By performing data analysis on the server side using prompt statements, the server can provide quick and accurate support to on-site workers and quickly respond to problems when they occur.

[0821] "Means for capturing images" refers to the function of acquiring visual information as digital data using a camera or sensor.

[0822] "Means for capturing audio" refers to a function for obtaining audio information as digital data using a microphone.

[0823] "Means for accepting text input" refers to the ability to input text information using a keyboard or touch screen.

[0824] "Means for analyzing captured video" refers to the function of processing and analyzing acquired video data using algorithms and AI.

[0825] The "means for analyzing captured audio" is a function for converting acquired audio data into text and analyzing the content.

[0826] "Means for analyzing input text" refers to a function that understands the character information input by the user and analyzes the context and meaning.

[0827] "Means for providing advice to the user based on the analysis results" is a function that provides the user with appropriate instructions and advice based on the analysis results of video, audio, and text.

[0828] "Means for notifying the user of the analysis results in real time" is a function that enables a quick response by immediately transmitting the analysis results to the user.

[0829] A "server" is a computer system that performs analysis and advice generation, and is often located in a cloud environment.

[0830] A "generative AI model" is an artificial intelligence model that generatively provides specific advice and answers in response to a user's questions.

[0831] A "prompt statement" is an instruction statement that gives AI instructions for analysis or generation.

[0832] This invention provides a system that provides an application called "Smart Robo Coach" that is installed on robots in factories. This system captures video, audio, and text, analyzes the data in real time, and provides appropriate advice to users.

[0833] The system has the following main components:

[0834] 1. Terminal

[0835] Camera: Built into the robot to acquire visual information.

[0836] Microphone: Captures audio information.

[0837] Text input interface: A user inputs text information by operating a keyboard or touch screen.

[0838] 2. Server

[0839] Video analysis method: The acquired video data is processed and analyzed using algorithms and AI.

[0840] Voice analysis means: Converts acquired voice data into text and analyzes the content.

[0841] Text analysis means: Understands the text information entered by the user and analyzes the context and meaning.

[0842] Advice generation method: Uses a generative AI model to generate appropriate advice based on the analysis results.

[0843] Data transmission and reception means: Data is transmitted from the terminal to the server, and the server notifies the terminal of the analysis results.

[0844] 3. Communication Network

[0845] It is the infrastructure for sending and receiving data between terminals and servers.

[0846] The specific operation is as follows.

[0847] Video, audio and text capture

[0848] 1. The user uses the robot's built-in camera to capture video of the work being done.

[0849] 2. Users can ask questions by voice about any points they do not understand or are unsure about. In noisy environments, a text input interface can also be used.

[0850] 3. The device temporarily stores the captured video data, audio data, and text data.

[0851] Data upload and analysis

[0852] 1. The device sends the captured data to the server in real time.

[0853] 2. The server analyzes the received video data and detects specific actions and parts in real time.

[0854] 3. The server converts the voice data into text and analyzes the user's question.

[0855] 4. The server analyzes the text input data and understands the question.

[0856] 5. The server integrates and analyzes the video, audio, and text data, and uses a generative AI model to generate optimal advice.

[0857] Advice generation and display

[0858] 1. The server generates advice and sends it to the device as data.

[0859] 2. The device notifies the user of the received advice and displays it to them. This advice is provided as text, audio, or in some cases as a guidance video.

[0860] The hardware used is a camera, microphone, and communication network built into the robot, and the software is video analysis, audio analysis, and generative AI models (such as GPT-3) implemented in Python and other languages.

[0861] As a specific example, consider a robot trying to remove a machine part. When an operator asks, "Please tell me how to remove this bolt," the robot's built-in camera captures an image of the bolt and the question is recorded by a microphone. This data is sent to a server, which analyzes it and generates advice such as, "Turn the bolt clockwise to remove it."

[0862] An example of a prompt for the generative AI model is as follows:

[0863] "Your task is to analyze video and audio data captured by robots in a factory and provide the best advice for a given task. For example, how to remove a bolt or the cause of an unusual noise."

[0864] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0865] Step 1:

[0866] The user captures video of the robot working using the built-in camera and asks questions by voice using the microphone. The input is video data and audio data. The output is the captured video file and audio file. Specifically, the camera captures video at a specified frame rate, and the microphone records audio.

[0867] Step 2:

[0868] The device temporarily stores captured video and audio data. The input is the captured video and audio files. The output is the video and audio files stored in the device's storage. Specifically, the video and audio files are stored in a specific directory.

[0869] Step 3:

[0870] The device sends stored video data, audio data, and text input data to the server in real time. The input is the video file, audio file, and text data. The output is the data sent to the server. Specifically, the data is uploaded to the server using the HTTPS protocol.

[0871] Step 4:

[0872] The server analyzes the received video data and detects specific actions and parts in real time. The input is a video file. The output is the analyzed video data, such as the location information of specific parts. Specifically, it uses OpenCV and machine learning models to analyze the video data frame by frame.

[0873] Step 5:

[0874] The server converts the received voice data into text and analyzes the question content. The input is an audio file. The output is text data and the analyzed question content. Specifically, the voice data is transcribed using the SpeechRecognition library and the content is analyzed using natural language processing (NLP) algorithms.

[0875] Step 6:

[0876] The server analyzes the input text data and understands the question. The input is text data. The output is the analyzed question. Specifically, it uses NLP algorithms to analyze the context and meaning of the text.

[0877] Step 7:

[0878] The server integrates and analyzes video data, audio data, and text data, and generates optimal advice using a generative AI model. The input is video data, audio data, and text data. The output is text or audio data as advice. Specifically, the server integrates the results of video analysis and text analysis, and generates advice using a generative AI model (e.g., GPT-3).

[0879] Step 8:

[0880] The server sends the generated advice to the terminal as data. The input is text or audio data as advice. The output is the advice data sent to the terminal. Specifically, the data is sent to the terminal using the HTTPS protocol.

[0881] Step 9:

[0882] The terminal notifies and displays the received advice to the user. This advice is provided as text or audio, or in some cases as a guidance video. The input is text or audio data as advice. The output is the advice notified to the user. Specifically, the text is displayed on the display and audio is played from the speaker.

[0883] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0884] ---

[0885] This invention is a system that provides efficient and intuitive training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. Furthermore, it aims to improve the user experience and training effectiveness by adding an emotion engine that recognizes the user's emotions.

[0886] System Configuration

[0887] The system of the present invention includes the following major components:

[0888] 1. Terminal: A mobile terminal such as a smartphone that provides an interface for users to operate.

[0889] Video capture method

[0890] Audio capture method

[0891] Text input method

[0892] Data transmission method

[0893] Emotion Recognition Module

[0894] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[0895] Video analysis methods

[0896] Voice analysis methods

[0897] Text Analysis Methods

[0898] Advice Generation Method

[0899] Data Receiving Method

[0900] Emotion Engine

[0901] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[0902] Specific actions

[0903] Video, audio, text and emotional capture

[0904] 1. The user takes a video of the machine operation using a device such as a smartphone.

[0905] 2. Users can ask questions by voice if they have any questions or concerns. In noisy environments, they can input text.

[0906] 3. The device temporarily stores the video data, audio data, and text data.

[0907] 4. The emotion recognition module analyzes the user's emotions in real time from the video and audio data.

[0908] Data upload and analysis

[0909] 1. The device sends the captured data to the server.

[0910] 2. The server analyzes the received video data and detects specific actions and parts in each frame.

[0911] 3. The server converts the voice data into text and analyzes the user's question.

[0912] 4. The server analyzes the text input data and understands the question.

[0913] 5. The emotion engine analyzes the received emotion data and detects the user's stress level and emotional state.

[0914] 6. The server integrates and analyzes the video, audio, text, and emotional data to generate optimal advice.

[0915] Advice generation and display

[0916] 1. The server formats the generated advice into text and speech. An emotion engine adjusts the advice content to the user's emotional state.

[0917] For example, advice like "Turn this bolt clockwise to remove it. Use the appropriate tool" becomes "Turn this bolt clockwise slowly. If you have trouble, use the appropriate tool" if the user is feeling stressed.

[0918] 2. The server sends the generated advice data to the terminal, securing an appropriate network path to transfer the data.

[0919] 3. The device displays the advice data received from the server by displaying a text message on the screen and playing the advice via audio output.

[0920] Specific examples

[0921] Example 1: Machine part replacement

[0922] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[0923] 2. The device sends video and audio data to the server.

[0924] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[0925] 4. The emotion engine analyzes the user's emotional state from video and audio.

[0926] 5. The server combines video, audio, text, and emotion data to generate advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[0927] 6. The device displays the received advice to the user and also plays it back aloud.

[0928] 7. The user replaces the part according to the advice.

[0929] Example 2: Checking for abnormal sounds

[0930] 1. A user records an abnormal machine sound on their smartphone and asks via text, "Is this sound normal?"

[0931] 2. The device sends the voice and text data to the server.

[0932] 3. The server uses voice analysis to detect abnormal sound patterns and uses text analysis to understand the content of the question.

[0933] 4. The emotion engine analyzes the user's emotional state.

[0934] 5. The server generates advice on the cause of the abnormal sound and how to deal with it, for example, by telling a nervous user, "This sound is caused by worn bearings. It would be a good idea to replace the bearings."

[0935] 6. The device displays the received advice to the user and plays it aloud.

[0936] 7. The user follows the advice to identify the cause of the abnormal sound and take action.

[0937] This concludes the description of the "Multimodal Factory Coach" system, which enables efficient and intuitive training in manufacturing workplaces and provides personalized support based on the user's emotions.

[0938] The processing flow will be explained below.

[0939] ---

[0940] In case of replacing machine parts

[0941] Step 1:

[0942] The user takes a video of the machine operation using a smartphone, and uses the device's camera function to record the specific operation points.

[0943] Step 2:

[0944] The user can ask questions by voice about things they don't understand, such as, "Please tell me how to remove this bolt."

[0945] Step 3:

[0946] Temporarily saves video and audio data captured by the device. Stores data in internal storage.

[0947] Step 4:

[0948] The device's emotion recognition module analyzes the audio and video data to determine the user's emotional state, for example, determining whether the user is nervous based on the tone of their voice or facial expression.

[0949] Step 5:

[0950] The device establishes an internet connection using Wi-Fi or a mobile data network to send video data, audio data, text data, and emotion data to the server.

[0951] Step 6:

[0952] The device uploads all data to the server, which calls the appropriate API to package and transmit the data.

[0953] Step 7:

[0954] The server analyzes the received video data and uses an image analysis algorithm to identify the location of the bolt in each video frame.

[0955] Step 8:

[0956] The server converts the voice data into text. Using voice recognition technology, the server converts the user's speech into a string of characters.

[0957] Step 9:

[0958] The server analyzes the converted text and understands the question, using natural language processing technology to analyze the intent of the question.

[0959] Step 10:

[0960] The server's emotion engine analyzes the emotion data to understand the user's stress level and emotional state.

[0961] Step 11:

[0962] The server integrates and analyzes video data, audio data, text data, and emotional data, and uses a multimodal AI model to generate optimal advice in response to the user's question.

[0963] Step 12:

[0964] The server adjusts the advice based on the user's emotional data. For example, if the user is nervous, the advice will be adjusted to something like, "Turn this bolt clockwise slowly. If you have trouble, use the appropriate tool."

[0965] Step 13:

[0966] Formatting server-generated advice into text and speech, e.g. "Turn this bolt clockwise to remove it. Use the special tool."

[0967] Step 14:

[0968] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[0969] Step 15:

[0970] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[0971] Step 16:

[0972] The user follows the displayed advice and continues the part replacement operation.

[0973] ---

[0974] Checking for abnormal machine noise

[0975] Step 1:

[0976] The user records abnormal sounds from the machine using their smartphone. The abnormal sounds are recorded using the device's microphone function.

[0977] Step 2:

[0978] The user texts in a question: "Is this sound normal?"

[0979] Step 3:

[0980] The device temporarily stores captured voice data and entered text. Stores data in internal storage.

[0981] Step 4:

[0982] The device's emotion recognition module analyzes the video and audio data to determine the user's emotional state.

[0983] Step 5:

[0984] The device establishes an internet connection using Wi-Fi or a mobile data network to send voice, text, and emotion data to the server.

[0985] Step 6:

[0986] The device uploads all data to the server, which calls the appropriate API to package and transmit the data.

[0987] Step 7:

[0988] The server analyzes the received audio data and uses an audio analysis algorithm to detect abnormal sound patterns.

[0989] Step 8:

[0990] The server analyzes the input text data and understands the question. It uses natural language processing technology to analyze the intent of the question.

[0991] Step 9:

[0992] The server's emotion engine analyzes the emotion data to understand the user's stress level and emotional state.

[0993] Step 10:

[0994] The server integrates and analyzes voice data, text data, and emotion data, and uses a multimodal AI model to generate optimal advice for the user's question.

[0995] Step 11:

[0996] The server adjusts the advice based on the user's emotional data. For example, if the user is nervous, the advice will be adjusted to say, "This noise is caused by worn bearings. It would be a good idea to replace the bearings."

[0997] Step 12:

[0998] Formatting server-generated advice into text and speech, e.g. "This sound is caused by worn bearings. We recommend replacing the bearings."

[0999] Step 13:

[1000] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[1001] Step 14:

[1002] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[1003] Step 15:

[1004] The user follows the displayed advice to identify the cause of the abnormal sound and take action.

[1005] ---

[1006] The above are the specific processing steps when combining the "Multimodal Factory Coach" system with an emotion engine.

[1007] Example 2

[1008] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1009] In modern manufacturing, a lack of trainers and a lack of sharing of know-how are serious problems. This makes it difficult for new or inexperienced employees to receive training quickly and effectively, resulting in reduced work efficiency. Furthermore, simple advice that does not take users' emotions into consideration makes it difficult to address the stress and tension they feel, preventing optimal performance. A system that can solve these problems and provide efficient and intuitive training is needed.

[1010] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1011] In this invention, the server

[1012] a means for capturing video;

[1013] a means for capturing audio;

[1014] means for accepting text input;

[1015] a means for analyzing emotions;

[1016] means for analyzing the captured video;

[1017] means for analyzing the captured audio;

[1018] means for parsing input text;

[1019] means for analyzing user emotion data;

[1020] and means for providing advice to the user based on the analysis results.

[1021] This makes it possible to provide effective and intuitive training that takes into account the stress and tension felt by the user. By analyzing the user's emotional state in real time and generating and providing appropriate advice based on that state, the user's understanding and work efficiency are improved.

[1022] "Means for capturing video" refers to devices or functions that capture the equipment operated by the user and the work environment in real time, and record and save the video data.

[1023] "Audio capturing means" refers to a device or function that collects the user's voice using a microphone or the like and records it as audio data.

[1024] "Means for accepting text input" refers to a function that allows a user to input text data using an interface such as a keyboard or touch screen.

[1025] "Means for analyzing emotions" refers to software or algorithms that use video and audio data of users to recognize and evaluate their emotional state, such as stress and tension, in real time.

[1026] "Means for analyzing captured video" refers to a function that analyzes received video data and detects and identifies specific actions, positions, parts, etc.

[1027] The "means for analyzing captured audio" is a function that converts collected audio data into text data and further analyzes the content of that text data.

[1028] "Means for analyzing input text" refers to software or algorithms that understand the content of the text data entered by the user and extract the necessary information.

[1029] The "means for analyzing the user's emotional data" is a function that analyzes the user's emotional state from the user's video data and audio data, and evaluates the stress level and tension.

[1030] The "means for providing advice to the user based on the analysis results" is a function for generating and providing optimal instructions and advice to the user based on the analysis results of video, audio, text, and emotional data.

[1031] This invention is a system for providing efficient and intuitive training, resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. The system aims to improve the effectiveness of training by recognizing the user's emotions and providing appropriate advice.

[1032] System Configuration

[1033] The system includes the following major components:

[1034] 1. Terminal: A mobile terminal operated by the user, equipped with the following functions:

[1035] Video capture method

[1036] Audio capture method

[1037] Text input method

[1038] Data transmission method

[1039] Emotion Recognition Module

[1040] 2. Server: A computer system for analysis and advice generation located in a cloud environment.

[1041] Video analysis methods

[1042] Voice analysis methods

[1043] Text Analysis Methods

[1044] Advice Generation Method

[1045] Data Receiving Method

[1046] Emotion Engine

[1047] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[1048] Specific actions

[1049] Video, audio, text and emotional capture

[1050] 1. The user uses a device such as a smartphone to record video of the machine operation. For example, while recording the operation, the user can ask a question by voice, such as "Please tell me how to remove this bolt."

[1051] 2. If the user is in a noisy environment, he or she inputs text. For example, the user inputs text such as "Is this sound normal?"

[1052] 3. The device temporarily stores the video data, audio data, and text data.

[1053] 4. The emotion recognition module analyzes the user's emotions in real time from their video and audio data, for example, determining whether they are nervous.

[1054] Sending data

[1055] 1. The device sends all captured data to the server.

[1056] 2. Use a high-speed and stable communication network for data transmission.

[1057] Data analysis on the server

[1058] 1. The server analyzes the video data received by the server using video analysis means to detect specific actions and parts in each frame. For example, it identifies the position of a bolt.

[1059] 2. The server converts the voice data into text and uses speech analysis to understand the question. For example, it analyzes the question, "Please tell me how to remove this bolt."

[1060] 3. The server analyzes the text input data and understands the question.

[1061] 4. The emotion engine analyzes the received emotion data and detects the user's stress level and emotional state. For example, it recognizes that the user is nervous.

[1062] Generating Advice

[1063] 1. The server integrates and analyzes video, audio, text, and emotion data to generate optimal advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[1064] 2. The emotion engine adjusts advice according to the user's emotional state.

[1065] Providing advice and feedback

[1066] 1. The server sends the generated advice to the device.

[1067] 2. The device displays the received advice data to the user and also plays it aloud. For example, the device might display the message "Turn this bolt clockwise slowly. Use a tool if necessary" on the screen and simultaneously play it aloud.

[1068] 3. The user follows the advice and performs the task. During this time, the device monitors the user's emotional state and sends feedback to the server as needed.

[1069] Specific examples

[1070] Example 1: Replacing machine parts

[1071] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[1072] 2. The device sends video and audio data to the server.

[1073] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[1074] 4. The emotion engine analyzes the user's emotional state from video and audio.

[1075] 5. The server combines video, audio, text, and emotion data to generate advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[1076] 6. The device displays the received advice to the user and also plays it back aloud.

[1077] 7. The user replaces the part according to the advice.

[1078] Example 2: Checking for abnormal sounds

[1079] 1. A user records an abnormal machine sound on their smartphone and asks via text, "Is this sound normal?"

[1080] 2. The device sends the voice and text data to the server.

[1081] 3. The server uses voice analysis to detect abnormal sound patterns and uses text analysis to understand the content of the question.

[1082] 4. The emotion engine analyzes the user's emotional state.

[1083] 5. The server generates advice on the cause of the abnormal sound and how to deal with it, for example, by telling a nervous user, "This sound is caused by worn bearings. It would be a good idea to replace the bearings."

[1084] 6. The device displays the received advice to the user and plays it aloud.

[1085] 7. The user follows the advice to identify the cause of the abnormal sound and take action.

[1086] Prompt Sentence Examples

[1087] "Please tell me how to replace the machine parts. I'm a little nervous."

[1088] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1089] Step 1:

[1090] The user captures video using a device. For example, the user may record video of a machine operation or part replacement using a smartphone, and then ask a question by voice, such as "Please tell me how to remove this bolt." Specifically, the user's operation of the equipment is recorded, and the audio is collected using the device's microphone.

[1091] Input: Video and audio data captured by a smartphone.

[1092] Output: The captured video and audio data is saved on the device.

[1093] Step 2:

[1094] When a user is in a noisy environment, he or she uses the device to input text. For example, he or she inputs text such as "Is this sound normal?". The input is performed using the device's keyboard or touch screen.

[1095] Input: Text data that a user types into a terminal.

[1096] Output: The entered text data is saved to the terminal.

[1097] Step 3:

[1098] The device temporarily stores video, audio, and text data and prepares it for later transmission to the server. The data is temporarily stored in the device's memory.

[1099] Input: Captured video, audio, and text data.

[1100] Output: Video, audio, and text data temporarily stored in the device's memory.

[1101] Step 4:

[1102] The emotion recognition module analyzes the user's emotions in real time from their video and audio data. For example, it analyzes the user's facial expressions and tone of voice to detect tension or stress. This emotion analysis is performed in real time and is executed on the device.

[1103] Input: Captured video and audio data.

[1104] Output: Parsed user emotion data.

[1105] Step 5:

[1106] All data captured by the device (video data, audio data, text data, and emotion data) is sent to the server. The data is transferred using a high-speed, stable communication network.

[1107] Input: Temporarily stored video, audio, and text data, as well as analyzed emotion data.

[1108] Output: Video, audio, text data and emotion data sent to the server.

[1109] Step 6:

[1110] The server analyzes the received video data. Using video analysis means, it detects specific actions and parts in each frame. For example, analysis is performed to identify the location and operation method of a bolt in the video.

[1111] Input: Video data sent to the server.

[1112] Output: Analyzed video data.

[1113] Step 7:

[1114] The server converts the received voice data into text and uses voice analysis means to understand the content of the user's question. For example, it analyzes voice data such as "Please tell me how to remove this bolt."

[1115] Input: The audio data sent to the server.

[1116] Output: The audio data converted to text and the parsed questions.

[1117] Step 8:

[1118] The server uses text analysis means to understand the content of the input text data, for example, analyzing the content of a question such as "Is this sound normal?"

[1119] Input: The text data sent to the server.

[1120] Output: Parsed text data and question content.

[1121] Step 9:

[1122] The emotion engine analyzes the received emotion data to detect the user's stress level and emotional state, for example, assessing whether the user is nervous.

[1123] Input: Emotion data sent to the server.

[1124] Output: The analyzed emotional state of the user.

[1125] Step 10:

[1126] The server combines video, audio, text, and emotion data to generate optimal advice. For example, if the user is nervous, the server generates advice such as, "Turn this bolt clockwise slowly. Use a tool if necessary."

[1127] Input: Analyzed video data, audio data, text data, and emotion data.

[1128] Output: The generated advice data.

[1129] Step 11:

[1130] The server sends the generated advice data to the terminal, securing an appropriate network path for rapid transfer.

[1131] Input: Generated advice data.

[1132] Output: Advice data sent to the terminal.

[1133] Step 12:

[1134] The device displays the received advice data to the user. It displays a text message on the screen and plays the advice aloud. For example, the device might display "Turn this bolt slowly clockwise. Use a tool if necessary" and play the advice aloud.

[1135] Input: Advice data sent by the server.

[1136] Output: The advice that was displayed and played back to the user.

[1137] The above is a flow of the specific processing steps of the system. It includes details of the data processing and calculations performed at each step, which allows for effective and intuitive training for users.

[1138] (Application example 2)

[1139] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1140] In the manufacturing industry, there are problems with efficient and intuitive training due to a shortage of trainers and insufficient sharing of know-how. Furthermore, training and advice that do not take into account the user's emotional state can reduce its effectiveness in the field.

[1141] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video, means for capturing audio, means for accepting text input, means for analyzing the user's emotional state, means for providing advice to the user based on the analysis results and the user's emotional state, means for transmitting the captured video and audio to the server, means for displaying the advice transmitted from the server on the user terminal, means for analyzing prompt sentences using a generative AI model and generating advice corresponding to the emotion, and means for integrating and analyzing the user's video data, audio data, and text data, inputting the prompt sentences into the generative AI model, and generating advice based on the analysis results. This enables efficient and intuitive training and advice provision while taking the user's emotional state into consideration.

[1142] A "video capture means" is a device or method for obtaining images or video of a user or the work environment.

[1143] An "audio capturing means" is a device or method for recording the sounds of a user or the work environment.

[1144] The "means for accepting text input" refers to a device or method that allows a user to input characters using a keyboard, an on-screen keyboard, or the like.

[1145] A "means for analyzing a user's emotional state" is a device or method for inferring a user's emotional state using audio, video, and other data.

[1146] The "means for providing advice to a user based on the analysis results and the emotional state of the user" is a device or method for combining the results of data analysis with the emotional state to provide the user with optimal advice.

[1147] The "means for transmitting captured video and audio to a server" refers to a device or method for transferring captured video and audio data to a server via a network.

[1148] The "means for displaying advice sent from the server on the user terminal" refers to a device or method for displaying instructions or advice received from the server in a form that can be confirmed by the user.

[1149] "Means for analyzing prompt sentences using a generative AI model and generating advice based on emotions" refers to a device or method that utilizes AI technology to analyze text data entered by a user and generate appropriate advice based on that data and the user's emotional state.

[1150] "Means for integrating and analyzing a user's video data, audio data, and text data, inputting prompt sentences into a generative AI model, and generating advice based on the analysis results" refers to a device or method for integrating and analyzing multiple data sources, analyzing prompt sentences using AI based on the results, and providing appropriate advice to the user.

[1151] The present invention is a system for a factory robot that provides specific advice while taking into account the emotional state of the user. The system includes the following main components:

[1152] System Configuration

[1153] 1. Terminal: This applies to robots that work in factories. Robots use cameras and microphones to capture video and audio of their work.

[1154] Video capture method

[1155] Audio capture method

[1156] Text input method

[1157] Emotion Recognition Module

[1158] Data transmission method

[1159] 2. Server: A system deployed in a cloud environment that performs analysis and generates advice. The server includes the following means:

[1160] Video analysis methods

[1161] Voice analysis methods

[1162] Text Analysis Methods

[1163] Emotion Recognition Module

[1164] Advice Generation Method

[1165] Data Receiving Method

[1166] Generative AI Models

[1167] 3. Communication network: Infrastructure for sending and receiving data between terminals and servers.

[1168] Operation overview

[1169] Data Capture and Transmission

[1170] As the user performs a task, the device (robot) captures video and audio using a camera and microphone. The device also includes a means for the user to input text using a keyboard or voice input. This data is sent to a server in real time.

[1171] Data analysis and emotion recognition

[1172] The server integrates and analyzes the received video, audio, and text data. The emotion recognition module analyzes the user's emotional state, and the generative AI model analyzes the prompt sentences to generate appropriate advice.

[1173] Advice generation and delivery

[1174] The generated advice is sent from the server to the terminal in text and audio format, and the terminal displays the received advice to the user and plays it back aloud.

[1175] Hardware and software used

[1176] Hardware: Camera, microphone, user input devices (keyboard, touch screen, etc.), robot body

[1177] Software: Video analysis software (OpenCV, etc.), audio analysis software (SpeechRecognition, etc.), emotion recognition software (EmotionRecognizer, etc.), generative AI models (cloud-based AI services, etc.)

[1178] Specific examples

[1179] Machine part replacement

[1180] If a user is unsure how to remove a bolt, they can enter "Please tell me how to remove this bolt" in text format, and the robot will send video and audio to the server. The server will analyze the data and use an emotion recognition module to confirm that the user is nervous. The advice generator will then generate advice such as "Turn this bolt slowly clockwise. Use a tool if necessary," which will then be displayed and played aloud on the device.

[1181] Check for abnormal sounds

[1182] If a user wants to check for abnormal machine sounds, they record the sound on their smartphone and then type "Is this sound normal?" into the robot while playing it back. The robot sends this to the server, which then analyzes the sound pattern using a voice analysis module and detects tension using an emotion recognition module. The server then generates advice such as "This sound is caused by worn bearings. It would be a good idea to replace the bearings," which the robot displays and plays back in audio.

[1183] Prompt Sentence Examples

[1184] Can you tell me how to remove this bolt?

[1185] "Is this sound normal?"

[1186] Through these specific operations, user support is realized through cooperation between the terminal and the server.

[1187] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1188] Step 1:

[1189] The device uses a camera and microphone to capture video and audio of the user working. As input, video data is obtained from the camera and audio data is obtained from the microphone. Specifically, the device records the user's work scene in real time and temporarily stores the data.

[1190] Step 2:

[1191] The user enters a question requesting clarification through text input. As input, the text data entered by the user via a keyboard or touch screen is acquired. Specifically, the device provides an interface for accepting the user's input and saves the entered text.

[1192] Step 3:

[1193] The device sends the captured video, audio, and text data to the server. The device uses the saved video, audio, and text data as input and sends this data to the server as output. Specifically, the device uploads the data to the cloud server using the network.

[1194] Step 4:

[1195] The server analyzes the video data and identifies the user's work. It uses the video data sent from the device as input and obtains the analysis results as output. Specifically, the server uses video analysis software to recognize specific actions and parts in each frame and compares them with a database.

[1196] Step 5:

[1197] The server analyzes the voice data and understands the user's question. It uses the voice data sent from the device as input and outputs the converted text string and the analysis results. Specifically, it uses voice recognition software to convert the voice into text and analyzes the content of the question.

[1198] Step 6:

[1199] The server analyzes the user's emotional state. It uses video and audio data as input and obtains the emotion analysis results as output. Specifically, it uses an emotion recognition module to estimate the user's emotional state from facial expressions, tone of voice, etc.

[1200] Step 7:

[1201] The server integrates the analysis results and the emotional state and generates advice using a generative AI model. The analysis results, emotion analysis results, and prompt sentences are used as input, and advice for the user is obtained as output. Specifically, the generative AI model analyzes the prompt sentence and generates appropriate advice based on the data.

[1202] Step 8:

[1203] The server sends the generated advice to the terminal. The generated advice data is used as input to obtain data to be sent to the terminal as output. In concrete terms, the server uploads the advice data to the terminal via the network.

[1204] Step 9:

[1205] The device displays the advice received from the server to the user and plays it back aloud. It uses the advice data sent from the server as input and provides visual and auditory feedback to the user as output. Specifically, the device displays a text message on the screen and plays back an audio message from the speaker.

[1206] Prompt Sentence Examples

[1207] Can you tell me how to remove this bolt?

[1208] "Is this sound normal?"

[1209] The above are the specific processing steps of this system.

[1210] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1211] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1212] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1213] [Third embodiment]

[1214] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1215] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1216] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1217] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1218] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1219] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1220] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1221] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1222] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1223] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1224] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1225] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1226] ---

[1227] The purpose of this invention is to provide efficient training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. Specifically, it analyzes multimodal data from video, audio, and text input, and provides intuitive and appropriate advice to users.

[1228] System Configuration

[1229] The system of the present invention includes the following major components:

[1230] 1. Terminal: A mobile terminal such as a smartphone that provides an interface for users to operate.

[1231] Video capture method

[1232] Audio capture method

[1233] Text input method

[1234] Data transmission method

[1235] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[1236] Video analysis methods

[1237] Voice analysis methods

[1238] Text Analysis Methods

[1239] Advice Generation Method

[1240] Data Receiving Method

[1241] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[1242] Specific actions

[1243] Video, audio and text capture

[1244] 1. The user takes a photo of the machine operation using a device such as a smartphone.

[1245] 2. Users can ask questions by voice about points they don't understand or points they're unsure about. In some cases, they can input text in a noisy environment.

[1246] 3. The device temporarily stores the video data, audio data, and text data.

[1247] Data upload and analysis

[1248] 1. The device sends the captured data to the server.

[1249] 2. The server analyzes the received video data and detects specific actions and parts in each frame.

[1250] 3. The server converts the voice data into text and analyzes the user's question.

[1251] 4. The server analyzes the text input data and understands the question.

[1252] 5. The server integrates and analyzes the video, audio, and text data to generate optimal advice.

[1253] Advice generation and display

[1254] 1. The server generates advice and sends it to the device as data.

[1255] 2. The device notifies the user of the received advice and displays it to them. This advice is provided as text, audio, or in some cases as a guidance video.

[1256] Specific examples

[1257] Example 1: Machine part replacement

[1258] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[1259] 2. The device sends video and audio data to the server.

[1260] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[1261] 4. The server generates advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool."

[1262] 5. The device displays the received advice to the user and also plays it back aloud.

[1263] 6. The user replaces the part following the appropriate procedure.

[1264] Example 2: Checking for abnormal sounds

[1265] 1. A user records an abnormal sound made by a machine on their smartphone and asks by text, "Is this sound normal?"

[1266] 2. The device sends video and audio data to the server.

[1267] 3. The server uses voice analysis to detect abnormal sounds and analyzes the text input to understand the question.

[1268] 4. The server generates the advice, "This noise is caused by worn bearings. We recommend replacing the bearings."

[1269] 5. The device displays the received advice to the user as text and plays it back aloud.

[1270] 6. The user follows the advice, checks the cause of the abnormal sound, and takes action.

[1271] This concludes the description of the "Multimodal Factory Coach" system, which enables efficient and intuitive training in manufacturing workplaces.

[1272] The processing flow will be explained below.

[1273] ---

[1274] In case of replacing machine parts

[1275] Step 1:

[1276] The user takes a video of the machine operation using a smartphone, and uses the device's camera function to record the specific operation points.

[1277] Step 2:

[1278] The user can ask questions by voice about things they don't understand, such as, "Please tell me how to remove this bolt."

[1279] Step 3:

[1280] Temporarily saves video and audio data captured by the device. Stores data in internal storage.

[1281] Step 4:

[1282] Your device establishes an internet connection, using Wi-Fi or your mobile data network, to send the stored data to the server.

[1283] Step 5:

[1284] The device uploads video and audio data to the server, which then calls the appropriate API to package and transmit the data.

[1285] Step 6:

[1286] The server analyzes the received video data and uses an image analysis algorithm to identify the location of the bolt in each video frame.

[1287] Step 7:

[1288] The server converts the voice data into text. Using voice recognition technology, the server converts the user's speech into a string of characters.

[1289] Step 8:

[1290] The server analyzes the converted text and understands the question, using natural language processing technology to analyze the intent of the question.

[1291] Step 9:

[1292] The server integrates and analyzes the video and text data, and uses a multimodal AI model to generate optimal advice in response to the user's question.

[1293] Step 10:

[1294] Formatting server-generated advice into text and speech, e.g. "Turn this bolt clockwise to remove it. Use the special tool."

[1295] Step 11:

[1296] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[1297] Step 12:

[1298] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[1299] Step 13:

[1300] The user follows the displayed advice and continues the part replacement operation.

[1301] ---

[1302] Checking for abnormal machine noise

[1303] Step 1:

[1304] The user records abnormal sounds from the machine using their smartphone. The abnormal sounds are recorded using the device's microphone function.

[1305] Step 2:

[1306] The user texts in a question: "Is this sound normal?"

[1307] Step 3:

[1308] The device temporarily stores captured voice data and entered text. Stores data in internal storage.

[1309] Step 4:

[1310] Your device establishes an internet connection, using Wi-Fi or your mobile data network, to send the stored data to the server.

[1311] Step 5:

[1312] The device uploads the voice and text data to the server, which calls the appropriate API to package and transmit the data.

[1313] Step 6:

[1314] The server analyzes the received audio data and uses an audio analysis algorithm to detect abnormal sound patterns.

[1315] Step 7:

[1316] The server analyzes the input text data and understands the question. It uses natural language processing technology to analyze the intent of the question.

[1317] Step 8:

[1318] The server integrates and analyzes the voice and text data, and uses a multimodal AI model to generate optimal advice for the user's question.

[1319] Step 9:

[1320] Formatting server-generated advice into text and speech, e.g. "This sound is caused by worn bearings. We recommend replacing the bearings."

[1321] Step 10:

[1322] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[1323] Step 11:

[1324] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[1325] Step 12:

[1326] The user follows the displayed advice to identify the cause of the abnormal sound and take action.

[1327] ---

[1328] The above are the specific processing steps of the "Multimodal Factory Coach" system.

[1329] Example 1

[1330] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1331] To resolve the shortage of trainers and lack of know-how sharing in the manufacturing industry, provide efficient and intuitive training, and enable users to receive appropriate advice in real time. There is a demand for time-saving and effortless responses to complex machine operations and abnormality detection on manufacturing sites. Conventional methods lack the mechanisms for integrating and analyzing multimodal data such as video, audio, and text to provide appropriate advice.

[1332] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1333] In this invention, the server includes means for acquiring video, means for acquiring audio, means for accepting text, means for analyzing the acquired video, means for analyzing the acquired audio, means for analyzing input text, means for providing advice to the user based on the analysis results, and means for integrating and analyzing the acquired data, thereby making it possible to perform an integrated analysis of multimodal data and provide appropriate advice to the user in real time.

[1334] The "means for acquiring images" is a function for capturing visual information of the surroundings using a sensor device such as a camera mounted on the equipment operated by the user.

[1335] The "means for acquiring audio" is a function for collecting surrounding audio information using a microphone or audio capture device installed in the device used by the user.

[1336] "Means for accepting text" refers to a function that receives and records text information entered by a user using a keyboard or touch screen.

[1337] "Means for analyzing the captured video" refers to software algorithms or machine learning models that process the captured video data and recognize specific actions or objects.

[1338] The "means for analyzing the captured voice" refers to a voice recognition system or natural language processing technology that converts the captured voice data into text and further understands its content.

[1339] The "means for analyzing input text" is a natural language processing algorithm that processes the character data entered by the user and understands its content.

[1340] The "means for providing advice to the user based on the analysis results" is a system that integrates the analyzed video, audio, and text data, generates appropriate advice, and notifies or displays that advice to the user.

[1341] "Means for integrating and analyzing acquired data" refers to data fusion technology and analysis algorithms that process video, audio, and text data in an integrated manner and make comprehensive judgments.

[1342] "Computer system" refers collectively to hardware and software for receiving, processing, analyzing data, and generating advice.

[1343] The "user interface" refers to an interface such as a display device or audio output device for directly providing the generated advice to the user.

[1344] The "Multimodal Factory Coach" system is a system that provides efficient training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. This system analyzes multimodal data from video, audio, and text input, and provides intuitive and appropriate advice to users.

[1345] System Configuration

[1346] The system of the present invention includes the following major components:

[1347] 1. Terminal: A mobile terminal that provides an interface for users to operate. Specifically, a smartphone or tablet is used.

[1348] Means of acquiring images (camera)

[1349] A means of capturing audio (microphone)

[1350] A means of accepting text (keyboard or touchscreen)

[1351] Data transmission method (Internet data upload function)

[1352] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[1353] A method for analyzing video (video analysis algorithm)

[1354] A means of analyzing the voice (voice recognition technology, e.g., Google Cloud Speech-to-Text API)

[1355] Means of analyzing text (natural language processing technology)

[1356] A means of generating advice (generative AI model)

[1357] Data reception method (data download function via the Internet)

[1358] 3. Communication network: The infrastructure for sending and receiving data between devices and servers, usually using an internet connection.

[1359] Specific actions

[1360] Data Capture

[1361] A user uses a device such as a smartphone to record video of a machine operation. For example, a user may record a video of a machine part replacement and ask a question by voice, such as, "Please tell me how to remove this bolt." In a noisy environment, text input is also possible, and the user may input, "Is this sound normal?" The device temporarily stores this video, audio, and text data.

[1362] Data transmission

[1363] The device sends the captured data to a server, using an internet connection and uploading the data securely.

[1364] Data reception and analysis

[1365] The server receives and analyzes the data sent from the terminal. It uses video analysis means to detect specific actions and components in each frame. It uses audio analysis means to convert the audio data into text and analyze the question. It also uses input text analysis means to understand the question.

[1366] Advice Generation

[1367] The server integrates the analyzed data and generates optimal advice. For example, it generates advice on specific steps and tool selection, such as "Turn this bolt clockwise to remove it. Use the appropriate tool." Using a generative AI model, it is possible to generate the optimal advice that the user desires.

[1368] Providing advice

[1369] The server sends the generated advice to the terminal, which receives it and notifies and displays it to the user. The advice is provided as text, audio, or in some cases, guidance video. The user can follow the instructions and perform the appropriate task.

[1370] Specific examples

[1371] Example 1: Machine part replacement

[1372] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[1373] 2. The device sends video and audio data to the server.

[1374] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[1375] 4. The server generates advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool."

[1376] 5. The device displays the received advice to the user and also plays it back aloud.

[1377] 6. The user replaces the part following the appropriate procedure.

[1378] Example 2: Checking for abnormal sounds

[1379] 1. A user records an abnormal sound made by a machine on their smartphone and asks by text, "Is this sound normal?"

[1380] 2. The device sends the voice data and text data to the server.

[1381] 3. The server uses voice analysis to detect abnormal sounds and analyzes the text input to understand the question.

[1382] 4. The server generates the advice, "This noise is caused by worn bearings. We recommend replacing the bearings."

[1383] 5. The device displays the received advice to the user as text and plays it back aloud.

[1384] 6. The user follows the advice, checks the cause of the abnormal sound, and takes action.

[1385] In this way, the system enables real-time support for users, supporting rapid response and efficient work on-site.

[1386] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1387] Step 1: Data Capture

[1388] A user uses a smartphone to record video of a machine being operated and ask questions by voice. The input is video data and audio data, and the output is data temporarily stored on the device. For example, if a user records a machine part replacement and asks, "Please tell me how to remove this bolt," this data is stored on the device.

[1389] Step 2: Send data

[1390] The device sends the stored video and audio data to the server. The input is the data stored on the device, and the output is the data uploaded to the server. The data is sent via the Internet and received by the server.

[1391] Step 3: Receiving data

[1392] The server receives the data sent from the device. The input is the video and audio data sent from the device, and the output is the data saved on the server. After receiving, the data is stored in an appropriate folder for analysis.

[1393] Step 4: Video analysis

[1394] The server analyzes the received video data. The input is the video data, and the output is analysis data that identifies specific actions and objects. For example, a video analysis algorithm is used to analyze the position of bolts and the user's hand movements in each frame.

[1395] Step 5: Audio analysis

[1396] The server converts the received voice data into text and analyzes the question. The input is voice data, and the output is text data and the analysis results. Using a voice recognition system, for example, the voice data is converted into text, such as "Please tell me how to remove this bolt," and the intent is understood.

[1397] Step 6: Text Analysis

[1398] The server analyzes the text data and understands the question. The input is text data, and the output is an analysis result that indicates the intent of the question. Natural language processing technology is used to determine what information the input text is seeking.

[1399] Step 7: Data Integration

[1400] The server integrates the results of video analysis, audio analysis, and text analysis. The input is the results of each analysis, and the output is the integrated analysis data. For example, the server can combine and analyze the location information of a bolt and the content of a voice question to determine how to remove the bolt.

[1401] Step 8: Advice Generation

[1402] The server generates advice based on the integrated analysis data. The input is the integrated analysis data, and the output is specific advice. Using a generative AI model, advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool" is generated.

[1403] Step 9: Send Advice

[1404] The server sends the generated advice to the terminal. The input is the advice content, and the output is the data sent to the terminal. The advice is sent to the terminal via the Internet.

[1405] Step 10: Providing advice

[1406] The device notifies and displays the received advice to the user. The input is advice data sent from the server, and the output is text, audio, and guidance video displayed to the user. For example, it may play audio such as "Turn this bolt clockwise to remove it," encouraging the user to take specific action.

[1407] The above is the specific processing flow of the "Multimodal Factory Coach" system program.

[1408] (Application example 1)

[1409] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1410] In factory operations, there is a need to provide appropriate training and advice in real time when operating or maintaining robots. However, conventional systems lack human trainers, making it difficult to efficiently share know-how. This results in problems that cannot be dealt with quickly, leading to a decline in production efficiency. In particular, there is a need for technology that can provide effective support when an immediate response is required, such as when complex operations or abnormal sounds are generated.

[1411] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1412] In this invention, the server includes a means for analyzing video, a means for analyzing audio, and a means for analyzing text. This allows the server to analyze the video, audio, and text data captured by the user and provide advice based on the analysis results in real time using a generative AI model. Specifically, the server identifies the work content from the video data, understands the user's questions and problems from the audio and text data, and generates appropriate advice. By performing data analysis on the server side using prompt statements, the server can provide quick and accurate support to on-site workers and quickly respond to problems when they occur.

[1413] "Means for capturing images" refers to the function of acquiring visual information as digital data using a camera or sensor.

[1414] "Means for capturing audio" refers to a function for obtaining audio information as digital data using a microphone.

[1415] "Means for accepting text input" refers to the ability to input text information using a keyboard or touch screen.

[1416] "Means for analyzing captured video" refers to the function of processing and analyzing acquired video data using algorithms and AI.

[1417] The "means for analyzing captured audio" is a function for converting acquired audio data into text and analyzing the content.

[1418] "Means for analyzing input text" refers to a function that understands the character information input by the user and analyzes the context and meaning.

[1419] "Means for providing advice to the user based on the analysis results" is a function that provides the user with appropriate instructions and advice based on the analysis results of video, audio, and text.

[1420] "Means for notifying the user of the analysis results in real time" is a function that enables a quick response by immediately transmitting the analysis results to the user.

[1421] A "server" is a computer system that performs analysis and advice generation, and is often located in a cloud environment.

[1422] A "generative AI model" is an artificial intelligence model that generatively provides specific advice and answers in response to a user's questions.

[1423] A "prompt statement" is an instruction statement that gives AI instructions for analysis or generation.

[1424] This invention provides a system that provides an application called "Smart Robo Coach" that is installed on robots in factories. This system captures video, audio, and text, analyzes the data in real time, and provides appropriate advice to users.

[1425] The system has the following main components:

[1426] 1. Terminal

[1427] Camera: Built into the robot to acquire visual information.

[1428] Microphone: Captures audio information.

[1429] Text input interface: A user inputs text information by operating a keyboard or touch screen.

[1430] 2. Server

[1431] Video analysis method: The acquired video data is processed and analyzed using algorithms and AI.

[1432] Voice analysis means: Converts acquired voice data into text and analyzes the content.

[1433] Text analysis means: Understands the text information entered by the user and analyzes the context and meaning.

[1434] Advice generation method: Uses a generative AI model to generate appropriate advice based on the analysis results.

[1435] Data transmission and reception means: Data is transmitted from the terminal to the server, and the server notifies the terminal of the analysis results.

[1436] 3. Communication Network

[1437] It is the infrastructure for sending and receiving data between terminals and servers.

[1438] The specific operation is as follows.

[1439] Video, audio and text capture

[1440] 1. The user uses the robot's built-in camera to capture video of the work being done.

[1441] 2. Users can ask questions by voice about any points they do not understand or are unsure about. In noisy environments, a text input interface can also be used.

[1442] 3. The device temporarily stores the captured video data, audio data, and text data.

[1443] Data upload and analysis

[1444] 1. The device sends the captured data to the server in real time.

[1445] 2. The server analyzes the received video data and detects specific actions and parts in real time.

[1446] 3. The server converts the voice data into text and analyzes the user's question.

[1447] 4. The server analyzes the text input data and understands the question.

[1448] 5. The server integrates and analyzes the video, audio, and text data, and uses a generative AI model to generate optimal advice.

[1449] Advice generation and display

[1450] 1. The server generates advice and sends it to the device as data.

[1451] 2. The device notifies the user of the received advice and displays it to them. This advice is provided as text, audio, or in some cases as a guidance video.

[1452] The hardware used is a camera, microphone, and communication network built into the robot, and the software is video analysis, audio analysis, and generative AI models (such as GPT-3) implemented in Python and other languages.

[1453] As a specific example, consider a robot trying to remove a machine part. When an operator asks, "Please tell me how to remove this bolt," the robot's built-in camera captures an image of the bolt and the question is recorded by a microphone. This data is sent to a server, which analyzes it and generates advice such as, "Turn the bolt clockwise to remove it."

[1454] An example of a prompt for the generative AI model is as follows:

[1455] "Your task is to analyze video and audio data captured by robots in a factory and provide the best advice for a given task. For example, how to remove a bolt or the cause of an unusual noise."

[1456] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1457] Step 1:

[1458] The user captures video of the robot working using the built-in camera and asks questions by voice using the microphone. The input is video data and audio data. The output is the captured video file and audio file. Specifically, the camera captures video at a specified frame rate, and the microphone records audio.

[1459] Step 2:

[1460] The device temporarily stores captured video and audio data. The input is the captured video and audio files. The output is the video and audio files stored in the device's storage. Specifically, the video and audio files are stored in a specific directory.

[1461] Step 3:

[1462] The device sends stored video data, audio data, and text input data to the server in real time. The input is the video file, audio file, and text data. The output is the data sent to the server. Specifically, the data is uploaded to the server using the HTTPS protocol.

[1463] Step 4:

[1464] The server analyzes the received video data and detects specific actions and parts in real time. The input is a video file. The output is the analyzed video data, such as the location information of specific parts. Specifically, it uses OpenCV and machine learning models to analyze the video data frame by frame.

[1465] Step 5:

[1466] The server converts the received voice data into text and analyzes the question content. The input is an audio file. The output is text data and the analyzed question content. Specifically, the voice data is transcribed using the SpeechRecognition library and the content is analyzed using natural language processing (NLP) algorithms.

[1467] Step 6:

[1468] The server analyzes the input text data and understands the question. The input is text data. The output is the analyzed question. Specifically, it uses NLP algorithms to analyze the context and meaning of the text.

[1469] Step 7:

[1470] The server integrates and analyzes video data, audio data, and text data, and generates optimal advice using a generative AI model. The input is video data, audio data, and text data. The output is text or audio data as advice. Specifically, the server integrates the results of video analysis and text analysis, and generates advice using a generative AI model (e.g., GPT-3).

[1471] Step 8:

[1472] The server sends the generated advice to the terminal as data. The input is text or audio data as advice. The output is the advice data sent to the terminal. Specifically, the data is sent to the terminal using the HTTPS protocol.

[1473] Step 9:

[1474] The terminal notifies and displays the received advice to the user. This advice is provided as text or audio, or in some cases as a guidance video. The input is text or audio data as advice. The output is the advice notified to the user. Specifically, the text is displayed on the display and audio is played from the speaker.

[1475] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1476] ---

[1477] This invention is a system that provides efficient and intuitive training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. Furthermore, it aims to improve the user experience and training effectiveness by adding an emotion engine that recognizes the user's emotions.

[1478] System Configuration

[1479] The system of the present invention includes the following major components:

[1480] 1. Terminal: A mobile terminal such as a smartphone that provides an interface for users to operate.

[1481] Video capture method

[1482] Audio capture method

[1483] Text input method

[1484] Data transmission method

[1485] Emotion Recognition Module

[1486] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[1487] Video analysis methods

[1488] Voice analysis methods

[1489] Text Analysis Methods

[1490] Advice Generation Method

[1491] Data Receiving Method

[1492] Emotion Engine

[1493] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[1494] Specific actions

[1495] Video, audio, text and emotional capture

[1496] 1. The user takes a video of the machine operation using a device such as a smartphone.

[1497] 2. Users can ask questions by voice if they have any questions or concerns. In noisy environments, they can input text.

[1498] 3. The device temporarily stores the video data, audio data, and text data.

[1499] 4. The emotion recognition module analyzes the user's emotions in real time from the video and audio data.

[1500] Data upload and analysis

[1501] 1. The device sends the captured data to the server.

[1502] 2. The server analyzes the received video data and detects specific actions and parts in each frame.

[1503] 3. The server converts the voice data into text and analyzes the user's question.

[1504] 4. The server analyzes the text input data and understands the question.

[1505] 5. The emotion engine analyzes the received emotion data and detects the user's stress level and emotional state.

[1506] 6. The server integrates and analyzes the video, audio, text, and emotional data to generate optimal advice.

[1507] Advice generation and display

[1508] 1. The server formats the generated advice into text and speech. An emotion engine adjusts the advice content to the user's emotional state.

[1509] For example, advice like "Turn this bolt clockwise to remove it. Use the appropriate tool" becomes "Turn this bolt clockwise slowly. If you have trouble, use the appropriate tool" if the user is feeling stressed.

[1510] 2. The server sends the generated advice data to the terminal, securing an appropriate network path to transfer the data.

[1511] 3. The device displays the advice data received from the server by displaying a text message on the screen and playing the advice via audio output.

[1512] Specific examples

[1513] Example 1: Machine part replacement

[1514] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[1515] 2. The device sends video and audio data to the server.

[1516] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[1517] 4. The emotion engine analyzes the user's emotional state from video and audio.

[1518] 5. The server combines video, audio, text, and emotion data to generate advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[1519] 6. The device displays the received advice to the user and also plays it back aloud.

[1520] 7. The user replaces the part according to the advice.

[1521] Example 2: Checking for abnormal sounds

[1522] 1. A user records an abnormal machine sound on their smartphone and asks via text, "Is this sound normal?"

[1523] 2. The device sends the voice and text data to the server.

[1524] 3. The server uses voice analysis to detect abnormal sound patterns and uses text analysis to understand the content of the question.

[1525] 4. The emotion engine analyzes the user's emotional state.

[1526] 5. The server generates advice on the cause of the abnormal sound and how to deal with it, for example, by telling a nervous user, "This sound is caused by worn bearings. It would be a good idea to replace the bearings."

[1527] 6. The device displays the received advice to the user and plays it aloud.

[1528] 7. The user follows the advice to identify the cause of the abnormal sound and take action.

[1529] This concludes the description of the "Multimodal Factory Coach" system, which enables efficient and intuitive training in manufacturing workplaces and provides personalized support based on the user's emotions.

[1530] The processing flow will be explained below.

[1531] ---

[1532] In case of replacing machine parts

[1533] Step 1:

[1534] The user takes a video of the machine operation using a smartphone, and uses the device's camera function to record the specific operation points.

[1535] Step 2:

[1536] The user can ask questions by voice about things they don't understand, such as, "Please tell me how to remove this bolt."

[1537] Step 3:

[1538] Temporarily saves video and audio data captured by the device. Stores data in internal storage.

[1539] Step 4:

[1540] The device's emotion recognition module analyzes the audio and video data to determine the user's emotional state, for example, determining whether the user is nervous based on the tone of their voice or facial expression.

[1541] Step 5:

[1542] The device establishes an internet connection using Wi-Fi or a mobile data network to send video data, audio data, text data, and emotion data to the server.

[1543] Step 6:

[1544] The device uploads all data to the server, which calls the appropriate API to package and transmit the data.

[1545] Step 7:

[1546] The server analyzes the received video data and uses an image analysis algorithm to identify the location of the bolt in each video frame.

[1547] Step 8:

[1548] The server converts the voice data into text. Using voice recognition technology, the server converts the user's speech into a string of characters.

[1549] Step 9:

[1550] The server analyzes the converted text and understands the question, using natural language processing technology to analyze the intent of the question.

[1551] Step 10:

[1552] The server's emotion engine analyzes the emotion data to understand the user's stress level and emotional state.

[1553] Step 11:

[1554] The server integrates and analyzes video data, audio data, text data, and emotional data, and uses a multimodal AI model to generate optimal advice in response to the user's question.

[1555] Step 12:

[1556] The server adjusts the advice based on the user's emotional data. For example, if the user is nervous, the advice will be adjusted to something like, "Turn this bolt clockwise slowly. If you have trouble, use the appropriate tool."

[1557] Step 13:

[1558] Formatting server-generated advice into text and speech, e.g. "Turn this bolt clockwise to remove it. Use the special tool."

[1559] Step 14:

[1560] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[1561] Step 15:

[1562] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[1563] Step 16:

[1564] The user follows the displayed advice and continues the part replacement operation.

[1565] ---

[1566] Checking for abnormal machine noise

[1567] Step 1:

[1568] The user records abnormal sounds from the machine using their smartphone. The abnormal sounds are recorded using the device's microphone function.

[1569] Step 2:

[1570] The user texts in a question: "Is this sound normal?"

[1571] Step 3:

[1572] The device temporarily stores captured voice data and entered text. Stores data in internal storage.

[1573] Step 4:

[1574] The device's emotion recognition module analyzes the video and audio data to determine the user's emotional state.

[1575] Step 5:

[1576] The device establishes an internet connection using Wi-Fi or a mobile data network to send voice, text, and emotion data to the server.

[1577] Step 6:

[1578] The device uploads all data to the server, which calls the appropriate API to package and transmit the data.

[1579] Step 7:

[1580] The server analyzes the received audio data and uses an audio analysis algorithm to detect abnormal sound patterns.

[1581] Step 8:

[1582] The server analyzes the input text data and understands the question. It uses natural language processing technology to analyze the intent of the question.

[1583] Step 9:

[1584] The server's emotion engine analyzes the emotion data to understand the user's stress level and emotional state.

[1585] Step 10:

[1586] The server integrates and analyzes voice data, text data, and emotion data, and uses a multimodal AI model to generate optimal advice for the user's question.

[1587] Step 11:

[1588] The server adjusts the advice based on the user's emotional data. For example, if the user is nervous, the advice will be adjusted to say, "This noise is caused by worn bearings. It would be a good idea to replace the bearings."

[1589] Step 12:

[1590] Formatting server-generated advice into text and speech, e.g. "This sound is caused by worn bearings. We recommend replacing the bearings."

[1591] Step 13:

[1592] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[1593] Step 14:

[1594] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[1595] Step 15:

[1596] The user follows the displayed advice to identify the cause of the abnormal sound and take action.

[1597] ---

[1598] The above are the specific processing steps when combining the "Multimodal Factory Coach" system with an emotion engine.

[1599] Example 2

[1600] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1601] In modern manufacturing, a lack of trainers and a lack of sharing of know-how are serious problems. This makes it difficult for new or inexperienced employees to receive training quickly and effectively, resulting in reduced work efficiency. Furthermore, simple advice that does not take users' emotions into consideration makes it difficult to address the stress and tension they feel, preventing optimal performance. A system that can solve these problems and provide efficient and intuitive training is needed.

[1602] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1603] In this invention, the server

[1604] a means for capturing video;

[1605] a means for capturing audio;

[1606] means for accepting text input;

[1607] a means for analyzing emotions;

[1608] means for analyzing the captured video;

[1609] means for analyzing the captured audio;

[1610] means for parsing input text;

[1611] means for analyzing user emotion data;

[1612] and means for providing advice to the user based on the analysis results.

[1613] This makes it possible to provide effective and intuitive training that takes into account the stress and tension felt by the user. By analyzing the user's emotional state in real time and generating and providing appropriate advice based on that state, the user's understanding and work efficiency are improved.

[1614] "Means for capturing video" refers to devices or functions that capture the equipment operated by the user and the work environment in real time, and record and save the video data.

[1615] "Audio capturing means" refers to a device or function that collects the user's voice using a microphone or the like and records it as audio data.

[1616] "Means for accepting text input" refers to a function that allows a user to input text data using an interface such as a keyboard or touch screen.

[1617] "Means for analyzing emotions" refers to software or algorithms that use video and audio data of users to recognize and evaluate their emotional state, such as stress and tension, in real time.

[1618] "Means for analyzing captured video" refers to a function that analyzes received video data and detects and identifies specific actions, positions, parts, etc.

[1619] The "means for analyzing captured audio" is a function that converts collected audio data into text data and further analyzes the content of that text data.

[1620] "Means for analyzing input text" refers to software or algorithms that understand the content of the text data entered by the user and extract the necessary information.

[1621] The "means for analyzing the user's emotional data" is a function that analyzes the user's emotional state from the user's video data and audio data, and evaluates the stress level and tension.

[1622] The "means for providing advice to the user based on the analysis results" is a function for generating and providing optimal instructions and advice to the user based on the analysis results of video, audio, text, and emotional data.

[1623] This invention is a system for providing efficient and intuitive training, resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. The system aims to improve the effectiveness of training by recognizing the user's emotions and providing appropriate advice.

[1624] System Configuration

[1625] The system includes the following major components:

[1626] 1. Terminal: A mobile terminal operated by the user, equipped with the following functions:

[1627] Video capture method

[1628] Audio capture method

[1629] Text input method

[1630] Data transmission method

[1631] Emotion Recognition Module

[1632] 2. Server: A computer system for analysis and advice generation located in a cloud environment.

[1633] Video analysis methods

[1634] Voice analysis methods

[1635] Text Analysis Methods

[1636] Advice Generation Method

[1637] Data Receiving Method

[1638] Emotion Engine

[1639] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[1640] Specific actions

[1641] Video, audio, text and emotional capture

[1642] 1. The user uses a device such as a smartphone to record video of the machine operation. For example, while recording the operation, the user can ask a question by voice, such as "Please tell me how to remove this bolt."

[1643] 2. If the user is in a noisy environment, he or she inputs text. For example, the user inputs text such as "Is this sound normal?"

[1644] 3. The device temporarily stores the video data, audio data, and text data.

[1645] 4. The emotion recognition module analyzes the user's emotions in real time from their video and audio data, for example, determining whether they are nervous.

[1646] Sending data

[1647] 1. The device sends all captured data to the server.

[1648] 2. Use a high-speed and stable communication network for data transmission.

[1649] Data analysis on the server

[1650] 1. The server analyzes the video data received by the server using video analysis means to detect specific actions and parts in each frame. For example, it identifies the position of a bolt.

[1651] 2. The server converts the voice data into text and uses speech analysis to understand the question. For example, it analyzes the question, "Please tell me how to remove this bolt."

[1652] 3. The server analyzes the text input data and understands the question.

[1653] 4. The emotion engine analyzes the received emotion data and detects the user's stress level and emotional state. For example, it recognizes that the user is nervous.

[1654] Generating Advice

[1655] 1. The server integrates and analyzes video, audio, text, and emotion data to generate optimal advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[1656] 2. The emotion engine adjusts advice according to the user's emotional state.

[1657] Providing advice and feedback

[1658] 1. The server sends the generated advice to the device.

[1659] 2. The device displays the received advice data to the user and also plays it aloud. For example, the device might display the message "Turn this bolt clockwise slowly. Use a tool if necessary" on the screen and simultaneously play it aloud.

[1660] 3. The user follows the advice and performs the task. During this time, the device monitors the user's emotional state and sends feedback to the server as needed.

[1661] Specific examples

[1662] Example 1: Replacing machine parts

[1663] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[1664] 2. The device sends video and audio data to the server.

[1665] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[1666] 4. The emotion engine analyzes the user's emotional state from video and audio.

[1667] 5. The server combines video, audio, text, and emotion data to generate advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[1668] 6. The device displays the received advice to the user and also plays it back aloud.

[1669] 7. The user replaces the part according to the advice.

[1670] Example 2: Checking for abnormal sounds

[1671] 1. A user records an abnormal machine sound on their smartphone and asks via text, "Is this sound normal?"

[1672] 2. The device sends the voice and text data to the server.

[1673] 3. The server uses voice analysis to detect abnormal sound patterns and uses text analysis to understand the content of the question.

[1674] 4. The emotion engine analyzes the user's emotional state.

[1675] 5. The server generates advice on the cause of the abnormal sound and how to deal with it, for example, by telling a nervous user, "This sound is caused by worn bearings. It would be a good idea to replace the bearings."

[1676] 6. The device displays the received advice to the user and plays it aloud.

[1677] 7. The user follows the advice to identify the cause of the abnormal sound and take action.

[1678] Prompt Sentence Examples

[1679] "Please tell me how to replace the machine parts. I'm a little nervous."

[1680] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1681] Step 1:

[1682] The user captures video using a device. For example, the user may record video of a machine operation or part replacement using a smartphone, and then ask a question by voice, such as "Please tell me how to remove this bolt." Specifically, the user's operation of the equipment is recorded, and the audio is collected using the device's microphone.

[1683] Input: Video and audio data captured by a smartphone.

[1684] Output: The captured video and audio data is saved on the device.

[1685] Step 2:

[1686] When a user is in a noisy environment, he or she uses the device to input text. For example, he or she inputs text such as "Is this sound normal?". The input is performed using the device's keyboard or touch screen.

[1687] Input: Text data that a user types into a terminal.

[1688] Output: The entered text data is saved to the terminal.

[1689] Step 3:

[1690] The device temporarily stores video, audio, and text data and prepares it for later transmission to the server. The data is temporarily stored in the device's memory.

[1691] Input: Captured video, audio, and text data.

[1692] Output: Video, audio, and text data temporarily stored in the device's memory.

[1693] Step 4:

[1694] The emotion recognition module analyzes the user's emotions in real time from their video and audio data. For example, it analyzes the user's facial expressions and tone of voice to detect tension or stress. This emotion analysis is performed in real time and is executed on the device.

[1695] Input: Captured video and audio data.

[1696] Output: Parsed user emotion data.

[1697] Step 5:

[1698] All data captured by the device (video data, audio data, text data, and emotion data) is sent to the server. The data is transferred using a high-speed, stable communication network.

[1699] Input: Temporarily stored video, audio, and text data, as well as analyzed emotion data.

[1700] Output: Video, audio, text data and emotion data sent to the server.

[1701] Step 6:

[1702] The server analyzes the received video data. Using video analysis means, it detects specific actions and parts in each frame. For example, analysis is performed to identify the location and operation method of a bolt in the video.

[1703] Input: Video data sent to the server.

[1704] Output: Analyzed video data.

[1705] Step 7:

[1706] The server converts the received voice data into text and uses voice analysis means to understand the content of the user's question. For example, it analyzes voice data such as "Please tell me how to remove this bolt."

[1707] Input: The audio data sent to the server.

[1708] Output: The audio data converted to text and the parsed questions.

[1709] Step 8:

[1710] The server uses text analysis means to understand the content of the input text data, for example, analyzing the content of a question such as "Is this sound normal?"

[1711] Input: The text data sent to the server.

[1712] Output: Parsed text data and question content.

[1713] Step 9:

[1714] The emotion engine analyzes the received emotion data to detect the user's stress level and emotional state, for example, assessing whether the user is nervous.

[1715] Input: Emotion data sent to the server.

[1716] Output: The analyzed emotional state of the user.

[1717] Step 10:

[1718] The server combines video, audio, text, and emotion data to generate optimal advice. For example, if the user is nervous, the server generates advice such as, "Turn this bolt clockwise slowly. Use a tool if necessary."

[1719] Input: Analyzed video data, audio data, text data, and emotion data.

[1720] Output: The generated advice data.

[1721] Step 11:

[1722] The server sends the generated advice data to the terminal, securing an appropriate network path for rapid transfer.

[1723] Input: Generated advice data.

[1724] Output: Advice data sent to the terminal.

[1725] Step 12:

[1726] The device displays the received advice data to the user. It displays a text message on the screen and plays the advice aloud. For example, the device might display "Turn this bolt slowly clockwise. Use a tool if necessary" and play the advice aloud.

[1727] Input: Advice data sent by the server.

[1728] Output: The advice that was displayed and played back to the user.

[1729] The above is a flow of the specific processing steps of the system. It includes details of the data processing and calculations performed at each step, which allows for effective and intuitive training for users.

[1730] (Application example 2)

[1731] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1732] In the manufacturing industry, there are problems with efficient and intuitive training due to a shortage of trainers and insufficient sharing of know-how. Furthermore, training and advice that do not take into account the user's emotional state can reduce its effectiveness in the field.

[1733] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video, means for capturing audio, means for accepting text input, means for analyzing the user's emotional state, means for providing advice to the user based on the analysis results and the user's emotional state, means for transmitting the captured video and audio to the server, means for displaying the advice transmitted from the server on the user terminal, means for analyzing prompt sentences using a generative AI model and generating advice corresponding to the emotion, and means for integrating and analyzing the user's video data, audio data, and text data, inputting the prompt sentences into the generative AI model, and generating advice based on the analysis results. This enables efficient and intuitive training and advice provision while taking the user's emotional state into consideration.

[1734] A "video capture means" is a device or method for obtaining images or video of a user or the work environment.

[1735] An "audio capturing means" is a device or method for recording the sounds of a user or the work environment.

[1736] The "means for accepting text input" refers to a device or method that allows a user to input characters using a keyboard, an on-screen keyboard, or the like.

[1737] A "means for analyzing a user's emotional state" is a device or method for inferring a user's emotional state using audio, video, and other data.

[1738] The "means for providing advice to a user based on the analysis results and the emotional state of the user" is a device or method for combining the results of data analysis with the emotional state to provide the user with optimal advice.

[1739] The "means for transmitting captured video and audio to a server" refers to a device or method for transferring captured video and audio data to a server via a network.

[1740] The "means for displaying advice sent from the server on the user terminal" refers to a device or method for displaying instructions or advice received from the server in a form that can be confirmed by the user.

[1741] "Means for analyzing prompt sentences using a generative AI model and generating advice based on emotions" refers to a device or method that utilizes AI technology to analyze text data entered by a user and generate appropriate advice based on that data and the user's emotional state.

[1742] "Means for integrating and analyzing a user's video data, audio data, and text data, inputting prompt sentences into a generative AI model, and generating advice based on the analysis results" refers to a device or method for integrating and analyzing multiple data sources, analyzing prompt sentences using AI based on the results, and providing appropriate advice to the user.

[1743] The present invention is a system for a factory robot that provides specific advice while taking into account the emotional state of the user. The system includes the following main components:

[1744] System Configuration

[1745] 1. Terminal: This applies to robots that work in factories. Robots use cameras and microphones to capture video and audio of their work.

[1746] Video capture method

[1747] Audio capture method

[1748] Text input method

[1749] Emotion Recognition Module

[1750] Data transmission method

[1751] 2. Server: A system deployed in a cloud environment that performs analysis and generates advice. The server includes the following means:

[1752] Video analysis methods

[1753] Voice analysis methods

[1754] Text Analysis Methods

[1755] Emotion Recognition Module

[1756] Advice Generation Method

[1757] Data Receiving Method

[1758] Generative AI Models

[1759] 3. Communication network: Infrastructure for sending and receiving data between terminals and servers.

[1760] Operation overview

[1761] Data Capture and Transmission

[1762] As the user performs a task, the device (robot) captures video and audio using a camera and microphone. The device also includes a means for the user to input text using a keyboard or voice input. This data is sent to a server in real time.

[1763] Data analysis and emotion recognition

[1764] The server integrates and analyzes the received video, audio, and text data. The emotion recognition module analyzes the user's emotional state, and the generative AI model analyzes the prompt sentences to generate appropriate advice.

[1765] Advice generation and delivery

[1766] The generated advice is sent from the server to the terminal in text and audio format, and the terminal displays the received advice to the user and plays it back aloud.

[1767] Hardware and software used

[1768] Hardware: Camera, microphone, user input devices (keyboard, touch screen, etc.), robot body

[1769] Software: Video analysis software (OpenCV, etc.), audio analysis software (SpeechRecognition, etc.), emotion recognition software (EmotionRecognizer, etc.), generative AI models (cloud-based AI services, etc.)

[1770] Specific examples

[1771] Machine part replacement

[1772] If a user is unsure how to remove a bolt, they can enter "Please tell me how to remove this bolt" in text format, and the robot will send video and audio to the server. The server will analyze the data and use an emotion recognition module to confirm that the user is nervous. The advice generator will then generate advice such as "Turn this bolt slowly clockwise. Use a tool if necessary," which will then be displayed and played aloud on the device.

[1773] Check for abnormal sounds

[1774] If a user wants to check for abnormal machine sounds, they record the sound on their smartphone and then type "Is this sound normal?" into the robot while playing it back. The robot sends this to the server, which then analyzes the sound pattern using a voice analysis module and detects tension using an emotion recognition module. The server then generates advice such as "This sound is caused by worn bearings. It would be a good idea to replace the bearings," which the robot displays and plays back in audio.

[1775] Prompt Sentence Examples

[1776] Can you tell me how to remove this bolt?

[1777] "Is this sound normal?"

[1778] Through these specific operations, user support is realized through cooperation between the terminal and the server.

[1779] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1780] Step 1:

[1781] The device uses a camera and microphone to capture video and audio of the user working. As input, video data is obtained from the camera and audio data is obtained from the microphone. Specifically, the device records the user's work scene in real time and temporarily stores the data.

[1782] Step 2:

[1783] The user enters a question requesting clarification through text input. As input, the text data entered by the user via a keyboard or touch screen is acquired. Specifically, the device provides an interface for accepting the user's input and saves the entered text.

[1784] Step 3:

[1785] The device sends the captured video, audio, and text data to the server. The device uses the saved video, audio, and text data as input and sends this data to the server as output. Specifically, the device uploads the data to the cloud server using the network.

[1786] Step 4:

[1787] The server analyzes the video data and identifies the user's work. It uses the video data sent from the device as input and obtains the analysis results as output. Specifically, the server uses video analysis software to recognize specific actions and parts in each frame and compares them with a database.

[1788] Step 5:

[1789] The server analyzes the voice data and understands the user's question. It uses the voice data sent from the device as input and outputs the converted text string and the analysis results. Specifically, it uses voice recognition software to convert the voice into text and analyzes the content of the question.

[1790] Step 6:

[1791] The server analyzes the user's emotional state. It uses video and audio data as input and obtains the emotion analysis results as output. Specifically, it uses an emotion recognition module to estimate the user's emotional state from facial expressions, tone of voice, etc.

[1792] Step 7:

[1793] The server integrates the analysis results and the emotional state and generates advice using a generative AI model. The analysis results, emotion analysis results, and prompt sentences are used as input, and advice for the user is obtained as output. Specifically, the generative AI model analyzes the prompt sentence and generates appropriate advice based on the data.

[1794] Step 8:

[1795] The server sends the generated advice to the terminal. The generated advice data is used as input to obtain data to be sent to the terminal as output. In concrete terms, the server uploads the advice data to the terminal via the network.

[1796] Step 9:

[1797] The device displays the advice received from the server to the user and plays it back aloud. It uses the advice data sent from the server as input and provides visual and auditory feedback to the user as output. Specifically, the device displays a text message on the screen and plays back an audio message from the speaker.

[1798] Prompt Sentence Examples

[1799] Can you tell me how to remove this bolt?

[1800] "Is this sound normal?"

[1801] The above are the specific processing steps of this system.

[1802] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1803] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1804] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1805] [Fourth embodiment]

[1806] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1807] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1808] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1809] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1810] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1811] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1812] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1813] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1814] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1815] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1816] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1817] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1818] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1819] ---

[1820] The purpose of this invention is to provide efficient training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. Specifically, it analyzes multimodal data from video, audio, and text input, and provides intuitive and appropriate advice to users.

[1821] System Configuration

[1822] The system of the present invention includes the following major components:

[1823] 1. Terminal: A mobile terminal such as a smartphone that provides an interface for users to operate.

[1824] Video capture method

[1825] Audio capture method

[1826] Text input method

[1827] Data transmission method

[1828] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[1829] Video analysis methods

[1830] Voice analysis methods

[1831] Text Analysis Methods

[1832] Advice Generation Method

[1833] Data Receiving Method

[1834] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[1835] Specific actions

[1836] Video, audio and text capture

[1837] 1. The user takes a photo of the machine operation using a device such as a smartphone.

[1838] 2. Users can ask questions by voice about points they don't understand or points they're unsure about. In some cases, they can input text in a noisy environment.

[1839] 3. The device temporarily stores the video data, audio data, and text data.

[1840] Data upload and analysis

[1841] 1. The device sends the captured data to the server.

[1842] 2. The server analyzes the received video data and detects specific actions and parts in each frame.

[1843] 3. The server converts the voice data into text and analyzes the user's question.

[1844] 4. The server analyzes the text input data and understands the question.

[1845] 5. The server integrates and analyzes the video, audio, and text data to generate optimal advice.

[1846] Advice generation and display

[1847] 1. The server generates advice and sends it to the device as data.

[1848] 2. The device notifies the user of the received advice and displays it to them. This advice is provided as text, audio, or in some cases as a guidance video.

[1849] Specific examples

[1850] Example 1: Machine part replacement

[1851] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[1852] 2. The device sends video and audio data to the server.

[1853] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[1854] 4. The server generates advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool."

[1855] 5. The device displays the received advice to the user and also plays it back aloud.

[1856] 6. The user replaces the part following the appropriate procedure.

[1857] Example 2: Checking for abnormal sounds

[1858] 1. A user records an abnormal sound made by a machine on their smartphone and asks by text, "Is this sound normal?"

[1859] 2. The device sends video and audio data to the server.

[1860] 3. The server uses voice analysis to detect abnormal sounds and analyzes the text input to understand the question.

[1861] 4. The server generates the advice, "This noise is caused by worn bearings. We recommend replacing the bearings."

[1862] 5. The device displays the received advice to the user as text and plays it back aloud.

[1863] 6. The user follows the advice, checks the cause of the abnormal sound, and takes action.

[1864] This concludes the description of the "Multimodal Factory Coach" system, which enables efficient and intuitive training in manufacturing workplaces.

[1865] The processing flow will be explained below.

[1866] ---

[1867] In case of replacing machine parts

[1868] Step 1:

[1869] The user takes a video of the machine operation using a smartphone, and uses the device's camera function to record the specific operation points.

[1870] Step 2:

[1871] The user can ask questions by voice about things they don't understand, such as, "Please tell me how to remove this bolt."

[1872] Step 3:

[1873] Temporarily saves video and audio data captured by the device. Stores data in internal storage.

[1874] Step 4:

[1875] Your device establishes an internet connection, using Wi-Fi or your mobile data network, to send the stored data to the server.

[1876] Step 5:

[1877] The device uploads video and audio data to the server, which then calls the appropriate API to package and transmit the data.

[1878] Step 6:

[1879] The server analyzes the received video data and uses an image analysis algorithm to identify the location of the bolt in each video frame.

[1880] Step 7:

[1881] The server converts the voice data into text. Using voice recognition technology, the server converts the user's speech into a string of characters.

[1882] Step 8:

[1883] The server analyzes the converted text and understands the question, using natural language processing technology to analyze the intent of the question.

[1884] Step 9:

[1885] The server integrates and analyzes the video and text data, and uses a multimodal AI model to generate optimal advice in response to the user's question.

[1886] Step 10:

[1887] Formatting server-generated advice into text and speech, e.g. "Turn this bolt clockwise to remove it. Use the special tool."

[1888] Step 11:

[1889] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[1890] Step 12:

[1891] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[1892] Step 13:

[1893] The user follows the displayed advice and continues the part replacement operation.

[1894] ---

[1895] Checking for abnormal machine noise

[1896] Step 1:

[1897] The user records abnormal sounds from the machine using their smartphone. The abnormal sounds are recorded using the device's microphone function.

[1898] Step 2:

[1899] The user texts in a question: "Is this sound normal?"

[1900] Step 3:

[1901] The device temporarily stores captured voice data and entered text. Stores data in internal storage.

[1902] Step 4:

[1903] Your device establishes an internet connection, using Wi-Fi or your mobile data network, to send the stored data to the server.

[1904] Step 5:

[1905] The device uploads the voice and text data to the server, which calls the appropriate API to package and transmit the data.

[1906] Step 6:

[1907] The server analyzes the received audio data and uses an audio analysis algorithm to detect abnormal sound patterns.

[1908] Step 7:

[1909] The server analyzes the input text data and understands the question. It uses natural language processing technology to analyze the intent of the question.

[1910] Step 8:

[1911] The server integrates and analyzes the voice and text data, and uses a multimodal AI model to generate optimal advice for the user's question.

[1912] Step 9:

[1913] Formatting server-generated advice into text and speech, e.g. "This sound is caused by worn bearings. We recommend replacing the bearings."

[1914] Step 10:

[1915] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[1916] Step 11:

[1917] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[1918] Step 12:

[1919] The user follows the displayed advice to identify the cause of the abnormal sound and take action.

[1920] ---

[1921] The above are the specific processing steps of the "Multimodal Factory Coach" system.

[1922] Example 1

[1923] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1924] To resolve the shortage of trainers and lack of know-how sharing in the manufacturing industry, provide efficient and intuitive training, and enable users to receive appropriate advice in real time. There is a demand for time-saving and effortless responses to complex machine operations and abnormality detection on manufacturing sites. Conventional methods lack the mechanisms for integrating and analyzing multimodal data such as video, audio, and text to provide appropriate advice.

[1925] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1926] In this invention, the server includes means for acquiring video, means for acquiring audio, means for accepting text, means for analyzing the acquired video, means for analyzing the acquired audio, means for analyzing input text, means for providing advice to the user based on the analysis results, and means for integrating and analyzing the acquired data, thereby making it possible to perform an integrated analysis of multimodal data and provide appropriate advice to the user in real time.

[1927] The "means for acquiring images" is a function for capturing visual information of the surroundings using a sensor device such as a camera mounted on the equipment operated by the user.

[1928] The "means for acquiring audio" is a function for collecting surrounding audio information using a microphone or audio capture device installed in the device used by the user.

[1929] "Means for accepting text" refers to a function that receives and records text information entered by a user using a keyboard or touch screen.

[1930] "Means for analyzing the captured video" refers to software algorithms or machine learning models that process the captured video data and recognize specific actions or objects.

[1931] The "means for analyzing the captured voice" refers to a voice recognition system or natural language processing technology that converts the captured voice data into text and further understands its content.

[1932] The "means for analyzing input text" is a natural language processing algorithm that processes the character data entered by the user and understands its content.

[1933] The "means for providing advice to the user based on the analysis results" is a system that integrates the analyzed video, audio, and text data, generates appropriate advice, and notifies or displays that advice to the user.

[1934] "Means for integrating and analyzing acquired data" refers to data fusion technology and analysis algorithms that process video, audio, and text data in an integrated manner and make comprehensive judgments.

[1935] "Computer system" refers collectively to hardware and software for receiving, processing, analyzing data, and generating advice.

[1936] The "user interface" refers to an interface such as a display device or audio output device for directly providing the generated advice to the user.

[1937] The "Multimodal Factory Coach" system is a system that provides efficient training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. This system analyzes multimodal data from video, audio, and text input, and provides intuitive and appropriate advice to users.

[1938] System Configuration

[1939] The system of the present invention includes the following major components:

[1940] 1. Terminal: A mobile terminal that provides an interface for users to operate. Specifically, a smartphone or tablet is used.

[1941] Means of acquiring images (camera)

[1942] A means of capturing audio (microphone)

[1943] A means of accepting text (keyboard or touchscreen)

[1944] Data transmission method (Internet data upload function)

[1945] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[1946] A method for analyzing video (video analysis algorithm)

[1947] A means of analyzing the voice (voice recognition technology, e.g., Google Cloud Speech-to-Text API)

[1948] Means of analyzing text (natural language processing technology)

[1949] A means of generating advice (generative AI model)

[1950] Data reception method (data download function via the Internet)

[1951] 3. Communication network: The infrastructure for sending and receiving data between devices and servers, usually using an internet connection.

[1952] Specific actions

[1953] Data Capture

[1954] A user uses a device such as a smartphone to record video of a machine operation. For example, a user may record a video of a machine part replacement and ask a question by voice, such as, "Please tell me how to remove this bolt." In a noisy environment, text input is also possible, and the user may input, "Is this sound normal?" The device temporarily stores this video, audio, and text data.

[1955] Data transmission

[1956] The device sends the captured data to a server, using an internet connection and uploading the data securely.

[1957] Data reception and analysis

[1958] The server receives and analyzes the data sent from the terminal. It uses video analysis means to detect specific actions and components in each frame. It uses audio analysis means to convert the audio data into text and analyze the question. It also uses input text analysis means to understand the question.

[1959] Advice Generation

[1960] The server integrates the analyzed data and generates optimal advice. For example, it generates advice on specific steps and tool selection, such as "Turn this bolt clockwise to remove it. Use the appropriate tool." Using a generative AI model, it is possible to generate the optimal advice that the user desires.

[1961] Providing advice

[1962] The server sends the generated advice to the terminal, which receives it and notifies and displays it to the user. The advice is provided as text, audio, or in some cases, guidance video. The user can follow the instructions and perform the appropriate task.

[1963] Specific examples

[1964] Example 1: Machine part replacement

[1965] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[1966] 2. The device sends video and audio data to the server.

[1967] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[1968] 4. The server generates advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool."

[1969] 5. The device displays the received advice to the user and also plays it back aloud.

[1970] 6. The user replaces the part following the appropriate procedure.

[1971] Example 2: Checking for abnormal sounds

[1972] 1. A user records an abnormal sound made by a machine on their smartphone and asks by text, "Is this sound normal?"

[1973] 2. The device sends the voice data and text data to the server.

[1974] 3. The server uses voice analysis to detect abnormal sounds and analyzes the text input to understand the question.

[1975] 4. The server generates the advice, "This noise is caused by worn bearings. We recommend replacing the bearings."

[1976] 5. The device displays the received advice to the user as text and plays it back aloud.

[1977] 6. The user follows the advice, checks the cause of the abnormal sound, and takes action.

[1978] In this way, the system enables real-time support for users, supporting rapid response and efficient work on-site.

[1979] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1980] Step 1: Data Capture

[1981] A user uses a smartphone to record video of a machine being operated and ask questions by voice. The input is video data and audio data, and the output is data temporarily stored on the device. For example, if a user records a machine part replacement and asks, "Please tell me how to remove this bolt," this data is stored on the device.

[1982] Step 2: Send data

[1983] The device sends the stored video and audio data to the server. The input is the data stored on the device, and the output is the data uploaded to the server. The data is sent via the Internet and received by the server.

[1984] Step 3: Receiving data

[1985] The server receives the data sent from the device. The input is the video and audio data sent from the device, and the output is the data saved on the server. After receiving, the data is stored in an appropriate folder for analysis.

[1986] Step 4: Video analysis

[1987] The server analyzes the received video data. The input is the video data, and the output is analysis data that identifies specific actions and objects. For example, a video analysis algorithm is used to analyze the position of bolts and the user's hand movements in each frame.

[1988] Step 5: Audio analysis

[1989] The server converts the received voice data into text and analyzes the question. The input is voice data, and the output is text data and the analysis results. Using a voice recognition system, for example, the voice data is converted into text, such as "Please tell me how to remove this bolt," and the intent is understood.

[1990] Step 6: Text Analysis

[1991] The server analyzes the text data and understands the question. The input is text data, and the output is an analysis result that indicates the intent of the question. Natural language processing technology is used to determine what information the input text is seeking.

[1992] Step 7: Data Integration

[1993] The server integrates the results of video analysis, audio analysis, and text analysis. The input is the results of each analysis, and the output is the integrated analysis data. For example, the server can combine and analyze the location information of a bolt and the content of a voice question to determine how to remove the bolt.

[1994] Step 8: Advice Generation

[1995] The server generates advice based on the integrated analysis data. The input is the integrated analysis data, and the output is specific advice. Using a generative AI model, advice such as "Turn this bolt clockwise to remove it. Use the appropriate tool" is generated.

[1996] Step 9: Send Advice

[1997] The server sends the generated advice to the terminal. The input is the advice content, and the output is the data sent to the terminal. The advice is sent to the terminal via the Internet.

[1998] Step 10: Providing advice

[1999] The device notifies and displays the received advice to the user. The input is advice data sent from the server, and the output is text, audio, and guidance video displayed to the user. For example, it may play audio such as "Turn this bolt clockwise to remove it," encouraging the user to take specific action.

[2000] The above is the specific processing flow of the "Multimodal Factory Coach" system program.

[2001] (Application example 1)

[2002] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2003] In factory operations, there is a need to provide appropriate training and advice in real time when operating or maintaining robots. However, conventional systems lack human trainers, making it difficult to efficiently share know-how. This results in problems that cannot be dealt with quickly, leading to a decline in production efficiency. In particular, there is a need for technology that can provide effective support when an immediate response is required, such as when complex operations or abnormal sounds are generated.

[2004] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2005] In this invention, the server includes a means for analyzing video, a means for analyzing audio, and a means for analyzing text. This allows the server to analyze the video, audio, and text data captured by the user and provide advice based on the analysis results in real time using a generative AI model. Specifically, the server identifies the work content from the video data, understands the user's questions and problems from the audio and text data, and generates appropriate advice. By performing data analysis on the server side using prompt statements, the server can provide quick and accurate support to on-site workers and quickly respond to problems when they occur.

[2006] "Means for capturing images" refers to the function of acquiring visual information as digital data using a camera or sensor.

[2007] "Means for capturing audio" refers to a function for obtaining audio information as digital data using a microphone.

[2008] "Means for accepting text input" refers to the ability to input text information using a keyboard or touch screen.

[2009] "Means for analyzing captured video" refers to the function of processing and analyzing acquired video data using algorithms and AI.

[2010] The "means for analyzing captured audio" is a function for converting acquired audio data into text and analyzing the content.

[2011] "Means for analyzing input text" refers to a function that understands the character information input by the user and analyzes the context and meaning.

[2012] "Means for providing advice to the user based on the analysis results" is a function that provides the user with appropriate instructions and advice based on the analysis results of video, audio, and text.

[2013] "Means for notifying the user of the analysis results in real time" is a function that enables a quick response by immediately transmitting the analysis results to the user.

[2014] A "server" is a computer system that performs analysis and advice generation, and is often located in a cloud environment.

[2015] A "generative AI model" is an artificial intelligence model that generatively provides specific advice and answers in response to a user's questions.

[2016] A "prompt statement" is an instruction statement that gives AI instructions for analysis or generation.

[2017] This invention provides a system that provides an application called "Smart Robo Coach" that is installed on robots in factories. This system captures video, audio, and text, analyzes the data in real time, and provides appropriate advice to users.

[2018] The system has the following main components:

[2019] 1. Terminal

[2020] Camera: Built into the robot to acquire visual information.

[2021] Microphone: Captures audio information.

[2022] Text input interface: A user inputs text information by operating a keyboard or touch screen.

[2023] 2. Server

[2024] Video analysis method: The acquired video data is processed and analyzed using algorithms and AI.

[2025] Voice analysis means: Converts acquired voice data into text and analyzes the content.

[2026] Text analysis means: Understands the text information entered by the user and analyzes the context and meaning.

[2027] Advice generation method: Uses a generative AI model to generate appropriate advice based on the analysis results.

[2028] Data transmission and reception means: Data is transmitted from the terminal to the server, and the server notifies the terminal of the analysis results.

[2029] 3. Communication Network

[2030] It is the infrastructure for sending and receiving data between terminals and servers.

[2031] The specific operation is as follows.

[2032] Video, audio and text capture

[2033] 1. The user uses the robot's built-in camera to capture video of the work being done.

[2034] 2. Users can ask questions by voice about any points they do not understand or are unsure about. In noisy environments, a text input interface can also be used.

[2035] 3. The device temporarily stores the captured video data, audio data, and text data.

[2036] Data upload and analysis

[2037] 1. The device sends the captured data to the server in real time.

[2038] 2. The server analyzes the received video data and detects specific actions and parts in real time.

[2039] 3. The server converts the voice data into text and analyzes the user's question.

[2040] 4. The server analyzes the text input data and understands the question.

[2041] 5. The server integrates and analyzes the video, audio, and text data, and uses a generative AI model to generate optimal advice.

[2042] Advice generation and display

[2043] 1. The server generates advice and sends it to the device as data.

[2044] 2. The device notifies the user of the received advice and displays it to them. This advice is provided as text, audio, or in some cases as a guidance video.

[2045] The hardware used is a camera, microphone, and communication network built into the robot, and the software is video analysis, audio analysis, and generative AI models (such as GPT-3) implemented in Python and other languages.

[2046] As a specific example, consider a robot trying to remove a machine part. When an operator asks, "Please tell me how to remove this bolt," the robot's built-in camera captures an image of the bolt and the question is recorded by a microphone. This data is sent to a server, which analyzes it and generates advice such as, "Turn the bolt clockwise to remove it."

[2047] An example of a prompt for the generative AI model is as follows:

[2048] "Your task is to analyze video and audio data captured by robots in a factory and provide the best advice for a given task. For example, how to remove a bolt or the cause of an unusual noise."

[2049] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2050] Step 1:

[2051] The user captures video of the robot working using the built-in camera and asks questions by voice using the microphone. The input is video data and audio data. The output is the captured video file and audio file. Specifically, the camera captures video at a specified frame rate, and the microphone records audio.

[2052] Step 2:

[2053] The device temporarily stores captured video and audio data. The input is the captured video and audio files. The output is the video and audio files stored in the device's storage. Specifically, the video and audio files are stored in a specific directory.

[2054] Step 3:

[2055] The device sends stored video data, audio data, and text input data to the server in real time. The input is the video file, audio file, and text data. The output is the data sent to the server. Specifically, the data is uploaded to the server using the HTTPS protocol.

[2056] Step 4:

[2057] The server analyzes the received video data and detects specific actions and parts in real time. The input is a video file. The output is the analyzed video data, such as the location information of specific parts. Specifically, it uses OpenCV and machine learning models to analyze the video data frame by frame.

[2058] Step 5:

[2059] The server converts the received voice data into text and analyzes the question content. The input is an audio file. The output is text data and the analyzed question content. Specifically, the voice data is transcribed using the SpeechRecognition library and the content is analyzed using natural language processing (NLP) algorithms.

[2060] Step 6:

[2061] The server analyzes the input text data and understands the question. The input is text data. The output is the analyzed question. Specifically, it uses NLP algorithms to analyze the context and meaning of the text.

[2062] Step 7:

[2063] The server integrates and analyzes video data, audio data, and text data, and generates optimal advice using a generative AI model. The input is video data, audio data, and text data. The output is text or audio data as advice. Specifically, the server integrates the results of video analysis and text analysis, and generates advice using a generative AI model (e.g., GPT-3).

[2064] Step 8:

[2065] The server sends the generated advice to the terminal as data. The input is text or audio data as advice. The output is the advice data sent to the terminal. Specifically, the data is sent to the terminal using the HTTPS protocol.

[2066] Step 9:

[2067] The terminal notifies and displays the received advice to the user. This advice is provided as text or audio, or in some cases as a guidance video. The input is text or audio data as advice. The output is the advice notified to the user. Specifically, the text is displayed on the display and audio is played from the speaker.

[2068] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2069] ---

[2070] This invention is a system that provides efficient and intuitive training by resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. Furthermore, it aims to improve the user experience and training effectiveness by adding an emotion engine that recognizes the user's emotions.

[2071] System Configuration

[2072] The system of the present invention includes the following major components:

[2073] 1. Terminal: A mobile terminal such as a smartphone that provides an interface for users to operate.

[2074] Video capture method

[2075] Audio capture method

[2076] Text input method

[2077] Data transmission method

[2078] Emotion Recognition Module

[2079] 2. Server: A computer system for analysis and advice generation deployed in a cloud environment.

[2080] Video analysis methods

[2081] Voice analysis methods

[2082] Text Analysis Methods

[2083] Advice Generation Method

[2084] Data Receiving Method

[2085] Emotion Engine

[2086] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[2087] Specific actions

[2088] Video, audio, text and emotional capture

[2089] 1. The user takes a video of the machine operation using a device such as a smartphone.

[2090] 2. Users can ask questions by voice if they have any questions or concerns. In noisy environments, they can input text.

[2091] 3. The device temporarily stores the video data, audio data, and text data.

[2092] 4. The emotion recognition module analyzes the user's emotions in real time from the video and audio data.

[2093] Data upload and analysis

[2094] 1. The device sends the captured data to the server.

[2095] 2. The server analyzes the received video data and detects specific actions and parts in each frame.

[2096] 3. The server converts the voice data into text and analyzes the user's question.

[2097] 4. The server analyzes the text input data and understands the question.

[2098] 5. The emotion engine analyzes the received emotion data and detects the user's stress level and emotional state.

[2099] 6. The server integrates and analyzes the video, audio, text, and emotional data to generate optimal advice.

[2100] Advice generation and display

[2101] 1. The server formats the generated advice into text and speech. An emotion engine adjusts the advice content to the user's emotional state.

[2102] For example, advice like "Turn this bolt clockwise to remove it. Use the appropriate tool" becomes "Turn this bolt clockwise slowly. If you have trouble, use the appropriate tool" if the user is feeling stressed.

[2103] 2. The server sends the generated advice data to the terminal, securing an appropriate network path to transfer the data.

[2104] 3. The device displays the advice data received from the server by displaying a text message on the screen and playing the advice via audio output.

[2105] Specific examples

[2106] Example 1: Machine part replacement

[2107] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[2108] 2. The device sends video and audio data to the server.

[2109] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[2110] 4. The emotion engine analyzes the user's emotional state from video and audio.

[2111] 5. The server combines video, audio, text, and emotion data to generate advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[2112] 6. The device displays the received advice to the user and also plays it back aloud.

[2113] 7. The user replaces the part according to the advice.

[2114] Example 2: Checking for abnormal sounds

[2115] 1. A user records an abnormal machine sound on their smartphone and asks via text, "Is this sound normal?"

[2116] 2. The device sends the voice and text data to the server.

[2117] 3. The server uses voice analysis to detect abnormal sound patterns and uses text analysis to understand the content of the question.

[2118] 4. The emotion engine analyzes the user's emotional state.

[2119] 5. The server generates advice on the cause of the abnormal sound and how to deal with it, for example, by telling a nervous user, "This sound is caused by worn bearings. It would be a good idea to replace the bearings."

[2120] 6. The device displays the received advice to the user and plays it aloud.

[2121] 7. The user follows the advice to identify the cause of the abnormal sound and take action.

[2122] This concludes the description of the "Multimodal Factory Coach" system, which enables efficient and intuitive training in manufacturing workplaces and provides personalized support based on the user's emotions.

[2123] The processing flow will be explained below.

[2124] ---

[2125] In case of replacing machine parts

[2126] Step 1:

[2127] The user takes a video of the machine operation using a smartphone, and uses the device's camera function to record the specific operation points.

[2128] Step 2:

[2129] The user can ask questions by voice about things they don't understand, such as, "Please tell me how to remove this bolt."

[2130] Step 3:

[2131] Temporarily saves video and audio data captured by the device. Stores data in internal storage.

[2132] Step 4:

[2133] The device's emotion recognition module analyzes the audio and video data to determine the user's emotional state, for example, determining whether the user is nervous based on the tone of their voice or facial expression.

[2134] Step 5:

[2135] The device establishes an internet connection using Wi-Fi or a mobile data network to send video data, audio data, text data, and emotion data to the server.

[2136] Step 6:

[2137] The device uploads all data to the server, which calls the appropriate API to package and transmit the data.

[2138] Step 7:

[2139] The server analyzes the received video data and uses an image analysis algorithm to identify the location of the bolt in each video frame.

[2140] Step 8:

[2141] The server converts the voice data into text. Using voice recognition technology, the server converts the user's speech into a string of characters.

[2142] Step 9:

[2143] The server analyzes the converted text and understands the question, using natural language processing technology to analyze the intent of the question.

[2144] Step 10:

[2145] The server's emotion engine analyzes the emotion data to understand the user's stress level and emotional state.

[2146] Step 11:

[2147] The server integrates and analyzes video data, audio data, text data, and emotional data, and uses a multimodal AI model to generate optimal advice in response to the user's question.

[2148] Step 12:

[2149] The server adjusts the advice based on the user's emotional data. For example, if the user is nervous, the advice will be adjusted to something like, "Turn this bolt clockwise slowly. If you have trouble, use the appropriate tool."

[2150] Step 13:

[2151] Formatting server-generated advice into text and speech, e.g. "Turn this bolt clockwise to remove it. Use the special tool."

[2152] Step 14:

[2153] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[2154] Step 15:

[2155] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[2156] Step 16:

[2157] The user follows the displayed advice and continues the part replacement operation.

[2158] ---

[2159] Checking for abnormal machine noise

[2160] Step 1:

[2161] The user records abnormal sounds from the machine using their smartphone. The abnormal sounds are recorded using the device's microphone function.

[2162] Step 2:

[2163] The user texts in a question: "Is this sound normal?"

[2164] Step 3:

[2165] The device temporarily stores captured voice data and entered text. Stores data in internal storage.

[2166] Step 4:

[2167] The device's emotion recognition module analyzes the video and audio data to determine the user's emotional state.

[2168] Step 5:

[2169] The device establishes an internet connection using Wi-Fi or a mobile data network to send voice, text, and emotion data to the server.

[2170] Step 6:

[2171] The device uploads all data to the server, which calls the appropriate API to package and transmit the data.

[2172] Step 7:

[2173] The server analyzes the received audio data and uses an audio analysis algorithm to detect abnormal sound patterns.

[2174] Step 8:

[2175] The server analyzes the input text data and understands the question. It uses natural language processing technology to analyze the intent of the question.

[2176] Step 9:

[2177] The server's emotion engine analyzes the emotion data to understand the user's stress level and emotional state.

[2178] Step 10:

[2179] The server integrates and analyzes voice data, text data, and emotion data, and uses a multimodal AI model to generate optimal advice for the user's question.

[2180] Step 11:

[2181] The server adjusts the advice based on the user's emotional data. For example, if the user is nervous, the advice will be adjusted to say, "This noise is caused by worn bearings. It would be a good idea to replace the bearings."

[2182] Step 12:

[2183] Formatting server-generated advice into text and speech, e.g. "This sound is caused by worn bearings. We recommend replacing the bearings."

[2184] Step 13:

[2185] The server sends the generated advice data to the terminal, and secures an appropriate network path to transfer the data.

[2186] Step 14:

[2187] The device displays the advice data received from the server, displaying text messages on the screen and playing advice via audio output.

[2188] Step 15:

[2189] The user follows the displayed advice to identify the cause of the abnormal sound and take action.

[2190] ---

[2191] The above are the specific processing steps when combining the "Multimodal Factory Coach" system with an emotion engine.

[2192] Example 2

[2193] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2194] In modern manufacturing, a lack of trainers and a lack of sharing of know-how are serious problems. This makes it difficult for new or inexperienced employees to receive training quickly and effectively, resulting in reduced work efficiency. Furthermore, simple advice that does not take users' emotions into consideration makes it difficult to address the stress and tension they feel, preventing optimal performance. A system that can solve these problems and provide efficient and intuitive training is needed.

[2195] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2196] In this invention, the server

[2197] a means for capturing video;

[2198] a means for capturing audio;

[2199] means for accepting text input;

[2200] a means for analyzing emotions;

[2201] means for analyzing the captured video;

[2202] means for analyzing the captured audio;

[2203] means for parsing input text;

[2204] means for analyzing user emotion data;

[2205] and means for providing advice to the user based on the analysis results.

[2206] This makes it possible to provide effective and intuitive training that takes into account the stress and tension felt by the user. By analyzing the user's emotional state in real time and generating and providing appropriate advice based on that state, the user's understanding and work efficiency are improved.

[2207] "Means for capturing video" refers to devices or functions that capture the equipment operated by the user and the work environment in real time, and record and save the video data.

[2208] "Audio capturing means" refers to a device or function that collects the user's voice using a microphone or the like and records it as audio data.

[2209] "Means for accepting text input" refers to a function that allows a user to input text data using an interface such as a keyboard or touch screen.

[2210] "Means for analyzing emotions" refers to software or algorithms that use video and audio data of users to recognize and evaluate their emotional state, such as stress and tension, in real time.

[2211] "Means for analyzing captured video" refers to a function that analyzes received video data and detects and identifies specific actions, positions, parts, etc.

[2212] The "means for analyzing captured audio" is a function that converts collected audio data into text data and further analyzes the content of that text data.

[2213] "Means for analyzing input text" refers to software or algorithms that understand the content of the text data entered by the user and extract the necessary information.

[2214] The "means for analyzing the user's emotional data" is a function that analyzes the user's emotional state from the user's video data and audio data, and evaluates the stress level and tension.

[2215] The "means for providing advice to the user based on the analysis results" is a function for generating and providing optimal instructions and advice to the user based on the analysis results of video, audio, text, and emotional data.

[2216] This invention is a system for providing efficient and intuitive training, resolving the lack of trainers and the lack of sharing of know-how in the manufacturing industry. The system aims to improve the effectiveness of training by recognizing the user's emotions and providing appropriate advice.

[2217] System Configuration

[2218] The system includes the following major components:

[2219] 1. Terminal: A mobile terminal operated by the user, equipped with the following functions:

[2220] Video capture method

[2221] Audio capture method

[2222] Text input method

[2223] Data transmission method

[2224] Emotion Recognition Module

[2225] 2. Server: A computer system for analysis and advice generation located in a cloud environment.

[2226] Video analysis methods

[2227] Voice analysis methods

[2228] Text Analysis Methods

[2229] Advice Generation Method

[2230] Data Receiving Method

[2231] Emotion Engine

[2232] 3. Communication network: The infrastructure for sending and receiving data between terminals and servers.

[2233] Specific actions

[2234] Video, audio, text and emotional capture

[2235] 1. The user uses a device such as a smartphone to record video of the machine operation. For example, while recording the operation, the user can ask a question by voice, such as "Please tell me how to remove this bolt."

[2236] 2. If the user is in a noisy environment, he or she inputs text. For example, the user inputs text such as "Is this sound normal?"

[2237] 3. The device temporarily stores the video data, audio data, and text data.

[2238] 4. The emotion recognition module analyzes the user's emotions in real time from their video and audio data, for example, determining whether they are nervous.

[2239] Sending data

[2240] 1. The device sends all captured data to the server.

[2241] 2. Use a high-speed and stable communication network for data transmission.

[2242] Data analysis on the server

[2243] 1. The server analyzes the video data received by the server using video analysis means to detect specific actions and parts in each frame. For example, it identifies the position of a bolt.

[2244] 2. The server converts the voice data into text and uses speech analysis to understand the question. For example, it analyzes the question, "Please tell me how to remove this bolt."

[2245] 3. The server analyzes the text input data and understands the question.

[2246] 4. The emotion engine analyzes the received emotion data and detects the user's stress level and emotional state. For example, it recognizes that the user is nervous.

[2247] Generating Advice

[2248] 1. The server integrates and analyzes video, audio, text, and emotion data to generate optimal advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[2249] 2. The emotion engine adjusts advice according to the user's emotional state.

[2250] Providing advice and feedback

[2251] 1. The server sends the generated advice to the device.

[2252] 2. The device displays the received advice data to the user and also plays it aloud. For example, the device might display the message "Turn this bolt clockwise slowly. Use a tool if necessary" on the screen and simultaneously play it aloud.

[2253] 3. The user follows the advice and performs the task. During this time, the device monitors the user's emotional state and sends feedback to the server as needed.

[2254] Specific examples

[2255] Example 1: Replacing machine parts

[2256] 1. The user takes a video of the machine part replacement using their smartphone and asks aloud, "Please tell me how to remove this bolt."

[2257] 2. The device sends video and audio data to the server.

[2258] 3. The server uses video analysis to identify the location of the bolt and uses audio analysis to understand the question.

[2259] 4. The emotion engine analyzes the user's emotional state from video and audio.

[2260] 5. The server combines video, audio, text, and emotion data to generate advice. For example, if the user is nervous, the server might advise, "Turn this bolt clockwise slowly. Use a tool if necessary."

[2261] 6. The device displays the received advice to the user and also plays it back aloud.

[2262] 7. The user replaces the part according to the advice.

[2263] Example 2: Checking for abnormal sounds

[2264] 1. A user records an abnormal machine sound on their smartphone and asks via text, "Is this sound normal?"

[2265] 2. The device sends the voice and text data to the server.

[2266] 3. The server uses voice analysis to detect abnormal sound patterns and uses text analysis to understand the content of the question.

[2267] 4. The emotion engine analyzes the user's emotional state.

[2268] 5. The server generates advice on the cause of the abnormal sound and how to deal with it, for example, by telling a nervous user, "This sound is caused by worn bearings. It would be a good idea to replace the bearings."

[2269] 6. The device displays the received advice to the user and plays it aloud.

[2270] 7. The user follows the advice to identify the cause of the abnormal sound and take action.

[2271] Prompt Sentence Examples

[2272] "Please tell me how to replace the machine parts. I'm a little nervous."

[2273] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2274] Step 1:

[2275] The user captures video using a device. For example, the user may record video of a machine operation or part replacement using a smartphone, and then ask a question by voice, such as "Please tell me how to remove this bolt." Specifically, the user's operation of the equipment is recorded, and the audio is collected using the device's microphone.

[2276] Input: Video and audio data captured by a smartphone.

[2277] Output: The captured video and audio data is saved on the device.

[2278] Step 2:

[2279] When a user is in a noisy environment, he or she uses the device to input text. For example, he or she inputs text such as "Is this sound normal?". The input is performed using the device's keyboard or touch screen.

[2280] Input: Text data that a user types into a terminal.

[2281] Output: The entered text data is saved to the terminal.

[2282] Step 3:

[2283] The device temporarily stores video, audio, and text data and prepares it for later transmission to the server. The data is temporarily stored in the device's memory.

[2284] Input: Captured video, audio, and text data.

[2285] Output: Video, audio, and text data temporarily stored in the device's memory.

[2286] Step 4:

[2287] The emotion recognition module analyzes the user's emotions in real time from their video and audio data. For example, it analyzes the user's facial expressions and tone of voice to detect tension or stress. This emotion analysis is performed in real time and is executed on the device.

[2288] Input: Captured video and audio data.

[2289] Output: Parsed user emotion data.

[2290] Step 5:

[2291] All data captured by the device (video data, audio data, text data, and emotion data) is sent to the server. The data is transferred using a high-speed, stable communication network.

[2292] Input: Temporarily stored video, audio, and text data, as well as analyzed emotion data.

[2293] Output: Video, audio, text data and emotion data sent to the server.

[2294] Step 6:

[2295] The server analyzes the received video data. Using video analysis means, it detects specific actions and parts in each frame. For example, analysis is performed to identify the location and operation method of a bolt in the video.

[2296] Input: Video data sent to the server.

[2297] Output: Analyzed video data.

[2298] Step 7:

[2299] The server converts the received voice data into text and uses voice analysis means to understand the content of the user's question. For example, it analyzes voice data such as "Please tell me how to remove this bolt."

[2300] Input: The audio data sent to the server.

[2301] Output: The audio data converted to text and the parsed questions.

[2302] Step 8:

[2303] The server uses text analysis means to understand the content of the input text data, for example, analyzing the content of a question such as "Is this sound normal?"

[2304] Input: The text data sent to the server.

[2305] Output: Parsed text data and question content.

[2306] Step 9:

[2307] The emotion engine analyzes the received emotion data to detect the user's stress level and emotional state, for example, assessing whether the user is nervous.

[2308] Input: Emotion data sent to the server.

[2309] Output: The analyzed emotional state of the user.

[2310] Step 10:

[2311] The server combines video, audio, text, and emotion data to generate optimal advice. For example, if the user is nervous, the server generates advice such as, "Turn this bolt clockwise slowly. Use a tool if necessary."

[2312] Input: Analyzed video data, audio data, text data, and emotion data.

[2313] Output: The generated advice data.

[2314] Step 11:

[2315] The server sends the generated advice data to the terminal, securing an appropriate network path for rapid transfer.

[2316] Input: Generated advice data.

[2317] Output: Advice data sent to the terminal.

[2318] Step 12:

[2319] The device displays the received advice data to the user. It displays a text message on the screen and plays the advice aloud. For example, the device might display "Turn this bolt slowly clockwise. Use a tool if necessary" and play the advice aloud.

[2320] Input: Advice data sent by the server.

[2321] Output: The advice that was displayed and played back to the user.

[2322] The above is a flow of the specific processing steps of the system. It includes details of the data processing and calculations performed at each step, which allows for effective and intuitive training for users.

[2323] (Application example 2)

[2324] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2325] In the manufacturing industry, there are problems with efficient and intuitive training due to a shortage of trainers and insufficient sharing of know-how. Furthermore, training and advice that do not take into account the user's emotional state can reduce its effectiveness in the field.

[2326] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video, means for capturing audio, means for accepting text input, means for analyzing the user's emotional state, means for providing advice to the user based on the analysis results and the user's emotional state, means for transmitting the captured video and audio to the server, means for displaying the advice transmitted from the server on the user terminal, means for analyzing prompt sentences using a generative AI model and generating advice corresponding to the emotion, and means for integrating and analyzing the user's video data, audio data, and text data, inputting the prompt sentences into the generative AI model, and generating advice based on the analysis results. This enables efficient and intuitive training and advice provision while taking the user's emotional state into consideration.

[2327] A "video capture means" is a device or method for obtaining images or video of a user or the work environment.

[2328] An "audio capturing means" is a device or method for recording the sounds of a user or the work environment.

[2329] The "means for accepting text input" refers to a device or method that allows a user to input characters using a keyboard, an on-screen keyboard, or the like.

[2330] A "means for analyzing a user's emotional state" is a device or method for inferring a user's emotional state using audio, video, and other data.

[2331] The "means for providing advice to a user based on the analysis results and the emotional state of the user" is a device or method for combining the results of data analysis with the emotional state to provide the user with optimal advice.

[2332] The "means for transmitting captured video and audio to a server" refers to a device or method for transferring captured video and audio data to a server via a network.

[2333] The "means for displaying advice sent from the server on the user terminal" refers to a device or method for displaying instructions or advice received from the server in a form that can be confirmed by the user.

[2334] "Means for analyzing prompt sentences using a generative AI model and generating advice based on emotions" refers to a device or method that utilizes AI technology to analyze text data entered by a user and generate appropriate advice based on that data and the user's emotional state.

[2335] "Means for integrating and analyzing a user's video data, audio data, and text data, inputting prompt sentences into a generative AI model, and generating advice based on the analysis results" refers to a device or method for integrating and analyzing multiple data sources, analyzing prompt sentences using AI based on the results, and providing appropriate advice to the user.

[2336] The present invention is a system for a factory robot that provides specific advice while taking into account the emotional state of the user. The system includes the following main components:

[2337] System Configuration

[2338] 1. Terminal: This applies to robots that work in factories. Robots use cameras and microphones to capture video and audio of their work.

[2339] Video capture method

[2340] Audio capture method

[2341] Text input method

[2342] Emotion Recognition Module

[2343] Data transmission method

[2344] 2. Server: A system deployed in a cloud environment that performs analysis and generates advice. The server includes the following means:

[2345] Video analysis methods

[2346] Voice analysis methods

[2347] Text Analysis Methods

[2348] Emotion Recognition Module

[2349] Advice Generation Method

[2350] Data Receiving Method

[2351] Generative AI Models

[2352] 3. Communication network: Infrastructure for sending and receiving data between terminals and servers.

[2353] Operation overview

[2354] Data Capture and Transmission

[2355] As the user performs a task, the device (robot) captures video and audio using a camera and microphone. The device also includes a means for the user to input text using a keyboard or voice input. This data is sent to a server in real time.

[2356] Data analysis and emotion recognition

[2357] The server integrates and analyzes the received video, audio, and text data. The emotion recognition module analyzes the user's emotional state, and the generative AI model analyzes the prompt sentences to generate appropriate advice.

[2358] Advice generation and delivery

[2359] The generated advice is sent from the server to the terminal in text and audio format, and the terminal displays the received advice to the user and plays it back aloud.

[2360] Hardware and software used

[2361] Hardware: Camera, microphone, user input devices (keyboard, touch screen, etc.), robot body

[2362] Software: Video analysis software (OpenCV, etc.), audio analysis software (SpeechRecognition, etc.), emotion recognition software (EmotionRecognizer, etc.), generative AI models (cloud-based AI services, etc.)

[2363] Specific examples

[2364] Machine part replacement

[2365] If a user is unsure how to remove a bolt, they can enter "Please tell me how to remove this bolt" in text format, and the robot will send video and audio to the server. The server will analyze the data and use an emotion recognition module to confirm that the user is nervous. The advice generator will then generate advice such as "Turn this bolt slowly clockwise. Use a tool if necessary," which will then be displayed and played aloud on the device.

[2366] Check for abnormal sounds

[2367] If a user wants to check for abnormal machine sounds, they record the sound on their smartphone and then type "Is this sound normal?" into the robot while playing it back. The robot sends this to the server, which then analyzes the sound pattern using a voice analysis module and detects tension using an emotion recognition module. The server then generates advice such as "This sound is caused by worn bearings. It would be a good idea to replace the bearings," which the robot displays and plays back in audio.

[2368] Prompt Sentence Examples

[2369] Can you tell me how to remove this bolt?

[2370] "Is this sound normal?"

[2371] Through these specific operations, user support is realized through cooperation between the terminal and the server.

[2372] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2373] Step 1:

[2374] The device uses a camera and microphone to capture video and audio of the user working. As input, video data is obtained from the camera and audio data is obtained from the microphone. Specifically, the device records the user's work scene in real time and temporarily stores the data.

[2375] Step 2:

[2376] The user enters a question requesting clarification through text input. As input, the text data entered by the user via a keyboard or touch screen is acquired. Specifically, the device provides an interface for accepting the user's input and saves the entered text.

[2377] Step 3:

[2378] The device sends the captured video, audio, and text data to the server. The device uses the saved video, audio, and text data as input and sends this data to the server as output. Specifically, the device uploads the data to the cloud server using the network.

[2379] Step 4:

[2380] The server analyzes the video data and identifies the user's work. It uses the video data sent from the device as input and obtains the analysis results as output. Specifically, the server uses video analysis software to recognize specific actions and parts in each frame and compares them with a database.

[2381] Step 5:

[2382] The server analyzes the voice data and understands the user's question. It uses the voice data sent from the device as input and outputs the converted text string and the analysis results. Specifically, it uses voice recognition software to convert the voice into text and analyzes the content of the question.

[2383] Step 6:

[2384] The server analyzes the user's emotional state. It uses video and audio data as input and obtains the emotion analysis results as output. Specifically, it uses an emotion recognition module to estimate the user's emotional state from facial expressions, tone of voice, etc.

[2385] Step 7:

[2386] The server integrates the analysis results and the emotional state and generates advice using a generative AI model. The analysis results, emotion analysis results, and prompt sentences are used as input, and advice for the user is obtained as output. Specifically, the generative AI model analyzes the prompt sentence and generates appropriate advice based on the data.

[2387] Step 8:

[2388] The server sends the generated advice to the terminal. The generated advice data is used as input to obtain data to be sent to the terminal as output. In concrete terms, the server uploads the advice data to the terminal via the network.

[2389] Step 9:

[2390] The device displays the advice received from the server to the user and plays it back aloud. It uses the advice data sent from the server as input and provides visual and auditory feedback to the user as output. Specifically, the device displays a text message on the screen and plays back an audio message from the speaker.

[2391] Prompt Sentence Examples

[2392] Can you tell me how to remove this bolt?

[2393] "Is this sound normal?"

[2394] The above are the specific processing steps of this system.

[2395] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2396] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2397] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2398] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2399] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2400] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2401] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2402] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2403] ...

Claims

1. a means for capturing video; a means for capturing audio; means for accepting text input; means for analyzing the captured video; means for analyzing the captured audio; means for parsing input text; The system includes a means for providing advice to the user based on the analysis results.

2. means for transmitting the captured video and audio to a server; means for displaying the advice sent from the server on the user terminal; The system of claim 1 further comprising:

3. 2. The system according to claim 1, further comprising means for integrating and analyzing the user's video data, audio data and text data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A