System

The system addresses real-time data collection and analysis challenges by preprocessing and AI-driven anomaly detection, enhancing production efficiency and quality through continuous feedback loops.

JP2026019141APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120550
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current manufacturing systems face challenges in achieving real-time data collection and analysis, separation of data collection and analysis processes, limited collaboration between humans and robots, and inefficient production line efficiency and quality control.

Method used

A system that collects text, voice, and image data in real-time from sensors, preprocesses the data, analyzes it using AI models for anomaly detection and quality evaluation, and provides real-time notifications and feedback loops for continuous improvement.

Benefits of technology

Enables real-time monitoring and optimization of production lines, improving efficiency and quality by integrating data collection, analysis, and human-robot collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019141000001_ABST
    Figure 2026019141000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting real-time text, audio, image, and video data from a variety of sensors; means for pre-processing and buffering the collected data in a database; means for analyzing the pre-processed data and making anomaly detection, quality assessment, and process optimization suggestions; means for communicating real-time analysis results and optimization suggestions to field operators and receiving feedback from the operators; and means for improving analysis algorithms based on feedback from the operators.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The purpose of this invention is to solve issues related to production efficiency and quality control in the manufacturing industry. In particular, there is a need to improve production line efficiency and quality by collecting and analyzing a variety of production process data in real time, detecting anomalies, evaluating quality, and proposing process optimization. However, in many current systems, data collection and analysis are separated, making it difficult to achieve real-time performance and feedback loops. Furthermore, there are limited means for effectively realizing collaboration between humans and robots. Therefore, there is a need to build an effective system for improving production efficiency and quality that can be widely applied, from small and medium-sized factories with limited resources to large corporations. [Means for solving the problem]

[0005] The present invention solves the above problems by providing the following means: A means for collecting text data, voice data, image data, and video data in real time from various sensors is provided. A means for preprocessing the collected data and temporarily storing it in a database is provided. A means for analyzing the preprocessed data to detect anomalies, evaluate quality, and propose process optimization is included. Furthermore, a means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers is provided. Additionally, by providing a means for improving the analysis algorithm based on feedback from the workers, the system can continuously evolve and improve production efficiency and quality. In particular, a means for inputting preprocessed data into an artificial intelligence model for analysis and generating improvement proposals for the production process based on the analysis results is provided. The system also includes a means for analyzing voice data from sensors to detect abnormal machine sounds and a means for performing visual inspections of products using image data from sensors, thereby achieving multifaceted data-driven quality improvement.

[0006] A "sensor" is a device that measures a specific physical quantity and outputs it as a signal.

[0007] "Data collection" is the process of collecting data such as text, audio, images, and video in real time using sensors.

[0008] "Preprocessing" is the process of preparing collected data for easier analysis, and includes noise removal, format conversion, and data integration.

[0009] A "database" is a system for temporarily storing large amounts of collected data and managing it in a form that can be accessed as needed.

[0010] "Analysis" is the process of using collected and pre-processed data to generate information for anomaly detection, quality assessment, and process optimization.

[0011] "Anomaly detection" is the process of automatically identifying abnormal situations that deviate from normal patterns in collected data.

[0012] "Quality evaluation" is the process of objectively evaluating the quality of a product or process based on data.

[0013] "Process optimization" is the process of generating specific improvement proposals based on data to improve the efficiency and quality of a production process.

[0014] "Feedback" is the process by which field workers provide opinions and data on analysis results and suggestions, which can be used to improve the system.

[0015] An "artificial intelligence model" is an algorithm or system that analyzes large amounts of data and learns patterns to make complex judgments and predictions.

[0016] "Real-time notification" refers to the process of instantly sending analysis results and optimization suggestions to field workers' devices immediately after they are collected.

[0017] "Audio data" refers to data that includes sound information collected using a sensor.

[0018] "Image data" refers to data that includes information about photographs and videos taken using a sensor.

[0019] "Visual inspection" is the process of analyzing collected image data to automatically detect the quality and defects of a product's appearance. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] The present invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry, and is specifically implemented as follows.

[0042] The central roles of the system are played by a server, various sensors, and terminals (used by users). Various sensors placed on the production line collect text data (numerical information such as temperature and pressure), audio data (machine sounds), image data (photos of products), and video data (video of the production line in operation) in real time. This data is sent to the server and managed centrally.

[0043] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, processes such as removing noise from audio data and standardizing the resolution of image data are performed. Next, the preprocessed data is analyzed using an AI model. This generates suggestions for anomaly detection, quality assessment, and process optimization.

[0044] Analysis results and optimization suggestions are sent to the terminal in real time, allowing the user (field worker) to check the analysis results and take necessary measures using a tablet or smartphone. Users can also enter feedback, which allows the server to continuously improve the analysis algorithm.

[0045] A specific example is shown below.

[0046] ---

[0047] Example 1: Detecting abnormal sounds from machinery

[0048] The server analyzes machine sound data collected from the audio sensor in real time. If an abnormal sound that differs from normal operating sounds is detected, the server notifies the on-site worker's device of the occurrence of the abnormal sound. The user receives this notification, checks the machine's status, and performs maintenance work if necessary. This makes it possible to minimize production line downtime.

[0049] ---

[0050] Example 2: Automatic detection of surface defects on products

[0051] The server analyzes product images collected from the camera sensor and automatically detects small scratches or stains on the surface. If an abnormality is detected, the image data and the location of the abnormality are notified to the user's device. The user can check the notification and take action to re-inspect or correct the product. This makes it possible to maintain a high level of product quality.

[0052] ---

[0053] Example 3: Process optimization proposal

[0054] The server collects and analyzes numerical data related to the production process, such as temperature, pressure, and speed. Based on the analysis results, it suggests fine-tuning the temperature or pressure, for example. These suggestions are sent to the user's device, and if the user accepts the suggestions, the settings are automatically changed. This improves the efficiency of the entire production process and ensures optimal resource utilization.

[0055] ---

[0056] This system enables real-time monitoring and optimization of manufacturing production lines, dramatically improving production efficiency and quality. By effectively linking data collection, analysis, and feedback, on-site workers and robots can work together to achieve more efficient, higher-quality production.

[0057] The processing flow will be explained below.

[0058] Step 1:

[0059] The server collects text data, audio data, image data, and video data in real time from various sensors, including text data from temperature and pressure sensors, audio data from microphones, and image and video data from cameras.

[0060] Step 2:

[0061] The server preprocesses the collected data. Preprocessing includes data cleansing (such as noise removal), format conversion, and timestamp synchronization. For example, it removes background noise from audio data, standardizes the resolution of image data, and aligns the timestamps of all data.

[0062] Step 3:

[0063] The server then inputs the preprocessed data into the AI ​​model for analysis. During the analysis phase, machine learning algorithms are used to generate anomaly detection, quality assessment, and process optimization recommendations. For example, an anomaly detection model can be used to detect abnormal sounds from audio data, and an image analysis algorithm can be used to automatically detect defects on the surface of a product.

[0064] Step 4:

[0065] The server sends analysis results and optimization proposals in real time to the devices used by field workers, such as tablets or smartphones, which the workers use to check the analysis results and optimization proposals.

[0066] Step 5:

[0067] Users can use their devices to check the analysis results and optimization suggestions, and provide feedback as needed. For example, they can check abnormality detection notifications, inspect the machine's status, restart it, or perform maintenance.

[0068] Step 6:

[0069] The server receives feedback data from users and improves the analysis algorithm, which will enable more accurate data analysis the next time. Specifically, the feedback data is used to adjust the parameters of the machine learning model, improving the accuracy of anomaly detection and quality assessment.

[0070] Example 1

[0071] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0072] Improving production efficiency and quality is an important issue in the manufacturing industry. In particular, there is a demand for the detection of abnormal sounds, early detection of surface defects on products, and real-time process optimization. However, these issues have not been adequately resolved with conventional methods. A system that reduces the burden on on-site workers and provides highly accurate data analysis and notifications in real time is needed.

[0073] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0074] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for inputting the preprocessed data into a generative AI model for analysis, means for notifying the mobile devices of on-site workers of analysis results and improvement suggestions in real time, and means for receiving feedback from workers and improving the analysis algorithm. This enables highly accurate anomaly detection, quality evaluation, and process optimization in real time, dramatically improving production efficiency and quality while reducing the burden on on-site workers.

[0075] A "sensor" is a device that detects physical or environmental data in real time and outputs that information as a signal.

[0076] "Text data" is data that expresses numerical information such as temperature, pressure, and speed as a string of characters.

[0077] "Audio data" refers to sound wave information recorded in digital format. Specifically, it refers to audio information such as machine operating sounds and abnormal sounds.

[0078] "Image data" is data that digitally represents still images captured by a camera or sensor.

[0079] "Video data" is a series of images recorded along a time axis, allowing for visual capture of movements and changes in processes.

[0080] "Preprocessing" refers to processes such as filtering, noise removal, and format standardization that are performed to improve the quality of collected data.

[0081] A "database" is a storage device or system for temporarily storing preprocessed data.

[0082] A "generative AI model" is an algorithm or software that uses artificial intelligence techniques to perform data analysis.

[0083] "Analysis" is a series of procedures for anomaly detection, quality assessment, and process optimization based on collected and pre-processed data.

[0084] "Notification" refers to the act of communicating analysis results and improvement suggestions to the mobile devices of field workers in real time.

[0085] "Feedback" is input from field workers that is used to continually improve the analysis algorithms.

[0086] An "algorithm" is a collection of computational procedures or mathematical formulas for the purpose of data analysis or process optimization.

[0087] A "field worker" is a user who is in charge of work on the production line and receives notifications and suggestions from the system.

[0088] The present invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry. Specifically, it is configured as follows.

[0089] The central roles of the system are played by the server, various sensors, and terminals (used by users). Various sensors (temperature sensors, audio sensors, camera sensors, etc.) placed on the production line collect numerical information (text data) such as temperature and pressure, machine sounds (audio data), product photos (image data), and video of the production line in operation (video data) in real time. This data is sent to the server and managed centrally.

[0090] Data Preprocessing

[0091] The server preprocesses the collected data and temporarily stores it. This preprocessing involves, for example, removing noise from audio data and standardizing the resolution of image data. The technologies used here include audio filtering and image processing algorithms.

[0092] Data analysis

[0093] The preprocessed data is then fed into a generative AI model for analysis. The AI ​​model uses TensorFlow or PyTorch, among others. The analysis generates anomaly detection, quality assessment, and recommendations for process optimization. For example, it can detect abnormal sounds from audio data or identify surface defects on products from image data.

[0094] Notification and feedback of analysis results

[0095] Analysis results and optimization suggestions are sent to the terminal in real time, allowing users (field workers) to check the analysis results using a tablet or smartphone and take any necessary measures. Users can also enter feedback, which allows the server to continuously improve the analysis algorithm.

[0096] Example 1: Detecting abnormal sounds from machinery

[0097] The server analyzes machine sound data collected from the audio sensor in real time. If an abnormal sound that differs from normal operating sounds is detected, the server notifies the on-site worker's device of the occurrence of the abnormal sound. The user receives this notification, checks the machine's status, and performs maintenance work if necessary. This makes it possible to minimize production line downtime.

[0098] Example 2: Automatic detection of surface defects on products

[0099] The server analyzes product images collected from the camera sensor and automatically detects small scratches or stains on the surface. If an abnormality is detected, the image data and the location of the abnormality are notified to the user's device. The user can check the notification and take action to re-inspect or correct the product. This makes it possible to maintain a high level of product quality.

[0100] Example 3: Process optimization proposal

[0101] The server collects and analyzes numerical data related to the production process, such as temperature, pressure, and speed. Based on the analysis results, it suggests fine-tuning the temperature or pressure, for example. These suggestions are sent to the user's device, and if the user accepts the suggestions, the settings are automatically changed. This improves the efficiency of the entire production process and ensures optimal resource utilization.

[0102] Prompt Sentence Examples

[0103] The server should analyze the audio data in real time and detect any abnormal sounds.

[0104] "The server should analyze the product image data and automatically detect surface defects."

[0105] "The server should analyze production process data such as temperature and pressure and suggest optimal settings."

[0106] This system enables real-time monitoring and optimization of manufacturing production lines, dramatically improving production efficiency and quality. A series of processes—data collection, preprocessing, analysis, notification, and feedback—work together to enable on-site workers and robots to work together and achieve highly efficient, high-quality production.

[0107] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0108] Step 1: Data collection

[0109] The server collects data in real time from various sensors installed on the production line, such as temperature sensors, audio sensors, and camera sensors. This data includes text data on temperature and pressure, audio data on machine operating sounds and abnormal sounds, and product images and video data.

[0110] Input: Real-time data from various sensors

[0111] Data processing: receiving and storing sensor data

[0112] Output: Collected raw sensor data

[0113] Specific behavior:

[0114] The server obtains temperature data from the temperature sensor every 30 seconds. For example, it collects data such as "Temperature: 27.5°C, 28.0°C, 27.8°C."

[0115] Audio data is streamed from the audio sensor every second to capture the sound of the machine operating.

[0116] The camera sensor captures images of the product every minute and collects video data every five minutes.

[0117] Step 2: Data Preprocessing

[0118] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, noise is removed from the audio data and the resolution of the image data is standardized.

[0119] Input: Collected raw sensor data

[0120] Data processing: noise filtering, resolution unification, format conversion, etc.

[0121] Output: Pre-processed, high-quality data

[0122] Specific behavior:

[0123] The server applies a noise filter to the audio data to generate clear audio data.

[0124] The server standardizes the resolution of image data to a uniform 1024x768 pixels and converts it from JPEG format to PNG format.

[0125] Step 3: Data analysis

[0126] The server inputs the preprocessed data into a generative AI model for analysis, which generates recommendations for anomaly detection, quality assessment, and process optimization. The AI ​​model uses TensorFlow and PyTorch.

[0127] Input: Preprocessed data

[0128] Data processing: Data analysis using AI models

[0129] Output: Analysis results and optimization suggestions

[0130] Specific behavior:

[0131] The server uses TensorFlow to detect abnormal sounds from the preprocessed audio data, and obtains a result such as "abnormal sound detected at timestamp '00:01:15'".

[0132] The server uses PyTorch to analyze the image data and identify defects on the product surface. For example, it obtains a result such as "Small scratches on the product surface detected at region (x: 200, y: 350)."

[0133] Step 4: Notification of analysis results and feedback

[0134] The server notifies the device in real time of analysis results and optimization suggestions, allowing the user to respond, and also receives user feedback to continuously improve the analysis algorithm.

[0135] Input: Analysis results and optimization proposals

[0136] Data processing: notification generation, feedback processing

[0137] Output: Notifying field workers, collecting feedback data

[0138] Specific behavior:

[0139] The server then sends a notification to the on-site worker's tablet regarding any abnormal sounds detected as a result of the analysis, such as "Abnormal sound detected at timestamp 00:01:15."

[0140] The user checks the notification displayed on the tablet and performs the necessary maintenance work.

[0141] Users input the results and feedback after maintenance work into a tablet and send it to the server, which uses this feedback to improve its analysis algorithm.

[0142] Through these steps, the system can improve the efficiency and quality of the production line in real time.

[0143] (Application example 1)

[0144] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0145] Improving production efficiency and quality is a key challenge in modern manufacturing. However, conventional systems have difficulty analyzing data in real time and responding immediately, making early detection of abnormalities and process optimization insufficient. This makes it difficult to minimize production line downtime and maintain a high level of quality.

[0146] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0147] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality evaluation, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for improving the analysis algorithm based on feedback from the workers, means for robots installed on the production line to collect operation data in real time and send it to the server, means for the robots to autonomously optimize themselves based on the analysis results from the server, and means for notifying on-site workers of the analysis results so that they can make manual adjustments. This enables real-time anomaly detection, quality evaluation, and process optimization, thereby improving the efficiency and quality of the production line.

[0148] The "various sensors" are various types of sensors used in manufacturing sites, and are devices that have the function of collecting text data, audio data, image data, and video data.

[0149] "Text data" is data that expresses the state of a machine or environment as numbers or strings of characters, and is primarily information indicating temperature, pressure, speed, etc.

[0150] "Audio data" refers to data collected as acoustic signals from sounds generated from production lines and machines.

[0151] "Image data" is digital data of still images captured by an imaging device such as a camera, and represents the state of a product or production line.

[0152] "Video data" refers to digital data of video captured by a camera or video device, and is data that includes a series of images.

[0153] "Preprocessing" refers to the process of converting collected data into a format suitable for analysis, and includes noise removal and resolution unification.

[0154] A "database" is a computer system for temporarily storing preprocessed data.

[0155] "Analysis" is the process of applying calculations and models to pre-processed data to detect anomalies, assess quality, and suggest process optimization.

[0156] "Abnormality detection" refers to detecting abnormal conditions that differ from normal operating conditions, such as detecting abnormal sounds or vibrations in machinery.

[0157] "Quality evaluation" is the process of evaluating the quality of a product or process to determine whether it meets specified standards.

[0158] "Process optimization" is the process of finding the optimal parameters and procedures in a production process to improve overall efficiency.

[0159] "Real-time" means that the system collects data and analyzes and notifies immediately, with little time delay.

[0160] "Feedback" refers to information that collects opinions and evaluations from workers and system users and is used to improve the system.

[0161] A "server" is a central processing unit that collects, stores, analyzes, and processes feedback data on a network.

[0162] A "robot" is a mechanical device that is installed in a production line and can operate autonomously.

[0163] A "terminal" is a device such as a tablet or smartphone used by a field worker to receive notifications from the server.

[0164] This invention is an AI-driven robotics assistance system for improving production efficiency and quality in the manufacturing industry. The system collects data in real time from various sensors installed on the production line, and a server analyzes the data to detect anomalies and optimize processes.

[0165] Hardware and software used

[0166] Hardware:

[0167] Sensors: Temperature sensors, pressure sensors, camera sensors, audio sensors, etc. These sensors are installed at various points along the production line to collect necessary data in real time.

[0168] Robot: A machine with autonomous capabilities installed on a production line that receives data from sensors and performs the required actions.

[0169] Terminal: A device such as a tablet or smartphone used by field workers. It receives notifications from the server and allows data to be checked and manually adjusted.

[0170] software:

[0171] Server (Central Processing Unit): Collects, pre-processes, analyzes, and processes feedback data.

[0172] OpenCV: A library for capturing and processing camera images, and is used to analyze image data.

[0173] TensorFlow: A library that performs data analysis using AI models, and is used for anomaly detection and quality assessment.

[0174] Requests: A data communication library used to send data from sensors and robots to a server.

[0175] Specific explanation for carrying out the invention

[0176] First, various sensors installed on the production line collect temperature, pressure, image, audio, and video data in real time. The collected data is immediately sent to a server, where it is preprocessed, for example, to remove noise from audio data and standardize the resolution of image data.

[0177] Once preprocessed, the data is analyzed using TensorFlow, which generates anomaly detection, quality assessment, and process optimization recommendations. The analysis results and optimization recommendations are sent to on-site workers' devices in real time. Workers can view the analysis results on their tablets or smartphones and take any necessary action.

[0178] Furthermore, robots installed on the production line collect data from various sensors in real time and send it to a server. The robots receive the analysis results from the server and can then optimize themselves autonomously. This system makes it possible to minimize production line downtime and maintain a high level of quality.

[0179] Specific examples (prompt sentence examples)

[0180] Below is an example of a prompt sentence to input to the generative AI model.

[0181] Create a program to analyze sensor data (images, temperature, pressure, etc.) collected by factory robots in real time, detect anomalies, and propose process optimization. Include a function for the robot to autonomously adjust itself based on the analysis results. Use OpenCV for image analysis, TensorFlow for AI analysis, and the Requests library for data communication.

[0182] According to the aspects of the present invention, real-time anomaly detection, quality evaluation, and process optimization become possible, thereby realizing improved efficiency and quality of the production line.

[0183] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0184] Step 1:

[0185] The sensors collect temperature, pressure, image, and audio data from the production line in real time. These data are collected in the form of sensors that each have. The input is the environmental condition of the production line, and the output is the raw data obtained from each sensor.

[0186] Step 2:

[0187] The server preprocesses the collected data, specifically removing noise from audio data, standardizing the resolution of image data, and standardizing other data formats. The input is raw data sent from the sensor, and the output is preprocessed data.

[0188] Step 3:

[0189] The server analyzes the preprocessed data. This analysis uses TensorFlow to perform anomaly detection, quality assessment, and process optimization recommendations. The analysis requires advanced computations and runs models to identify abnormal sounds and surface defects. The input is the preprocessed data, and the output is the analysis results and recommendations.

[0190] Step 4:

[0191] The server notifies the on-site worker of the analysis results and optimization proposals in real time. The worker's device is a tablet or smartphone, where they can check the analysis results and take any necessary measures. The input is the analysis results and proposals, and the output is the notification received by the on-site worker.

[0192] Step 5:

[0193] The user receives the notification and sends feedback from the device to the server. The feedback includes opinions about the analysis results and specific adjustments. The input is the user's feedback on the notification, and the output is the feedback data sent to the server.

[0194] Step 6:

[0195] The server improves the analysis algorithm based on the received feedback. This improves the accuracy of the analysis from the next time onwards, and further increases in production efficiency are expected. The input is feedback data from the user, and the output is a revised analysis algorithm.

[0196] Step 7:

[0197] The robot collects operational data in real time from various sensors installed on the production line and sends it to a server. The input is the operational data collected by the robot, and the output is the data sent to the server.

[0198] Step 8:

[0199] The server analyzes the operation data from the robot and feeds the results back to the robot. The robot receives the analysis results and performs optimization autonomously. The input is the robot's operation data and analysis results, and the output is the robot's optimized behavior.

[0200] This series of processing flows enables real-time monitoring and optimization of the entire production line.

[0201] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0202] This invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry, incorporating an emotion engine that recognizes the user's emotional state in real time and reflects it in optimization suggestions.

[0203] This system includes a server, various sensors, user devices, and an emotion engine. Sensors placed on the production line collect text data (numerical information such as temperature and pressure), audio data (machine noise, user voice), image data (photos of products, user facial expressions), and video data (video of the production line in operation) in real time. This data is sent to the server and managed centrally.

[0204] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, processes such as removing noise from audio data and standardizing the resolution of image data are performed. Next, the preprocessed data is analyzed using an AI model. This generates suggestions for anomaly detection, quality assessment, and process optimization.

[0205] Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotional state in real time. The emotion engine analyzes the user's emotions from voice and image data, detecting, for example, stress and fatigue. The analysis results are reflected in the server's analysis algorithm, resulting in the generation of more appropriate optimization suggestions.

[0206] The analysis results and optimization suggestions are sent to the device in real time. This allows the user (field worker) to check the analysis results and optimization suggestions using a tablet or smartphone and take any necessary action. For example, if the user is in a high-stress state, the notification may include a suggestion to take a break or redistribute tasks. User feedback is also analyzed through the emotion engine and used to improve the system as a whole.

[0207] The following is a specific example.

[0208] ---

[0209] Example 1: User stress detection

[0210] The server uses audio and image sensors to detect the user's stress level from their tone of voice and facial expressions. If the emotion engine analyzes the user's emotional state as "high stress," it sends the result to the server. Based on this information, the server notifies the device with suggestions for reducing the user's workload (e.g., taking a break or switching to a lighter task). This maximizes the user's performance while maintaining their physical and mental health.

[0211] ---

[0212] Example 2: Emotional analysis feedback

[0213] If a user feels dissatisfied or stressed during product inspection, the emotion engine detects that emotional state. The server incorporates this information into the analysis results and receives it as feedback. For example, if the system detects that the user is "tired and prone to making judgment errors," it will adjust the sensitivity of the analysis algorithm next time and strengthen settings to minimize human error.

[0214] ---

[0215] Example 3: Dynamic adjustment of operating environment

[0216] The server analyzes data from the entire production line and the user's emotional data to suggest dynamic adjustments to the operating environment. For example, by appropriately changing the temperature settings for the entire line or adjusting the operating speed of specific machines, overall efficiency can be improved. When making these suggestions based on the user's emotional data, the emotion engine also takes into account whether the working environment is comfortable.

[0217] ---

[0218] This system will enable manufacturing production lines to operate with even greater efficiency and quality. By combining an emotion engine with AI analysis, appropriate suggestions that take into account the user's condition will be made in real time, enabling more effective optimization of the production process.

[0219] The processing flow will be explained below.

[0220] Step 1:

[0221] The server collects text data, audio data, image data, and video data in real time from various sensors, including text data from temperature and pressure sensors, audio data from microphones, and image and video data from cameras.

[0222] Step 2:

[0223] The server preprocesses the collected data, removing noise from the audio data, standardizing the resolution of the image data, and aligning the timestamps of all data. During this process, the data is temporarily stored in a database.

[0224] Step 3:

[0225] The server then inputs the preprocessed data into an AI model for analysis. During the analysis stage, machine learning algorithms are used to generate anomaly detection, quality assessment, and process optimization recommendations. For example, an anomaly detection model can be used to detect abnormal sounds from audio data, and an image analysis algorithm can be used to automatically detect defects on the surface of a product.

[0226] Step 4:

[0227] The server notifies the user's device of the analysis results and optimization proposals in real time, allowing the device to allow on-site workers to immediately check the analysis results and optimization proposals.

[0228] Step 5:

[0229] The device equipped with the emotion engine analyzes the user's voice tone and facial expression data in real time to evaluate the user's emotional state (e.g., stress, fatigue). Specifically, it detects changes in tone and speed of voice from the voice data, and senses changes in facial expression from the image data.

[0230] Step 6:

[0231] The server receives the emotional state data sent from the emotion engine. Based on this information, the server adjusts the analysis algorithm and generates optimization suggestions according to the user's emotional state. For example, if the user is in a high-stress state, the server generates suggestions for taking a break or reducing workload.

[0232] Step 7:

[0233] The server then sends the newly generated optimization proposals to the user's device, where the user can view these proposals in real time. The user can then use the device to input feedback, such as accepting the proposals or requesting modifications.

[0234] Step 8:

[0235] The server receives feedback data from users and continuously improves the analysis algorithm. The feedback, including emotional data, is analyzed to improve the accuracy of the next data analysis.

[0236] ---

[0237] This system monitors and optimizes manufacturing production lines in real time, continuously improving production efficiency and quality. The introduction of an emotion engine enables flexible responses according to the user's condition, achieving both an improved work environment and increased production efficiency.

[0238] Example 2

[0239] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0240] Improving production efficiency and quality are important challenges in modern manufacturing. However, conventional systems focused on collecting and analyzing data from sensors, and did not fully consider the impact of workers' emotional states. As a result, poor performance and mistakes caused by worker stress and fatigue had a negative impact on production efficiency and quality. This made it difficult to manage worker health and optimize production processes.

[0241] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0242] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality assessment, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for analyzing the emotional state of the workers from their voice data and image data, means for generating optimization proposals based on their emotional states, and means for improving the analysis algorithm based on feedback from the workers. This improves production efficiency and quality and enables optimal proposals that take into account the health status of the workers.

[0243] A "sensor" is a device that collects different types of data in real time, such as temperature, pressure, sound, images, and video.

[0244] "Data preprocessing" refers to the process of performing preprocessing such as noise removal and resolution standardization on collected data, and then temporarily storing it.

[0245] A "database" is a storage device for temporarily storing preprocessed data.

[0246] An "AI model" is an artificial intelligence algorithm that performs analysis for anomaly detection, quality evaluation, and process optimization.

[0247] "Analysis results" are the results of anomaly detection and quality assessment generated by the AI ​​model using preprocessed data.

[0248] An "optimization proposal" is a proposal for improving the production process that is generated based on the analysis results.

[0249] "Field workers" are workers who work on the manufacturing floor and act based on data collected from sensors.

[0250] "Feedback" refers to opinions and reactions to the system analysis and suggestions provided by field workers.

[0251] "Emotional state" refers to the emotional state of a worker, such as stress or fatigue, analyzed from the worker's voice data and image data.

[0252] The "emotion engine" is a system that analyzes the emotional state of workers from collected voice and image data.

[0253] "Notification" is a communication method by which the server conveys analysis results and optimization suggestions to field workers in real time.

[0254] An "analysis algorithm" is a calculation procedure for detecting anomalies and evaluating quality based on collected data.

[0255] This invention is an AI-driven robot support system for improving production efficiency and quality in the manufacturing industry, incorporating an emotion engine that recognizes the user's emotional state in real time and reflects it in optimization proposals. This system is composed of a server, various sensors, user terminals, and the emotion engine.

[0256] The server is the central part that collects data in real time from sensors placed on the production line and performs preprocessing and analysis. The main hardware used is a high-performance server computer, and the software includes Python, TensorFlow, PyTorch, OpenCV, Librosa, etc.

[0257] Data collection

[0258] The server uses various sensors to collect temperature, pressure, audio, image, and video data. For example, the temperature sensor measures the temperature data of the line in real time, the microphone collects user voices and machine sounds, and the image sensor takes photos of the product and records video of the production line operation.

[0259] Data Preprocessing

[0260] The collected data is sent to a server where preprocessing such as noise removal and resolution unification is performed. For example, the Librosa library is used to remove environmental noise from audio data, and the resolution of image data is unified using OpenCV. This organizes the data into a form optimal for analysis.

[0261] Data analysis

[0262] The preprocessed data is then analyzed by the server using AI models. Anomaly detection models are run using TensorFlow and PyTorch. Quality assessments and recommendations for optimizing the production process are also generated at this stage. If an anomaly is detected, its details are recorded in a log.

[0263] Emotion analysis

[0264] Furthermore, the system incorporates an emotion engine that recognizes the user's emotional state in real time. The emotion engine uses natural language processing models (e.g., BERT) and facial recognition models (e.g., DeepFace) to analyze the user's emotions from voice and image data, thereby determining whether the user is in a stressful state.

[0265] Proposal generation and notification

[0266] Optimization suggestions are generated based on the analysis results and the user's emotional state. The server notifies the user of these suggestions in real time. For example, if a user is in a high-stress state, the device may be notified of suggestions such as taking a break or redistributing tasks. The device application uses React Native or Android Studio, allowing users to view these notifications on their tablets or smartphones.

[0267] Feedback collection and system improvement

[0268] Users can provide feedback on the suggestions provided to them through their devices. The server collects this feedback and uses it to improve the analysis algorithm. This allows the system to learn from the feedback and improve the accuracy of future suggestions.

[0269] Example: Detecting user stress

[0270] The server uses audio and image sensors to detect the user's stress level from their tone of voice and facial expressions. If the emotion engine analyzes the user's emotional state as "high stress," it sends the result to the server. Based on this information, the server notifies the device with suggestions for reducing the user's workload (e.g., taking a break or switching to a lighter task). This maximizes the user's performance while maintaining their physical and mental health.

[0271] Example prompt sentence:

[0272] "Analyze the user's stress level based on their voice and image data, and make optimization suggestions based on the results."

[0273] This system will enable manufacturing production lines to operate with even greater efficiency and quality. By combining an emotion engine with AI analysis, appropriate suggestions that take into account the user's condition will be made in real time, enabling more effective optimization of the production process.

[0274] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0275] Step 1: Data collection

[0276] Inputs: temperature, pressure, audio, image, and video data.

[0277] Processing: The server collects data in real time from various sensors placed on the production line.

[0278] Output: Collected multimodal data.

[0279] How it works: The server collects real-time temperature data every 10 seconds from the temperature sensor, and collects user voice and machine sounds from the microphone. At the same time, the image sensor captures the product's appearance, and the video sensor records the operation of the production line.

[0280] Step 2: Data Preprocessing

[0281] Input: Collected multimodal data.

[0282] Processing: The server preprocesses the collected data, removing noise and unifying the resolution for temporary storage.

[0283] Output: Preprocessed data.

[0284] Specific operation: Noise reduction is performed on the audio data using the Librosa library, and the resolution of the image data is unified using OpenCV. The preprocessed data is temporarily stored in a database.

[0285] Step 3: Data analysis

[0286] Input: Preprocessed data.

[0287] Processing: The server uses the preprocessed data to perform analysis using an AI model.

[0288] Output: Analysis results (anomaly detection, quality assessment, process optimization suggestions).

[0289] What it does: Runs anomaly detection models using TensorFlow or PyTorch to generate quality assessments and process optimization suggestions. If an anomaly is detected, details of the anomaly are logged.

[0290] Step 4: Sentiment Analysis

[0291] Input: Preprocessed audio and image data.

[0292] Processing: The server analyzes the user's emotional state using an emotion engine.

[0293] Output: Emotion analysis results (e.g. high stress, fatigue).

[0294] Specific operation: Emotions are analyzed from voice data using BERT, and facial expressions are analyzed from image data using DeepFace. The analysis results of the emotional state are stored on the server.

[0295] Step 5: Proposal Generation

[0296] Input: Data analysis results and sentiment analysis results.

[0297] Processing: The server generates optimization suggestions based on the data analysis results and sentiment analysis results.

[0298] Output: Optimization suggestions (e.g. break suggestions, task redistribution).

[0299] Specific operation: Based on the analysis results, suggestions to reduce the user's workload are automatically generated and prepared for notification in the next step.

[0300] Step 6: Notification of results

[0301] Input: Optimization proposal.

[0302] Processing: The server notifies the user terminal of the generated optimization proposal.

[0303] Output: A notification message on the terminal.

[0304] Specific operation: The server pushes JSON format data to the device, and the application on the device interprets the data and displays it to the user.

[0305] Step 7: Gather feedback and improve the system

[0306] Input: User feedback.

[0307] Processing: The user enters feedback on the suggestions provided, and the server collects the feedback and uses it to improve the system.

[0308] Output: Feedback data, improved AI model.

[0309] How it works: Users input feedback from their devices, and the server collects that information and reflects it in the next analysis, thereby improving the accuracy of the analysis algorithm.

[0310] These steps not only improve the efficiency and quality of the production process, but also optimize the working environment, maintaining worker health and performance through suggestions that take the user's emotional state into account.

[0311] (Application example 2)

[0312] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0313] Conventional production systems focus on detecting machine anomalies and evaluating quality, but rarely consider the emotional state of human workers. This increases the risk of workers feeling stressed or fatigued, leading to reduced production efficiency. Furthermore, the lack of work instructions or improvement suggestions based on the emotional state of workers poses a challenge, making it difficult to fully optimize the entire production process.

[0314] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0315] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality assessment, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for improving the analysis algorithm based on feedback from the workers, and means for analyzing the emotions of the workers in real time and providing work instructions based on their emotional states. This makes it possible to provide optimal work instructions and propose process improvements that take the emotional states of the workers into consideration, thereby achieving both improved production efficiency and maintaining the health of the workers.

[0316] "Multiple sensors" are various sensors used to collect text data, audio data, image data, and video data in real time.

[0317] "Preprocessing" refers to the process of converting collected data into a form suitable for analysis, and specifically includes noise removal and resolution unification.

[0318] A "database" is a system for temporarily storing preprocessed data.

[0319] "Analysis algorithms" are techniques for analyzing pre-processed data and generating suggestions for anomaly detection, quality assessment, and process optimization.

[0320] "Emotional state" refers to the emotional state of the worker analyzed from voice and image data, and specifically refers to stress and fatigue.

[0321] "Work instructions" refers to specific work tasks and break suggestions provided to workers based on their emotional state.

[0322] "Feedback" refers to responses and reactions from workers, and is information collected to help improve the system.

[0323] An "emotion analysis device" is a piece of equipment used to analyze the emotional state of a worker in real time.

[0324] "Production process improvement proposals" refer to specific proposals for improving production efficiency based on the analysis results.

[0325] "Stress state" refers to the degree of tension and strain felt by the worker, and is analyzed by the emotion engine.

[0326] To realize the present invention, the following detailed description of the system and its components is required. The system of the present invention collects and analyzes data in real time and provides appropriate feedback and work instructions. Specific components include a server, various sensors, a user's device, and an emotion analysis engine.

[0327] server

[0328] The server is the core processing unit of the system and performs the following functions:

[0329] Data collection: Collect text, audio, image, and video data in real time from a variety of sensors.

[0330] Data preprocessing: The collected data is preprocessed by noise removal and resolution standardization, and then temporarily stored in a database.

[0331] Data Analysis: The pre-processed data is analyzed using generative AI models to generate recommendations for anomaly detection, quality assessment, and process optimization.

[0332] Emotion analysis: An emotion analysis engine is used to analyze the user's emotional state from voice and image data.

[0333] Various sensors

[0334] The system uses the following sensors:

[0335] Audio sensor: Collects the user's voice tone and mechanical sounds to detect stress levels and abnormal sounds.

[0336] Image sensor: Collects footage of the user's facial expressions and production line operations to inspect the emotional state and appearance of the product.

[0337] User's device

[0338] The user's device (smartphone, tablet, etc.) performs the following functions:

[0339] Notifications: Analysis results and optimization suggestions from the server are displayed in real time.

[0340] Feedback: Collects feedback from users and sends it to the server.

[0341] Sentiment Analysis Engine

[0342] The emotion analysis engine analyzes the user's emotional state from voice data and image data and provides the data to the server.

[0343] This is implemented using the following software and libraries:

[0344] OpenCV: Face Recognition and Landmark Detection

[0345] dlib: Facial landmark analysis

[0346] Keras: Sentiment Analysis Model

[0347] requests: Send notification

[0348] Specific examples

[0349] For example, when developing an application to detect the stress level of workers at a logistics center, the smartphone's camera and microphone can be used to analyze the worker's facial expressions and tone of voice. If the worker's emotional state is determined to be "high stress," the server will send a notification to take a break. In this way, optimal work instructions that take the worker's emotions into consideration can be provided, thereby improving production efficiency and maintaining the worker's health at the same time.

[0350] Prompt Sentence Examples

[0351] Develop an application that evaluates workers' emotions through real-time analysis of camera footage and suggests breaks when they are stressed. The technologies used are OpenCV, dlib, and Keras. The goal is to analyze emotions from workers' facial expressions and suggest breaks when stress or fatigue is detected.

[0352] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0353] Step 1:

[0354] The server collects text data, voice data, image data, and video data in real time from various sensors (audio sensors, image sensors). The input is data from the sensors, and the output is raw data for preprocessing. Specifically, the audio sensor collects the tone of the user's voice, and the image sensor captures the user's facial expressions and footage of the production line operation.

[0355] Step 2:

[0356] The server preprocesses the collected data. Specifically, it removes noise from the audio data and standardizes the resolution of the image data. At this point, the input is raw data and the output is preprocessed data. The preprocessed data is temporarily stored in a database.

[0357] Step 3:

[0358] The server inputs the preprocessed data into the generative AI model for analysis. The input is the preprocessed data, and the output of the analysis is anomaly detection, quality assessment, and process optimization recommendations. Specifically, the AI ​​model uses image data to perform visual inspections of products and analyzes audio data to detect abnormal machine sounds.

[0359] Step 4:

[0360] The server analyzes the user's emotional state using an emotion analysis engine. The input is audio and image data, and the output is the user's emotional state (stress or fatigue). Specifically, it uses OpenCV and dlib to detect facial landmarks, and classifies the emotional state using an emotion analysis model using Keras.

[0361] Step 5:

[0362] The server notifies the user device of the analysis results and optimization suggestions in real time. The input is the analysis results and sentiment analysis results, and the output is a notification message. Specifically, it uses the requests library to send a message to the notification API, which is then displayed on the user device. For example, if a high stress state is detected, a notification such as "Take a break" is sent.

[0363] Step 6:

[0364] The user checks the suggestions through the terminal and takes the necessary action. The input is a notification message and the output is feedback information. The suggestions are displayed on the terminal, and the user can check them and take action.

[0365] Step 7:

[0366] The terminal collects feedback from the user and sends it to the server. The input is the feedback information from the user, and the output is the feedback data sent to the server. Specifically, an interface is provided for the user to report task completion notifications and new problems.

[0367] Step 8:

[0368] The server improves the analysis algorithm based on the collected feedback. The input is the feedback data, and the output is an improved analysis algorithm. Specifically, it reevaluates the analysis results and adjusts the model parameters.

[0369] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0370] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0371] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0372] [Second embodiment]

[0373] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0374] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0375] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0376] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0377] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0378] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0379] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0380] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0381] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0382] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0383] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0384] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0385] The present invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry, and is specifically implemented as follows.

[0386] The central roles of the system are played by a server, various sensors, and terminals (used by users). Various sensors placed on the production line collect text data (numerical information such as temperature and pressure), audio data (machine sounds), image data (photos of products), and video data (video of the production line in operation) in real time. This data is sent to the server and managed centrally.

[0387] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, processes such as removing noise from audio data and standardizing the resolution of image data are performed. Next, the preprocessed data is analyzed using an AI model. This generates suggestions for anomaly detection, quality assessment, and process optimization.

[0388] Analysis results and optimization suggestions are sent to the terminal in real time, allowing the user (field worker) to check the analysis results and take necessary measures using a tablet or smartphone. Users can also enter feedback, which allows the server to continuously improve the analysis algorithm.

[0389] A specific example is shown below.

[0390] ---

[0391] Example 1: Detecting abnormal sounds from machinery

[0392] The server analyzes machine sound data collected from the audio sensor in real time. If an abnormal sound that differs from normal operating sounds is detected, the server notifies the on-site worker's device of the occurrence of the abnormal sound. The user receives this notification, checks the machine's status, and performs maintenance work if necessary. This makes it possible to minimize production line downtime.

[0393] ---

[0394] Example 2: Automatic detection of surface defects on products

[0395] The server analyzes product images collected from the camera sensor and automatically detects small scratches or stains on the surface. If an abnormality is detected, the image data and the location of the abnormality are notified to the user's device. The user can check the notification and take action to re-inspect or correct the product. This makes it possible to maintain a high level of product quality.

[0396] ---

[0397] Example 3: Process optimization proposal

[0398] The server collects and analyzes numerical data related to the production process, such as temperature, pressure, and speed. Based on the analysis results, it suggests fine-tuning the temperature or pressure, for example. These suggestions are sent to the user's device, and if the user accepts the suggestions, the settings are automatically changed. This improves the efficiency of the entire production process and ensures optimal resource utilization.

[0399] ---

[0400] This system enables real-time monitoring and optimization of manufacturing production lines, dramatically improving production efficiency and quality. By effectively linking data collection, analysis, and feedback, on-site workers and robots can work together to achieve more efficient, higher-quality production.

[0401] The processing flow will be explained below.

[0402] Step 1:

[0403] The server collects text data, audio data, image data, and video data in real time from various sensors, including text data from temperature and pressure sensors, audio data from microphones, and image and video data from cameras.

[0404] Step 2:

[0405] The server preprocesses the collected data. Preprocessing includes data cleansing (such as noise removal), format conversion, and timestamp synchronization. For example, it removes background noise from audio data, standardizes the resolution of image data, and aligns the timestamps of all data.

[0406] Step 3:

[0407] The server then inputs the preprocessed data into the AI ​​model for analysis. During the analysis phase, machine learning algorithms are used to generate anomaly detection, quality assessment, and process optimization recommendations. For example, an anomaly detection model can be used to detect abnormal sounds from audio data, and an image analysis algorithm can be used to automatically detect defects on the surface of a product.

[0408] Step 4:

[0409] The server sends analysis results and optimization proposals in real time to the devices used by field workers, such as tablets or smartphones, which the workers use to check the analysis results and optimization proposals.

[0410] Step 5:

[0411] Users can use their devices to check the analysis results and optimization suggestions, and provide feedback as needed. For example, they can check abnormality detection notifications, inspect the machine's status, restart it, or perform maintenance.

[0412] Step 6:

[0413] The server receives feedback data from users and improves the analysis algorithm, which will enable more accurate data analysis the next time. Specifically, the feedback data is used to adjust the parameters of the machine learning model, improving the accuracy of anomaly detection and quality assessment.

[0414] Example 1

[0415] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0416] Improving production efficiency and quality is an important issue in the manufacturing industry. In particular, there is a demand for the detection of abnormal sounds, early detection of surface defects on products, and real-time process optimization. However, these issues have not been adequately resolved with conventional methods. A system that reduces the burden on on-site workers and provides highly accurate data analysis and notifications in real time is needed.

[0417] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0418] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for inputting the preprocessed data into a generative AI model for analysis, means for notifying the mobile devices of on-site workers of analysis results and improvement suggestions in real time, and means for receiving feedback from workers and improving the analysis algorithm. This enables highly accurate anomaly detection, quality evaluation, and process optimization in real time, dramatically improving production efficiency and quality while reducing the burden on on-site workers.

[0419] A "sensor" is a device that detects physical or environmental data in real time and outputs that information as a signal.

[0420] "Text data" is data that expresses numerical information such as temperature, pressure, and speed as a string of characters.

[0421] "Audio data" refers to sound wave information recorded in digital format. Specifically, it refers to audio information such as machine operating sounds and abnormal sounds.

[0422] "Image data" is data that digitally represents still images captured by a camera or sensor.

[0423] "Video data" is a series of images recorded along a time axis, allowing for visual capture of movements and changes in processes.

[0424] "Preprocessing" refers to processes such as filtering, noise removal, and format standardization that are performed to improve the quality of collected data.

[0425] A "database" is a storage device or system for temporarily storing preprocessed data.

[0426] A "generative AI model" is an algorithm or software that uses artificial intelligence techniques to perform data analysis.

[0427] "Analysis" is a series of procedures for anomaly detection, quality assessment, and process optimization based on collected and pre-processed data.

[0428] "Notification" refers to the act of communicating analysis results and improvement suggestions to the mobile devices of field workers in real time.

[0429] "Feedback" is input from field workers that is used to continually improve the analysis algorithms.

[0430] An "algorithm" is a collection of computational procedures or mathematical formulas for the purpose of data analysis or process optimization.

[0431] A "field worker" is a user who is in charge of work on the production line and receives notifications and suggestions from the system.

[0432] The present invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry. Specifically, it is configured as follows.

[0433] The central roles of the system are played by the server, various sensors, and terminals (used by users). Various sensors (temperature sensors, audio sensors, camera sensors, etc.) placed on the production line collect numerical information (text data) such as temperature and pressure, machine sounds (audio data), product photos (image data), and video of the production line in operation (video data) in real time. This data is sent to the server and managed centrally.

[0434] Data Preprocessing

[0435] The server preprocesses the collected data and temporarily stores it. This preprocessing involves, for example, removing noise from audio data and standardizing the resolution of image data. The technologies used here include audio filtering and image processing algorithms.

[0436] Data analysis

[0437] The preprocessed data is then fed into a generative AI model for analysis. The AI ​​model uses TensorFlow or PyTorch, among others. The analysis generates anomaly detection, quality assessment, and recommendations for process optimization. For example, it can detect abnormal sounds from audio data or identify surface defects on products from image data.

[0438] Notification and feedback of analysis results

[0439] Analysis results and optimization suggestions are sent to the terminal in real time, allowing users (field workers) to check the analysis results using a tablet or smartphone and take any necessary measures. Users can also enter feedback, which allows the server to continuously improve the analysis algorithm.

[0440] Example 1: Detecting abnormal sounds from machinery

[0441] The server analyzes machine sound data collected from the audio sensor in real time. If an abnormal sound that differs from normal operating sounds is detected, the server notifies the on-site worker's device of the occurrence of the abnormal sound. The user receives this notification, checks the machine's status, and performs maintenance work if necessary. This makes it possible to minimize production line downtime.

[0442] Example 2: Automatic detection of surface defects on products

[0443] The server analyzes product images collected from the camera sensor and automatically detects small scratches or stains on the surface. If an abnormality is detected, the image data and the location of the abnormality are notified to the user's device. The user can check the notification and take action to re-inspect or correct the product. This makes it possible to maintain a high level of product quality.

[0444] Example 3: Process optimization proposal

[0445] The server collects and analyzes numerical data related to the production process, such as temperature, pressure, and speed. Based on the analysis results, it suggests fine-tuning the temperature or pressure, for example. These suggestions are sent to the user's device, and if the user accepts the suggestions, the settings are automatically changed. This improves the efficiency of the entire production process and ensures optimal resource utilization.

[0446] Prompt Sentence Examples

[0447] The server should analyze the audio data in real time and detect any abnormal sounds.

[0448] "The server should analyze the product image data and automatically detect surface defects."

[0449] "The server should analyze production process data such as temperature and pressure and suggest optimal settings."

[0450] This system enables real-time monitoring and optimization of manufacturing production lines, dramatically improving production efficiency and quality. A series of processes—data collection, preprocessing, analysis, notification, and feedback—work together to enable on-site workers and robots to work together and achieve highly efficient, high-quality production.

[0451] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0452] Step 1: Data collection

[0453] The server collects data in real time from various sensors installed on the production line, such as temperature sensors, audio sensors, and camera sensors. This data includes text data on temperature and pressure, audio data on machine operating sounds and abnormal sounds, and product images and video data.

[0454] Input: Real-time data from various sensors

[0455] Data processing: receiving and storing sensor data

[0456] Output: Collected raw sensor data

[0457] Specific behavior:

[0458] The server obtains temperature data from the temperature sensor every 30 seconds. For example, it collects data such as "Temperature: 27.5°C, 28.0°C, 27.8°C."

[0459] Audio data is streamed from the audio sensor every second to capture the sound of the machine operating.

[0460] The camera sensor captures images of the product every minute and collects video data every five minutes.

[0461] Step 2: Data Preprocessing

[0462] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, noise is removed from the audio data and the resolution of the image data is standardized.

[0463] Input: Collected raw sensor data

[0464] Data processing: noise filtering, resolution unification, format conversion, etc.

[0465] Output: Pre-processed, high-quality data

[0466] Specific behavior:

[0467] The server applies a noise filter to the audio data to generate clear audio data.

[0468] The server standardizes the resolution of image data to a uniform 1024x768 pixels and converts it from JPEG format to PNG format.

[0469] Step 3: Data analysis

[0470] The server inputs the preprocessed data into a generative AI model for analysis, which generates recommendations for anomaly detection, quality assessment, and process optimization. The AI ​​model uses TensorFlow and PyTorch.

[0471] Input: Preprocessed data

[0472] Data processing: Data analysis using AI models

[0473] Output: Analysis results and optimization suggestions

[0474] Specific behavior:

[0475] The server uses TensorFlow to detect abnormal sounds from the preprocessed audio data, and obtains a result such as "abnormal sound detected at timestamp '00:01:15'".

[0476] The server uses PyTorch to analyze the image data and identify defects on the product surface. For example, it obtains a result such as "Small scratches on the product surface detected at region (x: 200, y: 350)."

[0477] Step 4: Notification of analysis results and feedback

[0478] The server notifies the device in real time of analysis results and optimization suggestions, allowing the user to respond, and also receives user feedback to continuously improve the analysis algorithm.

[0479] Input: Analysis results and optimization proposals

[0480] Data processing: notification generation, feedback processing

[0481] Output: Notifying field workers, collecting feedback data

[0482] Specific behavior:

[0483] The server then sends a notification to the on-site worker's tablet regarding any abnormal sounds detected as a result of the analysis, such as "Abnormal sound detected at timestamp 00:01:15."

[0484] The user checks the notification displayed on the tablet and performs the necessary maintenance work.

[0485] Users input the results and feedback after maintenance work into a tablet and send it to the server, which uses this feedback to improve its analysis algorithm.

[0486] Through these steps, the system can improve the efficiency and quality of the production line in real time.

[0487] (Application example 1)

[0488] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0489] Improving production efficiency and quality is a key challenge in modern manufacturing. However, conventional systems have difficulty analyzing data in real time and responding immediately, making early detection of abnormalities and process optimization insufficient. This makes it difficult to minimize production line downtime and maintain a high level of quality.

[0490] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0491] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality evaluation, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for improving the analysis algorithm based on feedback from the workers, means for robots installed on the production line to collect operation data in real time and send it to the server, means for the robots to autonomously optimize themselves based on the analysis results from the server, and means for notifying on-site workers of the analysis results so that they can make manual adjustments. This enables real-time anomaly detection, quality evaluation, and process optimization, thereby improving the efficiency and quality of the production line.

[0492] The "various sensors" are various types of sensors used in manufacturing sites, and are devices that have the function of collecting text data, audio data, image data, and video data.

[0493] "Text data" is data that expresses the state of a machine or environment as numbers or strings of characters, and is primarily information indicating temperature, pressure, speed, etc.

[0494] "Audio data" refers to data collected as acoustic signals from sounds generated from production lines and machines.

[0495] "Image data" is digital data of still images captured by an imaging device such as a camera, and represents the state of a product or production line.

[0496] "Video data" refers to digital data of video captured by a camera or video device, and is data that includes a series of images.

[0497] "Preprocessing" refers to the process of converting collected data into a format suitable for analysis, and includes noise removal and resolution unification.

[0498] A "database" is a computer system for temporarily storing preprocessed data.

[0499] "Analysis" is the process of applying calculations and models to pre-processed data to detect anomalies, assess quality, and suggest process optimization.

[0500] "Abnormality detection" refers to detecting abnormal conditions that differ from normal operating conditions, such as detecting abnormal sounds or vibrations in machinery.

[0501] "Quality evaluation" is the process of evaluating the quality of a product or process to determine whether it meets specified standards.

[0502] "Process optimization" is the process of finding the optimal parameters and procedures in a production process to improve overall efficiency.

[0503] "Real-time" means that the system collects data and analyzes and notifies immediately, with little time delay.

[0504] "Feedback" refers to information that collects opinions and evaluations from workers and system users and is used to improve the system.

[0505] A "server" is a central processing unit that collects, stores, analyzes, and processes feedback data on a network.

[0506] A "robot" is a mechanical device that is installed in a production line and can operate autonomously.

[0507] A "terminal" is a device such as a tablet or smartphone used by a field worker to receive notifications from the server.

[0508] This invention is an AI-driven robotics assistance system for improving production efficiency and quality in the manufacturing industry. The system collects data in real time from various sensors installed on the production line, and a server analyzes the data to detect anomalies and optimize processes.

[0509] Hardware and software used

[0510] Hardware:

[0511] Sensors: Temperature sensors, pressure sensors, camera sensors, audio sensors, etc. These sensors are installed at various points along the production line to collect necessary data in real time.

[0512] Robot: A machine with autonomous capabilities installed on a production line that receives data from sensors and performs the required actions.

[0513] Terminal: A device such as a tablet or smartphone used by field workers. It receives notifications from the server and allows data to be checked and manually adjusted.

[0514] software:

[0515] Server (Central Processing Unit): Collects, pre-processes, analyzes, and processes feedback data.

[0516] OpenCV: A library for capturing and processing camera images, and is used to analyze image data.

[0517] TensorFlow: A library that performs data analysis using AI models, and is used for anomaly detection and quality assessment.

[0518] Requests: A data communication library used to send data from sensors and robots to a server.

[0519] Specific explanation for carrying out the invention

[0520] First, various sensors installed on the production line collect temperature, pressure, image, audio, and video data in real time. The collected data is immediately sent to a server, where it is preprocessed, for example, to remove noise from audio data and standardize the resolution of image data.

[0521] Once preprocessed, the data is analyzed using TensorFlow, which generates anomaly detection, quality assessment, and process optimization recommendations. The analysis results and optimization recommendations are sent to on-site workers' devices in real time. Workers can view the analysis results on their tablets or smartphones and take any necessary action.

[0522] Furthermore, robots installed on the production line collect data from various sensors in real time and send it to a server. The robots receive the analysis results from the server and can then optimize themselves autonomously. This system makes it possible to minimize production line downtime and maintain a high level of quality.

[0523] Specific examples (prompt sentence examples)

[0524] Below is an example of a prompt sentence to input to the generative AI model.

[0525] Create a program to analyze sensor data (images, temperature, pressure, etc.) collected by factory robots in real time, detect anomalies, and propose process optimization. Include a function for the robot to autonomously adjust itself based on the analysis results. Use OpenCV for image analysis, TensorFlow for AI analysis, and the Requests library for data communication.

[0526] According to the aspects of the present invention, real-time anomaly detection, quality evaluation, and process optimization become possible, thereby realizing improved efficiency and quality of the production line.

[0527] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0528] Step 1:

[0529] The sensors collect temperature, pressure, image, and audio data from the production line in real time. These data are collected in the form of sensors that each have. The input is the environmental condition of the production line, and the output is the raw data obtained from each sensor.

[0530] Step 2:

[0531] The server preprocesses the collected data, specifically removing noise from audio data, standardizing the resolution of image data, and standardizing other data formats. The input is raw data sent from the sensor, and the output is preprocessed data.

[0532] Step 3:

[0533] The server analyzes the preprocessed data. This analysis uses TensorFlow to perform anomaly detection, quality assessment, and process optimization recommendations. The analysis requires advanced computations and runs models to identify abnormal sounds and surface defects. The input is the preprocessed data, and the output is the analysis results and recommendations.

[0534] Step 4:

[0535] The server notifies the on-site worker of the analysis results and optimization proposals in real time. The worker's device is a tablet or smartphone, where they can check the analysis results and take any necessary measures. The input is the analysis results and proposals, and the output is the notification received by the on-site worker.

[0536] Step 5:

[0537] The user receives the notification and sends feedback from the device to the server. The feedback includes opinions about the analysis results and specific adjustments. The input is the user's feedback on the notification, and the output is the feedback data sent to the server.

[0538] Step 6:

[0539] The server improves the analysis algorithm based on the received feedback. This improves the accuracy of the analysis from the next time onwards, and further increases in production efficiency are expected. The input is feedback data from the user, and the output is a revised analysis algorithm.

[0540] Step 7:

[0541] The robot collects operational data in real time from various sensors installed on the production line and sends it to a server. The input is the operational data collected by the robot, and the output is the data sent to the server.

[0542] Step 8:

[0543] The server analyzes the operation data from the robot and feeds the results back to the robot. The robot receives the analysis results and performs optimization autonomously. The input is the robot's operation data and analysis results, and the output is the robot's optimized behavior.

[0544] This series of processing flows enables real-time monitoring and optimization of the entire production line.

[0545] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0546] This invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry, incorporating an emotion engine that recognizes the user's emotional state in real time and reflects it in optimization suggestions.

[0547] This system includes a server, various sensors, user devices, and an emotion engine. Sensors placed on the production line collect text data (numerical information such as temperature and pressure), audio data (machine noise, user voice), image data (photos of products, user facial expressions), and video data (video of the production line in operation) in real time. This data is sent to the server and managed centrally.

[0548] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, processes such as removing noise from audio data and standardizing the resolution of image data are performed. Next, the preprocessed data is analyzed using an AI model. This generates suggestions for anomaly detection, quality assessment, and process optimization.

[0549] Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotional state in real time. The emotion engine analyzes the user's emotions from voice and image data, detecting, for example, stress and fatigue. The analysis results are reflected in the server's analysis algorithm, resulting in the generation of more appropriate optimization suggestions.

[0550] The analysis results and optimization suggestions are sent to the device in real time. This allows the user (field worker) to check the analysis results and optimization suggestions using a tablet or smartphone and take any necessary action. For example, if the user is in a high-stress state, the notification may include a suggestion to take a break or redistribute tasks. User feedback is also analyzed through the emotion engine and used to improve the system as a whole.

[0551] The following is a specific example.

[0552] ---

[0553] Example 1: User stress detection

[0554] The server uses audio and image sensors to detect the user's stress level from their tone of voice and facial expressions. If the emotion engine analyzes the user's emotional state as "high stress," it sends the result to the server. Based on this information, the server notifies the device with suggestions for reducing the user's workload (e.g., taking a break or switching to a lighter task). This maximizes the user's performance while maintaining their physical and mental health.

[0555] ---

[0556] Example 2: Emotional analysis feedback

[0557] If a user feels dissatisfied or stressed during product inspection, the emotion engine detects that emotional state. The server incorporates this information into the analysis results and receives it as feedback. For example, if the system detects that the user is "tired and prone to making judgment errors," it will adjust the sensitivity of the analysis algorithm next time and strengthen settings to minimize human error.

[0558] ---

[0559] Example 3: Dynamic adjustment of operating environment

[0560] The server analyzes data from the entire production line and the user's emotional data to suggest dynamic adjustments to the operating environment. For example, by appropriately changing the temperature settings for the entire line or adjusting the operating speed of specific machines, overall efficiency can be improved. When making these suggestions based on the user's emotional data, the emotion engine also takes into account whether the working environment is comfortable.

[0561] ---

[0562] This system will enable manufacturing production lines to operate with even greater efficiency and quality. By combining an emotion engine with AI analysis, appropriate suggestions that take into account the user's condition will be made in real time, enabling more effective optimization of the production process.

[0563] The processing flow will be explained below.

[0564] Step 1:

[0565] The server collects text data, audio data, image data, and video data in real time from various sensors, including text data from temperature and pressure sensors, audio data from microphones, and image and video data from cameras.

[0566] Step 2:

[0567] The server preprocesses the collected data, removing noise from the audio data, standardizing the resolution of the image data, and aligning the timestamps of all data. During this process, the data is temporarily stored in a database.

[0568] Step 3:

[0569] The server then inputs the preprocessed data into an AI model for analysis. During the analysis stage, machine learning algorithms are used to generate anomaly detection, quality assessment, and process optimization recommendations. For example, an anomaly detection model can be used to detect abnormal sounds from audio data, and an image analysis algorithm can be used to automatically detect defects on the surface of a product.

[0570] Step 4:

[0571] The server notifies the user's device of the analysis results and optimization proposals in real time, allowing the device to allow on-site workers to immediately check the analysis results and optimization proposals.

[0572] Step 5:

[0573] The device equipped with the emotion engine analyzes the user's voice tone and facial expression data in real time to evaluate the user's emotional state (e.g., stress, fatigue). Specifically, it detects changes in tone and speed of voice from the voice data, and senses changes in facial expression from the image data.

[0574] Step 6:

[0575] The server receives the emotional state data sent from the emotion engine. Based on this information, the server adjusts the analysis algorithm and generates optimization suggestions according to the user's emotional state. For example, if the user is in a high-stress state, the server generates suggestions for taking a break or reducing workload.

[0576] Step 7:

[0577] The server then sends the newly generated optimization proposals to the user's device, where the user can view these proposals in real time. The user can then use the device to input feedback, such as accepting the proposals or requesting modifications.

[0578] Step 8:

[0579] The server receives feedback data from users and continuously improves the analysis algorithm. The feedback, including emotional data, is analyzed to improve the accuracy of the next data analysis.

[0580] ---

[0581] This system monitors and optimizes manufacturing production lines in real time, continuously improving production efficiency and quality. The introduction of an emotion engine enables flexible responses according to the user's condition, achieving both an improved work environment and increased production efficiency.

[0582] Example 2

[0583] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0584] Improving production efficiency and quality are important challenges in modern manufacturing. However, conventional systems focused on collecting and analyzing data from sensors, and did not fully consider the impact of workers' emotional states. As a result, poor performance and mistakes caused by worker stress and fatigue had a negative impact on production efficiency and quality. This made it difficult to manage worker health and optimize production processes.

[0585] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0586] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality assessment, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for analyzing the emotional state of the workers from their voice data and image data, means for generating optimization proposals based on their emotional states, and means for improving the analysis algorithm based on feedback from the workers. This improves production efficiency and quality and enables optimal proposals that take into account the health status of the workers.

[0587] A "sensor" is a device that collects different types of data in real time, such as temperature, pressure, sound, images, and video.

[0588] "Data preprocessing" refers to the process of performing preprocessing such as noise removal and resolution standardization on collected data, and then temporarily storing it.

[0589] A "database" is a storage device for temporarily storing preprocessed data.

[0590] An "AI model" is an artificial intelligence algorithm that performs analysis for anomaly detection, quality evaluation, and process optimization.

[0591] "Analysis results" are the results of anomaly detection and quality assessment generated by the AI ​​model using preprocessed data.

[0592] An "optimization proposal" is a proposal for improving the production process that is generated based on the analysis results.

[0593] "Field workers" are workers who work on the manufacturing floor and act based on data collected from sensors.

[0594] "Feedback" refers to opinions and reactions to the system analysis and suggestions provided by field workers.

[0595] "Emotional state" refers to the emotional state of a worker, such as stress or fatigue, analyzed from the worker's voice data and image data.

[0596] The "emotion engine" is a system that analyzes the emotional state of workers from collected voice and image data.

[0597] "Notification" is a communication method by which the server conveys analysis results and optimization suggestions to field workers in real time.

[0598] An "analysis algorithm" is a calculation procedure for detecting anomalies and evaluating quality based on collected data.

[0599] This invention is an AI-driven robot support system for improving production efficiency and quality in the manufacturing industry, incorporating an emotion engine that recognizes the user's emotional state in real time and reflects it in optimization proposals. This system is composed of a server, various sensors, user terminals, and the emotion engine.

[0600] The server is the central part that collects data in real time from sensors placed on the production line and performs preprocessing and analysis. The main hardware used is a high-performance server computer, and the software includes Python, TensorFlow, PyTorch, OpenCV, Librosa, etc.

[0601] Data collection

[0602] The server uses various sensors to collect temperature, pressure, audio, image, and video data. For example, the temperature sensor measures the temperature data of the line in real time, the microphone collects user voices and machine sounds, and the image sensor takes photos of the product and records video of the production line operation.

[0603] Data Preprocessing

[0604] The collected data is sent to a server where preprocessing such as noise removal and resolution unification is performed. For example, the Librosa library is used to remove environmental noise from audio data, and the resolution of image data is unified using OpenCV. This organizes the data into a form optimal for analysis.

[0605] Data analysis

[0606] The preprocessed data is then analyzed by the server using AI models. Anomaly detection models are run using TensorFlow and PyTorch. Quality assessments and recommendations for optimizing the production process are also generated at this stage. If an anomaly is detected, its details are recorded in a log.

[0607] Emotion analysis

[0608] Furthermore, the system incorporates an emotion engine that recognizes the user's emotional state in real time. The emotion engine uses natural language processing models (e.g., BERT) and facial recognition models (e.g., DeepFace) to analyze the user's emotions from voice and image data, thereby determining whether the user is in a stressful state.

[0609] Proposal generation and notification

[0610] Optimization suggestions are generated based on the analysis results and the user's emotional state. The server notifies the user of these suggestions in real time. For example, if a user is in a high-stress state, the device may be notified of suggestions such as taking a break or redistributing tasks. The device application uses React Native or Android Studio, allowing users to view these notifications on their tablets or smartphones.

[0611] Feedback collection and system improvement

[0612] Users can provide feedback on the suggestions provided to them through their devices. The server collects this feedback and uses it to improve the analysis algorithm. This allows the system to learn from the feedback and improve the accuracy of future suggestions.

[0613] Example: Detecting user stress

[0614] The server uses audio and image sensors to detect the user's stress level from their tone of voice and facial expressions. If the emotion engine analyzes the user's emotional state as "high stress," it sends the result to the server. Based on this information, the server notifies the device with suggestions for reducing the user's workload (e.g., taking a break or switching to a lighter task). This maximizes the user's performance while maintaining their physical and mental health.

[0615] Example prompt sentence:

[0616] "Analyze the user's stress level based on their voice and image data, and make optimization suggestions based on the results."

[0617] This system will enable manufacturing production lines to operate with even greater efficiency and quality. By combining an emotion engine with AI analysis, appropriate suggestions that take into account the user's condition will be made in real time, enabling more effective optimization of the production process.

[0618] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0619] Step 1: Data collection

[0620] Inputs: temperature, pressure, audio, image, and video data.

[0621] Processing: The server collects data in real time from various sensors placed on the production line.

[0622] Output: Collected multimodal data.

[0623] How it works: The server collects real-time temperature data every 10 seconds from the temperature sensor, and collects user voice and machine sounds from the microphone. At the same time, the image sensor captures the product's appearance, and the video sensor records the operation of the production line.

[0624] Step 2: Data Preprocessing

[0625] Input: Collected multimodal data.

[0626] Processing: The server preprocesses the collected data, removing noise and unifying the resolution for temporary storage.

[0627] Output: Preprocessed data.

[0628] Specific operation: Noise reduction is performed on the audio data using the Librosa library, and the resolution of the image data is unified using OpenCV. The preprocessed data is temporarily stored in a database.

[0629] Step 3: Data analysis

[0630] Input: Preprocessed data.

[0631] Processing: The server uses the preprocessed data to perform analysis using an AI model.

[0632] Output: Analysis results (anomaly detection, quality assessment, process optimization suggestions).

[0633] What it does: Runs anomaly detection models using TensorFlow or PyTorch to generate quality assessments and process optimization suggestions. If an anomaly is detected, details of the anomaly are logged.

[0634] Step 4: Sentiment Analysis

[0635] Input: Preprocessed audio and image data.

[0636] Processing: The server analyzes the user's emotional state using an emotion engine.

[0637] Output: Emotion analysis results (e.g. high stress, fatigue).

[0638] Specific operation: Emotions are analyzed from voice data using BERT, and facial expressions are analyzed from image data using DeepFace. The analysis results of the emotional state are stored on the server.

[0639] Step 5: Proposal Generation

[0640] Input: Data analysis results and sentiment analysis results.

[0641] Processing: The server generates optimization suggestions based on the data analysis results and sentiment analysis results.

[0642] Output: Optimization suggestions (e.g. break suggestions, task redistribution).

[0643] Specific operation: Based on the analysis results, suggestions to reduce the user's workload are automatically generated and prepared for notification in the next step.

[0644] Step 6: Notification of results

[0645] Input: Optimization proposal.

[0646] Processing: The server notifies the user terminal of the generated optimization proposal.

[0647] Output: A notification message on the terminal.

[0648] Specific operation: The server pushes JSON format data to the device, and the application on the device interprets the data and displays it to the user.

[0649] Step 7: Gather feedback and improve the system

[0650] Input: User feedback.

[0651] Processing: The user enters feedback on the suggestions provided, and the server collects the feedback and uses it to improve the system.

[0652] Output: Feedback data, improved AI model.

[0653] How it works: Users input feedback from their devices, and the server collects that information and reflects it in the next analysis, thereby improving the accuracy of the analysis algorithm.

[0654] These steps not only improve the efficiency and quality of the production process, but also optimize the working environment, maintaining worker health and performance through suggestions that take the user's emotional state into account.

[0655] (Application example 2)

[0656] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0657] Conventional production systems focus on detecting machine anomalies and evaluating quality, but rarely consider the emotional state of human workers. This increases the risk of workers feeling stressed or fatigued, leading to reduced production efficiency. Furthermore, the lack of work instructions or improvement suggestions based on the emotional state of workers poses a challenge, making it difficult to fully optimize the entire production process.

[0658] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0659] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality assessment, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for improving the analysis algorithm based on feedback from the workers, and means for analyzing the emotions of the workers in real time and providing work instructions based on their emotional states. This makes it possible to provide optimal work instructions and propose process improvements that take the emotional states of the workers into consideration, thereby achieving both improved production efficiency and maintaining the health of the workers.

[0660] "Multiple sensors" are various sensors used to collect text data, audio data, image data, and video data in real time.

[0661] "Preprocessing" refers to the process of converting collected data into a form suitable for analysis, and specifically includes noise removal and resolution unification.

[0662] A "database" is a system for temporarily storing preprocessed data.

[0663] "Analysis algorithms" are techniques for analyzing pre-processed data and generating suggestions for anomaly detection, quality assessment, and process optimization.

[0664] "Emotional state" refers to the emotional state of the worker analyzed from voice and image data, and specifically refers to stress and fatigue.

[0665] "Work instructions" refers to specific work tasks and break suggestions provided to workers based on their emotional state.

[0666] "Feedback" refers to responses and reactions from workers, and is information collected to help improve the system.

[0667] An "emotion analysis device" is a piece of equipment used to analyze the emotional state of a worker in real time.

[0668] "Production process improvement proposals" refer to specific proposals for improving production efficiency based on the analysis results.

[0669] "Stress state" refers to the degree of tension and strain felt by the worker, and is analyzed by the emotion engine.

[0670] To realize the present invention, the following detailed description of the system and its components is required. The system of the present invention collects and analyzes data in real time and provides appropriate feedback and work instructions. Specific components include a server, various sensors, a user's device, and an emotion analysis engine.

[0671] server

[0672] The server is the core processing unit of the system and performs the following functions:

[0673] Data collection: Collect text, audio, image, and video data in real time from a variety of sensors.

[0674] Data preprocessing: The collected data is preprocessed by noise removal and resolution standardization, and then temporarily stored in a database.

[0675] Data Analysis: The pre-processed data is analyzed using generative AI models to generate recommendations for anomaly detection, quality assessment, and process optimization.

[0676] Emotion analysis: An emotion analysis engine is used to analyze the user's emotional state from voice and image data.

[0677] Various sensors

[0678] The system uses the following sensors:

[0679] Audio sensor: Collects the user's voice tone and mechanical sounds to detect stress levels and abnormal sounds.

[0680] Image sensor: Collects footage of the user's facial expressions and production line operations to inspect the emotional state and appearance of the product.

[0681] User's device

[0682] The user's device (smartphone, tablet, etc.) performs the following functions:

[0683] Notifications: Analysis results and optimization suggestions from the server are displayed in real time.

[0684] Feedback: Collects feedback from users and sends it to the server.

[0685] Sentiment Analysis Engine

[0686] The emotion analysis engine analyzes the user's emotional state from voice data and image data and provides the data to the server.

[0687] This is implemented using the following software and libraries:

[0688] OpenCV: Face Recognition and Landmark Detection

[0689] dlib: Facial landmark analysis

[0690] Keras: Sentiment Analysis Model

[0691] requests: Send notification

[0692] Specific examples

[0693] For example, when developing an application to detect the stress level of workers at a logistics center, the smartphone's camera and microphone can be used to analyze the worker's facial expressions and tone of voice. If the worker's emotional state is determined to be "high stress," the server will send a notification to take a break. In this way, optimal work instructions that take the worker's emotions into consideration can be provided, thereby improving production efficiency and maintaining the worker's health at the same time.

[0694] Prompt Sentence Examples

[0695] Develop an application that evaluates workers' emotions through real-time analysis of camera footage and suggests breaks when they are stressed. The technologies used are OpenCV, dlib, and Keras. The goal is to analyze emotions from workers' facial expressions and suggest breaks when stress or fatigue is detected.

[0696] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0697] Step 1:

[0698] The server collects text data, voice data, image data, and video data in real time from various sensors (audio sensors, image sensors). The input is data from the sensors, and the output is raw data for preprocessing. Specifically, the audio sensor collects the tone of the user's voice, and the image sensor captures the user's facial expressions and footage of the production line operation.

[0699] Step 2:

[0700] The server preprocesses the collected data. Specifically, it removes noise from the audio data and standardizes the resolution of the image data. At this point, the input is raw data and the output is preprocessed data. The preprocessed data is temporarily stored in a database.

[0701] Step 3:

[0702] The server inputs the preprocessed data into the generative AI model for analysis. The input is the preprocessed data, and the output of the analysis is anomaly detection, quality assessment, and process optimization recommendations. Specifically, the AI ​​model uses image data to perform visual inspections of products and analyzes audio data to detect abnormal machine sounds.

[0703] Step 4:

[0704] The server analyzes the user's emotional state using an emotion analysis engine. The input is audio and image data, and the output is the user's emotional state (stress or fatigue). Specifically, it uses OpenCV and dlib to detect facial landmarks, and classifies the emotional state using an emotion analysis model using Keras.

[0705] Step 5:

[0706] The server notifies the user device of the analysis results and optimization suggestions in real time. The input is the analysis results and sentiment analysis results, and the output is a notification message. Specifically, it uses the requests library to send a message to the notification API, which is then displayed on the user device. For example, if a high stress state is detected, a notification such as "Take a break" is sent.

[0707] Step 6:

[0708] The user checks the suggestions through the terminal and takes the necessary action. The input is a notification message and the output is feedback information. The suggestions are displayed on the terminal, and the user can check them and take action.

[0709] Step 7:

[0710] The terminal collects feedback from the user and sends it to the server. The input is the feedback information from the user, and the output is the feedback data sent to the server. Specifically, an interface is provided for the user to report task completion notifications and new problems.

[0711] Step 8:

[0712] The server improves the analysis algorithm based on the collected feedback. The input is the feedback data, and the output is an improved analysis algorithm. Specifically, it reevaluates the analysis results and adjusts the model parameters.

[0713] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0714] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0715] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0716] [Third embodiment]

[0717] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0718] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0719] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0720] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0721] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0722] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0723] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0724] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0725] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0726] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0727] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0728] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0729] The present invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry, and is specifically implemented as follows.

[0730] The central roles of the system are played by a server, various sensors, and terminals (used by users). Various sensors placed on the production line collect text data (numerical information such as temperature and pressure), audio data (machine sounds), image data (photos of products), and video data (video of the production line in operation) in real time. This data is sent to the server and managed centrally.

[0731] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, processes such as removing noise from audio data and standardizing the resolution of image data are performed. Next, the preprocessed data is analyzed using an AI model. This generates suggestions for anomaly detection, quality assessment, and process optimization.

[0732] Analysis results and optimization suggestions are sent to the terminal in real time, allowing the user (field worker) to check the analysis results and take necessary measures using a tablet or smartphone. Users can also enter feedback, which allows the server to continuously improve the analysis algorithm.

[0733] A specific example is shown below.

[0734] ---

[0735] Example 1: Detecting abnormal sounds from machinery

[0736] The server analyzes machine sound data collected from the audio sensor in real time. If an abnormal sound that differs from normal operating sounds is detected, the server notifies the on-site worker's device of the occurrence of the abnormal sound. The user receives this notification, checks the machine's status, and performs maintenance work if necessary. This makes it possible to minimize production line downtime.

[0737] ---

[0738] Example 2: Automatic detection of surface defects on products

[0739] The server analyzes product images collected from the camera sensor and automatically detects small scratches or stains on the surface. If an abnormality is detected, the image data and the location of the abnormality are notified to the user's device. The user can check the notification and take action to re-inspect or correct the product. This makes it possible to maintain a high level of product quality.

[0740] ---

[0741] Example 3: Process optimization proposal

[0742] The server collects and analyzes numerical data related to the production process, such as temperature, pressure, and speed. Based on the analysis results, it suggests fine-tuning the temperature or pressure, for example. These suggestions are sent to the user's device, and if the user accepts the suggestions, the settings are automatically changed. This improves the efficiency of the entire production process and ensures optimal resource utilization.

[0743] ---

[0744] This system enables real-time monitoring and optimization of manufacturing production lines, dramatically improving production efficiency and quality. By effectively linking data collection, analysis, and feedback, on-site workers and robots can work together to achieve more efficient, higher-quality production.

[0745] The processing flow will be explained below.

[0746] Step 1:

[0747] The server collects text data, audio data, image data, and video data in real time from various sensors, including text data from temperature and pressure sensors, audio data from microphones, and image and video data from cameras.

[0748] Step 2:

[0749] The server preprocesses the collected data. Preprocessing includes data cleansing (such as noise removal), format conversion, and timestamp synchronization. For example, it removes background noise from audio data, standardizes the resolution of image data, and aligns the timestamps of all data.

[0750] Step 3:

[0751] The server then inputs the preprocessed data into the AI ​​model for analysis. During the analysis phase, machine learning algorithms are used to generate anomaly detection, quality assessment, and process optimization recommendations. For example, an anomaly detection model can be used to detect abnormal sounds from audio data, and an image analysis algorithm can be used to automatically detect defects on the surface of a product.

[0752] Step 4:

[0753] The server sends analysis results and optimization proposals in real time to the devices used by field workers, such as tablets or smartphones, which the workers use to check the analysis results and optimization proposals.

[0754] Step 5:

[0755] Users can use their devices to check the analysis results and optimization suggestions, and provide feedback as needed. For example, they can check abnormality detection notifications, inspect the machine's status, restart it, or perform maintenance.

[0756] Step 6:

[0757] The server receives feedback data from users and improves the analysis algorithm, which will enable more accurate data analysis the next time. Specifically, the feedback data is used to adjust the parameters of the machine learning model, improving the accuracy of anomaly detection and quality assessment.

[0758] Example 1

[0759] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0760] Improving production efficiency and quality is an important issue in the manufacturing industry. In particular, there is a demand for the detection of abnormal sounds, early detection of surface defects on products, and real-time process optimization. However, these issues have not been adequately resolved with conventional methods. A system that reduces the burden on on-site workers and provides highly accurate data analysis and notifications in real time is needed.

[0761] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0762] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for inputting the preprocessed data into a generative AI model for analysis, means for notifying the mobile devices of on-site workers of analysis results and improvement suggestions in real time, and means for receiving feedback from workers and improving the analysis algorithm. This enables highly accurate anomaly detection, quality evaluation, and process optimization in real time, dramatically improving production efficiency and quality while reducing the burden on on-site workers.

[0763] A "sensor" is a device that detects physical or environmental data in real time and outputs that information as a signal.

[0764] "Text data" is data that expresses numerical information such as temperature, pressure, and speed as a string of characters.

[0765] "Audio data" refers to sound wave information recorded in digital format. Specifically, it refers to audio information such as machine operating sounds and abnormal sounds.

[0766] "Image data" is data that digitally represents still images captured by a camera or sensor.

[0767] "Video data" is a series of images recorded along a time axis, allowing for visual capture of movements and changes in processes.

[0768] "Preprocessing" refers to processes such as filtering, noise removal, and format standardization that are performed to improve the quality of collected data.

[0769] A "database" is a storage device or system for temporarily storing preprocessed data.

[0770] A "generative AI model" is an algorithm or software that uses artificial intelligence techniques to perform data analysis.

[0771] "Analysis" is a series of procedures for anomaly detection, quality assessment, and process optimization based on collected and pre-processed data.

[0772] "Notification" refers to the act of communicating analysis results and improvement suggestions to the mobile devices of field workers in real time.

[0773] "Feedback" is input from field workers that is used to continually improve the analysis algorithms.

[0774] An "algorithm" is a collection of computational procedures or mathematical formulas for the purpose of data analysis or process optimization.

[0775] A "field worker" is a user who is in charge of work on the production line and receives notifications and suggestions from the system.

[0776] The present invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry. Specifically, it is configured as follows.

[0777] The central roles of the system are played by the server, various sensors, and terminals (used by users). Various sensors (temperature sensors, audio sensors, camera sensors, etc.) placed on the production line collect numerical information (text data) such as temperature and pressure, machine sounds (audio data), product photos (image data), and video of the production line in operation (video data) in real time. This data is sent to the server and managed centrally.

[0778] Data Preprocessing

[0779] The server preprocesses the collected data and temporarily stores it. This preprocessing involves, for example, removing noise from audio data and standardizing the resolution of image data. The technologies used here include audio filtering and image processing algorithms.

[0780] Data analysis

[0781] The preprocessed data is then fed into a generative AI model for analysis. The AI ​​model uses TensorFlow or PyTorch, among others. The analysis generates anomaly detection, quality assessment, and recommendations for process optimization. For example, it can detect abnormal sounds from audio data or identify surface defects on products from image data.

[0782] Notification and feedback of analysis results

[0783] Analysis results and optimization suggestions are sent to the terminal in real time, allowing users (field workers) to check the analysis results using a tablet or smartphone and take any necessary measures. Users can also enter feedback, which allows the server to continuously improve the analysis algorithm.

[0784] Example 1: Detecting abnormal sounds from machinery

[0785] The server analyzes machine sound data collected from the audio sensor in real time. If an abnormal sound that differs from normal operating sounds is detected, the server notifies the on-site worker's device of the occurrence of the abnormal sound. The user receives this notification, checks the machine's status, and performs maintenance work if necessary. This makes it possible to minimize production line downtime.

[0786] Example 2: Automatic detection of surface defects on products

[0787] The server analyzes product images collected from the camera sensor and automatically detects small scratches or stains on the surface. If an abnormality is detected, the image data and the location of the abnormality are notified to the user's device. The user can check the notification and take action to re-inspect or correct the product. This makes it possible to maintain a high level of product quality.

[0788] Example 3: Process optimization proposal

[0789] The server collects and analyzes numerical data related to the production process, such as temperature, pressure, and speed. Based on the analysis results, it suggests fine-tuning the temperature or pressure, for example. These suggestions are sent to the user's device, and if the user accepts the suggestions, the settings are automatically changed. This improves the efficiency of the entire production process and ensures optimal resource utilization.

[0790] Prompt Sentence Examples

[0791] The server should analyze the audio data in real time and detect any abnormal sounds.

[0792] "The server should analyze the product image data and automatically detect surface defects."

[0793] "The server should analyze production process data such as temperature and pressure and suggest optimal settings."

[0794] This system enables real-time monitoring and optimization of manufacturing production lines, dramatically improving production efficiency and quality. A series of processes—data collection, preprocessing, analysis, notification, and feedback—work together to enable on-site workers and robots to work together and achieve highly efficient, high-quality production.

[0795] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0796] Step 1: Data collection

[0797] The server collects data in real time from various sensors installed on the production line, such as temperature sensors, audio sensors, and camera sensors. This data includes text data on temperature and pressure, audio data on machine operating sounds and abnormal sounds, and product images and video data.

[0798] Input: Real-time data from various sensors

[0799] Data processing: receiving and storing sensor data

[0800] Output: Collected raw sensor data

[0801] Specific behavior:

[0802] The server obtains temperature data from the temperature sensor every 30 seconds. For example, it collects data such as "Temperature: 27.5°C, 28.0°C, 27.8°C."

[0803] Audio data is streamed from the audio sensor every second to capture the sound of the machine operating.

[0804] The camera sensor captures images of the product every minute and collects video data every five minutes.

[0805] Step 2: Data Preprocessing

[0806] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, noise is removed from the audio data and the resolution of the image data is standardized.

[0807] Input: Collected raw sensor data

[0808] Data processing: noise filtering, resolution unification, format conversion, etc.

[0809] Output: Pre-processed, high-quality data

[0810] Specific behavior:

[0811] The server applies a noise filter to the audio data to generate clear audio data.

[0812] The server standardizes the resolution of image data to a uniform 1024x768 pixels and converts it from JPEG format to PNG format.

[0813] Step 3: Data analysis

[0814] The server inputs the preprocessed data into a generative AI model for analysis, which generates recommendations for anomaly detection, quality assessment, and process optimization. The AI ​​model uses TensorFlow and PyTorch.

[0815] Input: Preprocessed data

[0816] Data processing: Data analysis using AI models

[0817] Output: Analysis results and optimization suggestions

[0818] Specific behavior:

[0819] The server uses TensorFlow to detect abnormal sounds from the preprocessed audio data, and obtains a result such as "abnormal sound detected at timestamp '00:01:15'".

[0820] The server uses PyTorch to analyze the image data and identify defects on the product surface. For example, it obtains a result such as "Small scratches on the product surface detected at region (x: 200, y: 350)."

[0821] Step 4: Notification of analysis results and feedback

[0822] The server notifies the device in real time of analysis results and optimization suggestions, allowing the user to respond, and also receives user feedback to continuously improve the analysis algorithm.

[0823] Input: Analysis results and optimization proposals

[0824] Data processing: notification generation, feedback processing

[0825] Output: Notifying field workers, collecting feedback data

[0826] Specific behavior:

[0827] The server then sends a notification to the on-site worker's tablet regarding any abnormal sounds detected as a result of the analysis, such as "Abnormal sound detected at timestamp 00:01:15."

[0828] The user checks the notification displayed on the tablet and performs the necessary maintenance work.

[0829] Users input the results and feedback after maintenance work into a tablet and send it to the server, which uses this feedback to improve its analysis algorithm.

[0830] Through these steps, the system can improve the efficiency and quality of the production line in real time.

[0831] (Application example 1)

[0832] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0833] Improving production efficiency and quality is a key challenge in modern manufacturing. However, conventional systems have difficulty analyzing data in real time and responding immediately, making early detection of abnormalities and process optimization insufficient. This makes it difficult to minimize production line downtime and maintain a high level of quality.

[0834] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0835] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality evaluation, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for improving the analysis algorithm based on feedback from the workers, means for robots installed on the production line to collect operation data in real time and send it to the server, means for the robots to autonomously optimize themselves based on the analysis results from the server, and means for notifying on-site workers of the analysis results so that they can make manual adjustments. This enables real-time anomaly detection, quality evaluation, and process optimization, thereby improving the efficiency and quality of the production line.

[0836] The "various sensors" are various types of sensors used in manufacturing sites, and are devices that have the function of collecting text data, audio data, image data, and video data.

[0837] "Text data" is data that expresses the state of a machine or environment as numbers or strings of characters, and is primarily information indicating temperature, pressure, speed, etc.

[0838] "Audio data" refers to data collected as acoustic signals from sounds generated from production lines and machines.

[0839] "Image data" is digital data of still images captured by an imaging device such as a camera, and represents the state of a product or production line.

[0840] "Video data" refers to digital data of video captured by a camera or video device, and is data that includes a series of images.

[0841] "Preprocessing" refers to the process of converting collected data into a format suitable for analysis, and includes noise removal and resolution unification.

[0842] A "database" is a computer system for temporarily storing preprocessed data.

[0843] "Analysis" is the process of applying calculations and models to pre-processed data to detect anomalies, assess quality, and suggest process optimization.

[0844] "Abnormality detection" refers to detecting abnormal conditions that differ from normal operating conditions, such as detecting abnormal sounds or vibrations in machinery.

[0845] "Quality evaluation" is the process of evaluating the quality of a product or process to determine whether it meets specified standards.

[0846] "Process optimization" is the process of finding the optimal parameters and procedures in a production process to improve overall efficiency.

[0847] "Real-time" means that the system collects data and analyzes and notifies immediately, with little time delay.

[0848] "Feedback" refers to information that collects opinions and evaluations from workers and system users and is used to improve the system.

[0849] A "server" is a central processing unit that collects, stores, analyzes, and processes feedback data on a network.

[0850] A "robot" is a mechanical device that is installed in a production line and can operate autonomously.

[0851] A "terminal" is a device such as a tablet or smartphone used by a field worker to receive notifications from the server.

[0852] This invention is an AI-driven robotics assistance system for improving production efficiency and quality in the manufacturing industry. The system collects data in real time from various sensors installed on the production line, and a server analyzes the data to detect anomalies and optimize processes.

[0853] Hardware and software used

[0854] Hardware:

[0855] Sensors: Temperature sensors, pressure sensors, camera sensors, audio sensors, etc. These sensors are installed at various points along the production line to collect necessary data in real time.

[0856] Robot: A machine with autonomous capabilities installed on a production line that receives data from sensors and performs the required actions.

[0857] Terminal: A device such as a tablet or smartphone used by field workers. It receives notifications from the server and allows data to be checked and manually adjusted.

[0858] software:

[0859] Server (Central Processing Unit): Collects, pre-processes, analyzes, and processes feedback data.

[0860] OpenCV: A library for capturing and processing camera images, and is used to analyze image data.

[0861] TensorFlow: A library that performs data analysis using AI models, and is used for anomaly detection and quality assessment.

[0862] Requests: A data communication library used to send data from sensors and robots to a server.

[0863] Specific explanation for carrying out the invention

[0864] First, various sensors installed on the production line collect temperature, pressure, image, audio, and video data in real time. The collected data is immediately sent to a server, where it is preprocessed, for example, to remove noise from audio data and standardize the resolution of image data.

[0865] Once preprocessed, the data is analyzed using TensorFlow, which generates anomaly detection, quality assessment, and process optimization recommendations. The analysis results and optimization recommendations are sent to on-site workers' devices in real time. Workers can view the analysis results on their tablets or smartphones and take any necessary action.

[0866] Furthermore, robots installed on the production line collect data from various sensors in real time and send it to a server. The robots receive the analysis results from the server and can then optimize themselves autonomously. This system makes it possible to minimize production line downtime and maintain a high level of quality.

[0867] Specific examples (prompt sentence examples)

[0868] Below is an example of a prompt sentence to input to the generative AI model.

[0869] Create a program to analyze sensor data (images, temperature, pressure, etc.) collected by factory robots in real time, detect anomalies, and propose process optimization. Include a function for the robot to autonomously adjust itself based on the analysis results. Use OpenCV for image analysis, TensorFlow for AI analysis, and the Requests library for data communication.

[0870] According to the aspects of the present invention, real-time anomaly detection, quality evaluation, and process optimization become possible, thereby realizing improved efficiency and quality of the production line.

[0871] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0872] Step 1:

[0873] The sensors collect temperature, pressure, image, and audio data from the production line in real time. These data are collected in the form of sensors that each have. The input is the environmental condition of the production line, and the output is the raw data obtained from each sensor.

[0874] Step 2:

[0875] The server preprocesses the collected data, specifically removing noise from audio data, standardizing the resolution of image data, and standardizing other data formats. The input is raw data sent from the sensor, and the output is preprocessed data.

[0876] Step 3:

[0877] The server analyzes the preprocessed data. This analysis uses TensorFlow to perform anomaly detection, quality assessment, and process optimization recommendations. The analysis requires advanced computations and runs models to identify abnormal sounds and surface defects. The input is the preprocessed data, and the output is the analysis results and recommendations.

[0878] Step 4:

[0879] The server notifies the on-site worker of the analysis results and optimization proposals in real time. The worker's device is a tablet or smartphone, where they can check the analysis results and take any necessary measures. The input is the analysis results and proposals, and the output is the notification received by the on-site worker.

[0880] Step 5:

[0881] The user receives the notification and sends feedback from the device to the server. The feedback includes opinions about the analysis results and specific adjustments. The input is the user's feedback on the notification, and the output is the feedback data sent to the server.

[0882] Step 6:

[0883] The server improves the analysis algorithm based on the received feedback. This improves the accuracy of the analysis from the next time onwards, and further increases in production efficiency are expected. The input is feedback data from the user, and the output is a revised analysis algorithm.

[0884] Step 7:

[0885] The robot collects operational data in real time from various sensors installed on the production line and sends it to a server. The input is the operational data collected by the robot, and the output is the data sent to the server.

[0886] Step 8:

[0887] The server analyzes the operation data from the robot and feeds the results back to the robot. The robot receives the analysis results and performs optimization autonomously. The input is the robot's operation data and analysis results, and the output is the robot's optimized behavior.

[0888] This series of processing flows enables real-time monitoring and optimization of the entire production line.

[0889] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0890] This invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry, incorporating an emotion engine that recognizes the user's emotional state in real time and reflects it in optimization suggestions.

[0891] This system includes a server, various sensors, user devices, and an emotion engine. Sensors placed on the production line collect text data (numerical information such as temperature and pressure), audio data (machine noise, user voice), image data (photos of products, user facial expressions), and video data (video of the production line in operation) in real time. This data is sent to the server and managed centrally.

[0892] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, processes such as removing noise from audio data and standardizing the resolution of image data are performed. Next, the preprocessed data is analyzed using an AI model. This generates suggestions for anomaly detection, quality assessment, and process optimization.

[0893] Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotional state in real time. The emotion engine analyzes the user's emotions from voice and image data, detecting, for example, stress and fatigue. The analysis results are reflected in the server's analysis algorithm, resulting in the generation of more appropriate optimization suggestions.

[0894] The analysis results and optimization suggestions are sent to the device in real time. This allows the user (field worker) to check the analysis results and optimization suggestions using a tablet or smartphone and take any necessary action. For example, if the user is in a high-stress state, the notification may include a suggestion to take a break or redistribute tasks. User feedback is also analyzed through the emotion engine and used to improve the system as a whole.

[0895] The following is a specific example.

[0896] ---

[0897] Example 1: User stress detection

[0898] The server uses audio and image sensors to detect the user's stress level from their tone of voice and facial expressions. If the emotion engine analyzes the user's emotional state as "high stress," it sends the result to the server. Based on this information, the server notifies the device with suggestions for reducing the user's workload (e.g., taking a break or switching to a lighter task). This maximizes the user's performance while maintaining their physical and mental health.

[0899] ---

[0900] Example 2: Emotional analysis feedback

[0901] If a user feels dissatisfied or stressed during product inspection, the emotion engine detects that emotional state. The server incorporates this information into the analysis results and receives it as feedback. For example, if the system detects that the user is "tired and prone to making judgment errors," it will adjust the sensitivity of the analysis algorithm next time and strengthen settings to minimize human error.

[0902] ---

[0903] Example 3: Dynamic adjustment of operating environment

[0904] The server analyzes data from the entire production line and the user's emotional data to suggest dynamic adjustments to the operating environment. For example, by appropriately changing the temperature settings for the entire line or adjusting the operating speed of specific machines, overall efficiency can be improved. When making these suggestions based on the user's emotional data, the emotion engine also takes into account whether the working environment is comfortable.

[0905] ---

[0906] This system will enable manufacturing production lines to operate with even greater efficiency and quality. By combining an emotion engine with AI analysis, appropriate suggestions that take into account the user's condition will be made in real time, enabling more effective optimization of the production process.

[0907] The processing flow will be explained below.

[0908] Step 1:

[0909] The server collects text data, audio data, image data, and video data in real time from various sensors, including text data from temperature and pressure sensors, audio data from microphones, and image and video data from cameras.

[0910] Step 2:

[0911] The server preprocesses the collected data, removing noise from the audio data, standardizing the resolution of the image data, and aligning the timestamps of all data. During this process, the data is temporarily stored in a database.

[0912] Step 3:

[0913] The server then inputs the preprocessed data into an AI model for analysis. During the analysis stage, machine learning algorithms are used to generate anomaly detection, quality assessment, and process optimization recommendations. For example, an anomaly detection model can be used to detect abnormal sounds from audio data, and an image analysis algorithm can be used to automatically detect defects on the surface of a product.

[0914] Step 4:

[0915] The server notifies the user's device of the analysis results and optimization proposals in real time, allowing the device to allow on-site workers to immediately check the analysis results and optimization proposals.

[0916] Step 5:

[0917] The device equipped with the emotion engine analyzes the user's voice tone and facial expression data in real time to evaluate the user's emotional state (e.g., stress, fatigue). Specifically, it detects changes in tone and speed of voice from the voice data, and senses changes in facial expression from the image data.

[0918] Step 6:

[0919] The server receives the emotional state data sent from the emotion engine. Based on this information, the server adjusts the analysis algorithm and generates optimization suggestions according to the user's emotional state. For example, if the user is in a high-stress state, the server generates suggestions for taking a break or reducing workload.

[0920] Step 7:

[0921] The server then sends the newly generated optimization proposals to the user's device, where the user can view these proposals in real time. The user can then use the device to input feedback, such as accepting the proposals or requesting modifications.

[0922] Step 8:

[0923] The server receives feedback data from users and continuously improves the analysis algorithm. The feedback, including emotional data, is analyzed to improve the accuracy of the next data analysis.

[0924] ---

[0925] This system monitors and optimizes manufacturing production lines in real time, continuously improving production efficiency and quality. The introduction of an emotion engine enables flexible responses according to the user's condition, achieving both an improved work environment and increased production efficiency.

[0926] Example 2

[0927] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0928] Improving production efficiency and quality are important challenges in modern manufacturing. However, conventional systems focused on collecting and analyzing data from sensors, and did not fully consider the impact of workers' emotional states. As a result, poor performance and mistakes caused by worker stress and fatigue had a negative impact on production efficiency and quality. This made it difficult to manage worker health and optimize production processes.

[0929] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0930] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality assessment, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for analyzing the emotional state of the workers from their voice data and image data, means for generating optimization proposals based on their emotional states, and means for improving the analysis algorithm based on feedback from the workers. This improves production efficiency and quality and enables optimal proposals that take into account the health status of the workers.

[0931] A "sensor" is a device that collects different types of data in real time, such as temperature, pressure, sound, images, and video.

[0932] "Data preprocessing" refers to the process of performing preprocessing such as noise removal and resolution standardization on collected data, and then temporarily storing it.

[0933] A "database" is a storage device for temporarily storing preprocessed data.

[0934] An "AI model" is an artificial intelligence algorithm that performs analysis for anomaly detection, quality evaluation, and process optimization.

[0935] "Analysis results" are the results of anomaly detection and quality assessment generated by the AI ​​model using preprocessed data.

[0936] An "optimization proposal" is a proposal for improving the production process that is generated based on the analysis results.

[0937] "Field workers" are workers who work on the manufacturing floor and act based on data collected from sensors.

[0938] "Feedback" refers to opinions and reactions to the system analysis and suggestions provided by field workers.

[0939] "Emotional state" refers to the emotional state of a worker, such as stress or fatigue, analyzed from the worker's voice data and image data.

[0940] The "emotion engine" is a system that analyzes the emotional state of workers from collected voice and image data.

[0941] "Notification" is a communication method by which the server conveys analysis results and optimization suggestions to field workers in real time.

[0942] An "analysis algorithm" is a calculation procedure for detecting anomalies and evaluating quality based on collected data.

[0943] This invention is an AI-driven robot support system for improving production efficiency and quality in the manufacturing industry, incorporating an emotion engine that recognizes the user's emotional state in real time and reflects it in optimization proposals. This system is composed of a server, various sensors, user terminals, and the emotion engine.

[0944] The server is the central part that collects data in real time from sensors placed on the production line and performs preprocessing and analysis. The main hardware used is a high-performance server computer, and the software includes Python, TensorFlow, PyTorch, OpenCV, Librosa, etc.

[0945] Data collection

[0946] The server uses various sensors to collect temperature, pressure, audio, image, and video data. For example, the temperature sensor measures the temperature data of the line in real time, the microphone collects user voices and machine sounds, and the image sensor takes photos of the product and records video of the production line operation.

[0947] Data Preprocessing

[0948] The collected data is sent to a server where preprocessing such as noise removal and resolution unification is performed. For example, the Librosa library is used to remove environmental noise from audio data, and the resolution of image data is unified using OpenCV. This organizes the data into a form optimal for analysis.

[0949] Data analysis

[0950] The preprocessed data is then analyzed by the server using AI models. Anomaly detection models are run using TensorFlow and PyTorch. Quality assessments and recommendations for optimizing the production process are also generated at this stage. If an anomaly is detected, its details are recorded in a log.

[0951] Emotion analysis

[0952] Furthermore, the system incorporates an emotion engine that recognizes the user's emotional state in real time. The emotion engine uses natural language processing models (e.g., BERT) and facial recognition models (e.g., DeepFace) to analyze the user's emotions from voice and image data, thereby determining whether the user is in a stressful state.

[0953] Proposal generation and notification

[0954] Optimization suggestions are generated based on the analysis results and the user's emotional state. The server notifies the user of these suggestions in real time. For example, if a user is in a high-stress state, the device may be notified of suggestions such as taking a break or redistributing tasks. The device application uses React Native or Android Studio, allowing users to view these notifications on their tablets or smartphones.

[0955] Feedback collection and system improvement

[0956] Users can provide feedback on the suggestions provided to them through their devices. The server collects this feedback and uses it to improve the analysis algorithm. This allows the system to learn from the feedback and improve the accuracy of future suggestions.

[0957] Example: Detecting user stress

[0958] The server uses audio and image sensors to detect the user's stress level from their tone of voice and facial expressions. If the emotion engine analyzes the user's emotional state as "high stress," it sends the result to the server. Based on this information, the server notifies the device with suggestions for reducing the user's workload (e.g., taking a break or switching to a lighter task). This maximizes the user's performance while maintaining their physical and mental health.

[0959] Example prompt sentence:

[0960] "Analyze the user's stress level based on their voice and image data, and make optimization suggestions based on the results."

[0961] This system will enable manufacturing production lines to operate with even greater efficiency and quality. By combining an emotion engine with AI analysis, appropriate suggestions that take into account the user's condition will be made in real time, enabling more effective optimization of the production process.

[0962] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0963] Step 1: Data collection

[0964] Inputs: temperature, pressure, audio, image, and video data.

[0965] Processing: The server collects data in real time from various sensors placed on the production line.

[0966] Output: Collected multimodal data.

[0967] How it works: The server collects real-time temperature data every 10 seconds from the temperature sensor, and collects user voice and machine sounds from the microphone. At the same time, the image sensor captures the product's appearance, and the video sensor records the operation of the production line.

[0968] Step 2: Data Preprocessing

[0969] Input: Collected multimodal data.

[0970] Processing: The server preprocesses the collected data, removing noise and unifying the resolution for temporary storage.

[0971] Output: Preprocessed data.

[0972] Specific operation: Noise reduction is performed on the audio data using the Librosa library, and the resolution of the image data is unified using OpenCV. The preprocessed data is temporarily stored in a database.

[0973] Step 3: Data analysis

[0974] Input: Preprocessed data.

[0975] Processing: The server uses the preprocessed data to perform analysis using an AI model.

[0976] Output: Analysis results (anomaly detection, quality assessment, process optimization suggestions).

[0977] What it does: Runs anomaly detection models using TensorFlow or PyTorch to generate quality assessments and process optimization suggestions. If an anomaly is detected, details of the anomaly are logged.

[0978] Step 4: Sentiment Analysis

[0979] Input: Preprocessed audio and image data.

[0980] Processing: The server analyzes the user's emotional state using an emotion engine.

[0981] Output: Emotion analysis results (e.g. high stress, fatigue).

[0982] Specific operation: Emotions are analyzed from voice data using BERT, and facial expressions are analyzed from image data using DeepFace. The analysis results of the emotional state are stored on the server.

[0983] Step 5: Proposal Generation

[0984] Input: Data analysis results and sentiment analysis results.

[0985] Processing: The server generates optimization suggestions based on the data analysis results and sentiment analysis results.

[0986] Output: Optimization suggestions (e.g. break suggestions, task redistribution).

[0987] Specific operation: Based on the analysis results, suggestions to reduce the user's workload are automatically generated and prepared for notification in the next step.

[0988] Step 6: Notification of results

[0989] Input: Optimization proposal.

[0990] Processing: The server notifies the user terminal of the generated optimization proposal.

[0991] Output: A notification message on the terminal.

[0992] Specific operation: The server pushes JSON format data to the device, and the application on the device interprets the data and displays it to the user.

[0993] Step 7: Gather feedback and improve the system

[0994] Input: User feedback.

[0995] Processing: The user enters feedback on the suggestions provided, and the server collects the feedback and uses it to improve the system.

[0996] Output: Feedback data, improved AI model.

[0997] How it works: Users input feedback from their devices, and the server collects that information and reflects it in the next analysis, thereby improving the accuracy of the analysis algorithm.

[0998] These steps not only improve the efficiency and quality of the production process, but also optimize the working environment, maintaining worker health and performance through suggestions that take the user's emotional state into account.

[0999] (Application example 2)

[1000] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1001] Conventional production systems focus on detecting machine anomalies and evaluating quality, but rarely consider the emotional state of human workers. This increases the risk of workers feeling stressed or fatigued, leading to reduced production efficiency. Furthermore, the lack of work instructions or improvement suggestions based on the emotional state of workers poses a challenge, making it difficult to fully optimize the entire production process.

[1002] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1003] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality assessment, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for improving the analysis algorithm based on feedback from the workers, and means for analyzing the emotions of the workers in real time and providing work instructions based on their emotional states. This makes it possible to provide optimal work instructions and propose process improvements that take the emotional states of the workers into consideration, thereby achieving both improved production efficiency and maintaining the health of the workers.

[1004] "Multiple sensors" are various sensors used to collect text data, audio data, image data, and video data in real time.

[1005] "Preprocessing" refers to the process of converting collected data into a form suitable for analysis, and specifically includes noise removal and resolution unification.

[1006] A "database" is a system for temporarily storing preprocessed data.

[1007] "Analysis algorithms" are techniques for analyzing pre-processed data and generating suggestions for anomaly detection, quality assessment, and process optimization.

[1008] "Emotional state" refers to the emotional state of the worker analyzed from voice and image data, and specifically refers to stress and fatigue.

[1009] "Work instructions" refers to specific work tasks and break suggestions provided to workers based on their emotional state.

[1010] "Feedback" refers to responses and reactions from workers, and is information collected to help improve the system.

[1011] An "emotion analysis device" is a piece of equipment used to analyze the emotional state of a worker in real time.

[1012] "Production process improvement proposals" refer to specific proposals for improving production efficiency based on the analysis results.

[1013] "Stress state" refers to the degree of tension and strain felt by the worker, and is analyzed by the emotion engine.

[1014] To realize the present invention, the following detailed description of the system and its components is required. The system of the present invention collects and analyzes data in real time and provides appropriate feedback and work instructions. Specific components include a server, various sensors, a user's device, and an emotion analysis engine.

[1015] server

[1016] The server is the core processing unit of the system and performs the following functions:

[1017] Data collection: Collect text, audio, image, and video data in real time from a variety of sensors.

[1018] Data preprocessing: The collected data is preprocessed by noise removal and resolution standardization, and then temporarily stored in a database.

[1019] Data Analysis: The pre-processed data is analyzed using generative AI models to generate recommendations for anomaly detection, quality assessment, and process optimization.

[1020] Emotion analysis: An emotion analysis engine is used to analyze the user's emotional state from voice and image data.

[1021] Various sensors

[1022] The system uses the following sensors:

[1023] Audio sensor: Collects the user's voice tone and mechanical sounds to detect stress levels and abnormal sounds.

[1024] Image sensor: Collects footage of the user's facial expressions and production line operations to inspect the emotional state and appearance of the product.

[1025] User's device

[1026] The user's device (smartphone, tablet, etc.) performs the following functions:

[1027] Notifications: Analysis results and optimization suggestions from the server are displayed in real time.

[1028] Feedback: Collects feedback from users and sends it to the server.

[1029] Sentiment Analysis Engine

[1030] The emotion analysis engine analyzes the user's emotional state from voice data and image data and provides the data to the server.

[1031] This is implemented using the following software and libraries:

[1032] OpenCV: Face Recognition and Landmark Detection

[1033] dlib: Facial landmark analysis

[1034] Keras: Sentiment Analysis Model

[1035] requests: Send notification

[1036] Specific examples

[1037] For example, when developing an application to detect the stress level of workers at a logistics center, the smartphone's camera and microphone can be used to analyze the worker's facial expressions and tone of voice. If the worker's emotional state is determined to be "high stress," the server will send a notification to take a break. In this way, optimal work instructions that take the worker's emotions into consideration can be provided, thereby improving production efficiency and maintaining the worker's health at the same time.

[1038] Prompt Sentence Examples

[1039] Develop an application that evaluates workers' emotions through real-time analysis of camera footage and suggests breaks when they are stressed. The technologies used are OpenCV, dlib, and Keras. The goal is to analyze emotions from workers' facial expressions and suggest breaks when stress or fatigue is detected.

[1040] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1041] Step 1:

[1042] The server collects text data, voice data, image data, and video data in real time from various sensors (audio sensors, image sensors). The input is data from the sensors, and the output is raw data for preprocessing. Specifically, the audio sensor collects the tone of the user's voice, and the image sensor captures the user's facial expressions and footage of the production line operation.

[1043] Step 2:

[1044] The server preprocesses the collected data. Specifically, it removes noise from the audio data and standardizes the resolution of the image data. At this point, the input is raw data and the output is preprocessed data. The preprocessed data is temporarily stored in a database.

[1045] Step 3:

[1046] The server inputs the preprocessed data into the generative AI model for analysis. The input is the preprocessed data, and the output of the analysis is anomaly detection, quality assessment, and process optimization recommendations. Specifically, the AI ​​model uses image data to perform visual inspections of products and analyzes audio data to detect abnormal machine sounds.

[1047] Step 4:

[1048] The server analyzes the user's emotional state using an emotion analysis engine. The input is audio and image data, and the output is the user's emotional state (stress or fatigue). Specifically, it uses OpenCV and dlib to detect facial landmarks, and classifies the emotional state using an emotion analysis model using Keras.

[1049] Step 5:

[1050] The server notifies the user device of the analysis results and optimization suggestions in real time. The input is the analysis results and sentiment analysis results, and the output is a notification message. Specifically, it uses the requests library to send a message to the notification API, which is then displayed on the user device. For example, if a high stress state is detected, a notification such as "Take a break" is sent.

[1051] Step 6:

[1052] The user checks the suggestions through the terminal and takes the necessary action. The input is a notification message and the output is feedback information. The suggestions are displayed on the terminal, and the user can check them and take action.

[1053] Step 7:

[1054] The terminal collects feedback from the user and sends it to the server. The input is the feedback information from the user, and the output is the feedback data sent to the server. Specifically, an interface is provided for the user to report task completion notifications and new problems.

[1055] Step 8:

[1056] The server improves the analysis algorithm based on the collected feedback. The input is the feedback data, and the output is an improved analysis algorithm. Specifically, it reevaluates the analysis results and adjusts the model parameters.

[1057] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1058] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1059] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1060] [Fourth embodiment]

[1061] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1062] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1063] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1064] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1065] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1066] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1067] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1068] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1069] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1070] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1071] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1072] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1073] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1074] The present invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry, and is specifically implemented as follows.

[1075] The central roles of the system are played by a server, various sensors, and terminals (used by users). Various sensors placed on the production line collect text data (numerical information such as temperature and pressure), audio data (machine sounds), image data (photos of products), and video data (video of the production line in operation) in real time. This data is sent to the server and managed centrally.

[1076] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, processes such as removing noise from audio data and standardizing the resolution of image data are performed. Next, the preprocessed data is analyzed using an AI model. This generates suggestions for anomaly detection, quality assessment, and process optimization.

[1077] Analysis results and optimization suggestions are sent to the terminal in real time, allowing the user (field worker) to check the analysis results and take necessary measures using a tablet or smartphone. Users can also enter feedback, which allows the server to continuously improve the analysis algorithm.

[1078] A specific example is shown below.

[1079] ---

[1080] Example 1: Detecting abnormal sounds from machinery

[1081] The server analyzes machine sound data collected from the audio sensor in real time. If an abnormal sound that differs from normal operating sounds is detected, the server notifies the on-site worker's device of the occurrence of the abnormal sound. The user receives this notification, checks the machine's status, and performs maintenance work if necessary. This makes it possible to minimize production line downtime.

[1082] ---

[1083] Example 2: Automatic detection of surface defects on products

[1084] The server analyzes product images collected from the camera sensor and automatically detects small scratches or stains on the surface. If an abnormality is detected, the image data and the location of the abnormality are notified to the user's device. The user can check the notification and take action to re-inspect or correct the product. This makes it possible to maintain a high level of product quality.

[1085] ---

[1086] Example 3: Process optimization proposal

[1087] The server collects and analyzes numerical data related to the production process, such as temperature, pressure, and speed. Based on the analysis results, it suggests fine-tuning the temperature or pressure, for example. These suggestions are sent to the user's device, and if the user accepts the suggestions, the settings are automatically changed. This improves the efficiency of the entire production process and ensures optimal resource utilization.

[1088] ---

[1089] This system enables real-time monitoring and optimization of manufacturing production lines, dramatically improving production efficiency and quality. By effectively linking data collection, analysis, and feedback, on-site workers and robots can work together to achieve more efficient, higher-quality production.

[1090] The processing flow will be explained below.

[1091] Step 1:

[1092] The server collects text data, audio data, image data, and video data in real time from various sensors, including text data from temperature and pressure sensors, audio data from microphones, and image and video data from cameras.

[1093] Step 2:

[1094] The server preprocesses the collected data. Preprocessing includes data cleansing (such as noise removal), format conversion, and timestamp synchronization. For example, it removes background noise from audio data, standardizes the resolution of image data, and aligns the timestamps of all data.

[1095] Step 3:

[1096] The server then inputs the preprocessed data into the AI ​​model for analysis. During the analysis phase, machine learning algorithms are used to generate anomaly detection, quality assessment, and process optimization recommendations. For example, an anomaly detection model can be used to detect abnormal sounds from audio data, and an image analysis algorithm can be used to automatically detect defects on the surface of a product.

[1097] Step 4:

[1098] The server sends analysis results and optimization proposals in real time to the devices used by field workers, such as tablets or smartphones, which the workers use to check the analysis results and optimization proposals.

[1099] Step 5:

[1100] Users can use their devices to check the analysis results and optimization suggestions, and provide feedback as needed. For example, they can check abnormality detection notifications, inspect the machine's status, restart it, or perform maintenance.

[1101] Step 6:

[1102] The server receives feedback data from users and improves the analysis algorithm, which will enable more accurate data analysis the next time. Specifically, the feedback data is used to adjust the parameters of the machine learning model, improving the accuracy of anomaly detection and quality assessment.

[1103] Example 1

[1104] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1105] Improving production efficiency and quality is an important issue in the manufacturing industry. In particular, there is a demand for the detection of abnormal sounds, early detection of surface defects on products, and real-time process optimization. However, these issues have not been adequately resolved with conventional methods. A system that reduces the burden on on-site workers and provides highly accurate data analysis and notifications in real time is needed.

[1106] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1107] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for inputting the preprocessed data into a generative AI model for analysis, means for notifying the mobile devices of on-site workers of analysis results and improvement suggestions in real time, and means for receiving feedback from workers and improving the analysis algorithm. This enables highly accurate anomaly detection, quality evaluation, and process optimization in real time, dramatically improving production efficiency and quality while reducing the burden on on-site workers.

[1108] A "sensor" is a device that detects physical or environmental data in real time and outputs that information as a signal.

[1109] "Text data" is data that expresses numerical information such as temperature, pressure, and speed as a string of characters.

[1110] "Audio data" refers to sound wave information recorded in digital format. Specifically, it refers to audio information such as machine operating sounds and abnormal sounds.

[1111] "Image data" is data that digitally represents still images captured by a camera or sensor.

[1112] "Video data" is a series of images recorded along a time axis, allowing for visual capture of movements and changes in processes.

[1113] "Preprocessing" refers to processes such as filtering, noise removal, and format standardization that are performed to improve the quality of collected data.

[1114] A "database" is a storage device or system for temporarily storing preprocessed data.

[1115] A "generative AI model" is an algorithm or software that uses artificial intelligence techniques to perform data analysis.

[1116] "Analysis" is a series of procedures for anomaly detection, quality assessment, and process optimization based on collected and pre-processed data.

[1117] "Notification" refers to the act of communicating analysis results and improvement suggestions to the mobile devices of field workers in real time.

[1118] "Feedback" is input from field workers that is used to continually improve the analysis algorithms.

[1119] An "algorithm" is a collection of computational procedures or mathematical formulas for the purpose of data analysis or process optimization.

[1120] A "field worker" is a user who is in charge of work on the production line and receives notifications and suggestions from the system.

[1121] The present invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry. Specifically, it is configured as follows.

[1122] The central roles of the system are played by the server, various sensors, and terminals (used by users). Various sensors (temperature sensors, audio sensors, camera sensors, etc.) placed on the production line collect numerical information (text data) such as temperature and pressure, machine sounds (audio data), product photos (image data), and video of the production line in operation (video data) in real time. This data is sent to the server and managed centrally.

[1123] Data Preprocessing

[1124] The server preprocesses the collected data and temporarily stores it. This preprocessing involves, for example, removing noise from audio data and standardizing the resolution of image data. The technologies used here include audio filtering and image processing algorithms.

[1125] Data analysis

[1126] The preprocessed data is then fed into a generative AI model for analysis. The AI ​​model uses TensorFlow or PyTorch, among others. The analysis generates anomaly detection, quality assessment, and recommendations for process optimization. For example, it can detect abnormal sounds from audio data or identify surface defects on products from image data.

[1127] Notification and feedback of analysis results

[1128] Analysis results and optimization suggestions are sent to the terminal in real time, allowing users (field workers) to check the analysis results using a tablet or smartphone and take any necessary measures. Users can also enter feedback, which allows the server to continuously improve the analysis algorithm.

[1129] Example 1: Detecting abnormal sounds from machinery

[1130] The server analyzes machine sound data collected from the audio sensor in real time. If an abnormal sound that differs from normal operating sounds is detected, the server notifies the on-site worker's device of the occurrence of the abnormal sound. The user receives this notification, checks the machine's status, and performs maintenance work if necessary. This makes it possible to minimize production line downtime.

[1131] Example 2: Automatic detection of surface defects on products

[1132] The server analyzes product images collected from the camera sensor and automatically detects small scratches or stains on the surface. If an abnormality is detected, the image data and the location of the abnormality are notified to the user's device. The user can check the notification and take action to re-inspect or correct the product. This makes it possible to maintain a high level of product quality.

[1133] Example 3: Process optimization proposal

[1134] The server collects and analyzes numerical data related to the production process, such as temperature, pressure, and speed. Based on the analysis results, it suggests fine-tuning the temperature or pressure, for example. These suggestions are sent to the user's device, and if the user accepts the suggestions, the settings are automatically changed. This improves the efficiency of the entire production process and ensures optimal resource utilization.

[1135] Prompt Sentence Examples

[1136] The server should analyze the audio data in real time and detect any abnormal sounds.

[1137] "The server should analyze the product image data and automatically detect surface defects."

[1138] "The server should analyze production process data such as temperature and pressure and suggest optimal settings."

[1139] This system enables real-time monitoring and optimization of manufacturing production lines, dramatically improving production efficiency and quality. A series of processes—data collection, preprocessing, analysis, notification, and feedback—work together to enable on-site workers and robots to work together and achieve highly efficient, high-quality production.

[1140] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1141] Step 1: Data collection

[1142] The server collects data in real time from various sensors installed on the production line, such as temperature sensors, audio sensors, and camera sensors. This data includes text data on temperature and pressure, audio data on machine operating sounds and abnormal sounds, and product images and video data.

[1143] Input: Real-time data from various sensors

[1144] Data processing: receiving and storing sensor data

[1145] Output: Collected raw sensor data

[1146] Specific behavior:

[1147] The server obtains temperature data from the temperature sensor every 30 seconds. For example, it collects data such as "Temperature: 27.5°C, 28.0°C, 27.8°C."

[1148] Audio data is streamed from the audio sensor every second to capture the sound of the machine operating.

[1149] The camera sensor captures images of the product every minute and collects video data every five minutes.

[1150] Step 2: Data Preprocessing

[1151] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, noise is removed from the audio data and the resolution of the image data is standardized.

[1152] Input: Collected raw sensor data

[1153] Data processing: noise filtering, resolution unification, format conversion, etc.

[1154] Output: Pre-processed, high-quality data

[1155] Specific behavior:

[1156] The server applies a noise filter to the audio data to generate clear audio data.

[1157] The server standardizes the resolution of image data to a uniform 1024x768 pixels and converts it from JPEG format to PNG format.

[1158] Step 3: Data analysis

[1159] The server inputs the preprocessed data into a generative AI model for analysis, which generates recommendations for anomaly detection, quality assessment, and process optimization. The AI ​​model uses TensorFlow and PyTorch.

[1160] Input: Preprocessed data

[1161] Data processing: Data analysis using AI models

[1162] Output: Analysis results and optimization suggestions

[1163] Specific behavior:

[1164] The server uses TensorFlow to detect abnormal sounds from the preprocessed audio data, and obtains a result such as "abnormal sound detected at timestamp '00:01:15'".

[1165] The server uses PyTorch to analyze the image data and identify defects on the product surface. For example, it obtains a result such as "Small scratches on the product surface detected at region (x: 200, y: 350)."

[1166] Step 4: Notification of analysis results and feedback

[1167] The server notifies the device in real time of analysis results and optimization suggestions, allowing the user to respond, and also receives user feedback to continuously improve the analysis algorithm.

[1168] Input: Analysis results and optimization proposals

[1169] Data processing: notification generation, feedback processing

[1170] Output: Notifying field workers, collecting feedback data

[1171] Specific behavior:

[1172] The server then sends a notification to the on-site worker's tablet regarding any abnormal sounds detected as a result of the analysis, such as "Abnormal sound detected at timestamp 00:01:15."

[1173] The user checks the notification displayed on the tablet and performs the necessary maintenance work.

[1174] Users input the results and feedback after maintenance work into a tablet and send it to the server, which uses this feedback to improve its analysis algorithm.

[1175] Through these steps, the system can improve the efficiency and quality of the production line in real time.

[1176] (Application example 1)

[1177] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1178] Improving production efficiency and quality is a key challenge in modern manufacturing. However, conventional systems have difficulty analyzing data in real time and responding immediately, making early detection of abnormalities and process optimization insufficient. This makes it difficult to minimize production line downtime and maintain a high level of quality.

[1179] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1180] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality evaluation, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for improving the analysis algorithm based on feedback from the workers, means for robots installed on the production line to collect operation data in real time and send it to the server, means for the robots to autonomously optimize themselves based on the analysis results from the server, and means for notifying on-site workers of the analysis results so that they can make manual adjustments. This enables real-time anomaly detection, quality evaluation, and process optimization, thereby improving the efficiency and quality of the production line.

[1181] The "various sensors" are various types of sensors used in manufacturing sites, and are devices that have the function of collecting text data, audio data, image data, and video data.

[1182] "Text data" is data that expresses the state of a machine or environment as numbers or strings of characters, and is primarily information indicating temperature, pressure, speed, etc.

[1183] "Audio data" refers to data collected as acoustic signals from sounds generated from production lines and machines.

[1184] "Image data" is digital data of still images captured by an imaging device such as a camera, and represents the state of a product or production line.

[1185] "Video data" refers to digital data of video captured by a camera or video device, and is data that includes a series of images.

[1186] "Preprocessing" refers to the process of converting collected data into a format suitable for analysis, and includes noise removal and resolution unification.

[1187] A "database" is a computer system for temporarily storing preprocessed data.

[1188] "Analysis" is the process of applying calculations and models to pre-processed data to detect anomalies, assess quality, and suggest process optimization.

[1189] "Abnormality detection" refers to detecting abnormal conditions that differ from normal operating conditions, such as detecting abnormal sounds or vibrations in machinery.

[1190] "Quality evaluation" is the process of evaluating the quality of a product or process to determine whether it meets specified standards.

[1191] "Process optimization" is the process of finding the optimal parameters and procedures in a production process to improve overall efficiency.

[1192] "Real-time" means that the system collects data and analyzes and notifies immediately, with little time delay.

[1193] "Feedback" refers to information that collects opinions and evaluations from workers and system users and is used to improve the system.

[1194] A "server" is a central processing unit that collects, stores, analyzes, and processes feedback data on a network.

[1195] A "robot" is a mechanical device that is installed in a production line and can operate autonomously.

[1196] A "terminal" is a device such as a tablet or smartphone used by a field worker to receive notifications from the server.

[1197] This invention is an AI-driven robotics assistance system for improving production efficiency and quality in the manufacturing industry. The system collects data in real time from various sensors installed on the production line, and a server analyzes the data to detect anomalies and optimize processes.

[1198] Hardware and software used

[1199] Hardware:

[1200] Sensors: Temperature sensors, pressure sensors, camera sensors, audio sensors, etc. These sensors are installed at various points along the production line to collect necessary data in real time.

[1201] Robot: A machine with autonomous capabilities installed on a production line that receives data from sensors and performs the required actions.

[1202] Terminal: A device such as a tablet or smartphone used by field workers. It receives notifications from the server and allows data to be checked and manually adjusted.

[1203] software:

[1204] Server (Central Processing Unit): Collects, pre-processes, analyzes, and processes feedback data.

[1205] OpenCV: A library for capturing and processing camera images, and is used to analyze image data.

[1206] TensorFlow: A library that performs data analysis using AI models, and is used for anomaly detection and quality assessment.

[1207] Requests: A data communication library used to send data from sensors and robots to a server.

[1208] Specific explanation for carrying out the invention

[1209] First, various sensors installed on the production line collect temperature, pressure, image, audio, and video data in real time. The collected data is immediately sent to a server, where it is preprocessed, for example, to remove noise from audio data and standardize the resolution of image data.

[1210] Once preprocessed, the data is analyzed using TensorFlow, which generates anomaly detection, quality assessment, and process optimization recommendations. The analysis results and optimization recommendations are sent to on-site workers' devices in real time. Workers can view the analysis results on their tablets or smartphones and take any necessary action.

[1211] Furthermore, robots installed on the production line collect data from various sensors in real time and send it to a server. The robots receive the analysis results from the server and can then optimize themselves autonomously. This system makes it possible to minimize production line downtime and maintain a high level of quality.

[1212] Specific examples (prompt sentence examples)

[1213] Below is an example of a prompt sentence to input to the generative AI model.

[1214] Create a program to analyze sensor data (images, temperature, pressure, etc.) collected by factory robots in real time, detect anomalies, and propose process optimization. Include a function for the robot to autonomously adjust itself based on the analysis results. Use OpenCV for image analysis, TensorFlow for AI analysis, and the Requests library for data communication.

[1215] According to the aspects of the present invention, real-time anomaly detection, quality evaluation, and process optimization become possible, thereby realizing improved efficiency and quality of the production line.

[1216] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1217] Step 1:

[1218] The sensors collect temperature, pressure, image, and audio data from the production line in real time. These data are collected in the form of sensors that each have. The input is the environmental condition of the production line, and the output is the raw data obtained from each sensor.

[1219] Step 2:

[1220] The server preprocesses the collected data, specifically removing noise from audio data, standardizing the resolution of image data, and standardizing other data formats. The input is raw data sent from the sensor, and the output is preprocessed data.

[1221] Step 3:

[1222] The server analyzes the preprocessed data. This analysis uses TensorFlow to perform anomaly detection, quality assessment, and process optimization recommendations. The analysis requires advanced computations and runs models to identify abnormal sounds and surface defects. The input is the preprocessed data, and the output is the analysis results and recommendations.

[1223] Step 4:

[1224] The server notifies the on-site worker of the analysis results and optimization proposals in real time. The worker's device is a tablet or smartphone, where they can check the analysis results and take any necessary measures. The input is the analysis results and proposals, and the output is the notification received by the on-site worker.

[1225] Step 5:

[1226] The user receives the notification and sends feedback from the device to the server. The feedback includes opinions about the analysis results and specific adjustments. The input is the user's feedback on the notification, and the output is the feedback data sent to the server.

[1227] Step 6:

[1228] The server improves the analysis algorithm based on the received feedback. This improves the accuracy of the analysis from the next time onwards, and further increases in production efficiency are expected. The input is feedback data from the user, and the output is a revised analysis algorithm.

[1229] Step 7:

[1230] The robot collects operational data in real time from various sensors installed on the production line and sends it to a server. The input is the operational data collected by the robot, and the output is the data sent to the server.

[1231] Step 8:

[1232] The server analyzes the operation data from the robot and feeds the results back to the robot. The robot receives the analysis results and performs optimization autonomously. The input is the robot's operation data and analysis results, and the output is the robot's optimized behavior.

[1233] This series of processing flows enables real-time monitoring and optimization of the entire production line.

[1234] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1235] This invention is an AI-driven robotic assistance system for improving production efficiency and quality in the manufacturing industry, incorporating an emotion engine that recognizes the user's emotional state in real time and reflects it in optimization suggestions.

[1236] This system includes a server, various sensors, user devices, and an emotion engine. Sensors placed on the production line collect text data (numerical information such as temperature and pressure), audio data (machine noise, user voice), image data (photos of products, user facial expressions), and video data (video of the production line in operation) in real time. This data is sent to the server and managed centrally.

[1237] The server preprocesses the collected data and temporarily stores it. During the preprocessing stage, processes such as removing noise from audio data and standardizing the resolution of image data are performed. Next, the preprocessed data is analyzed using an AI model. This generates suggestions for anomaly detection, quality assessment, and process optimization.

[1238] Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotional state in real time. The emotion engine analyzes the user's emotions from voice and image data, detecting, for example, stress and fatigue. The analysis results are reflected in the server's analysis algorithm, resulting in the generation of more appropriate optimization suggestions.

[1239] The analysis results and optimization suggestions are sent to the device in real time. This allows the user (field worker) to check the analysis results and optimization suggestions using a tablet or smartphone and take any necessary action. For example, if the user is in a high-stress state, the notification may include a suggestion to take a break or redistribute tasks. User feedback is also analyzed through the emotion engine and used to improve the system as a whole.

[1240] The following is a specific example.

[1241] ---

[1242] Example 1: User stress detection

[1243] The server uses audio and image sensors to detect the user's stress level from their tone of voice and facial expressions. If the emotion engine analyzes the user's emotional state as "high stress," it sends the result to the server. Based on this information, the server notifies the device with suggestions for reducing the user's workload (e.g., taking a break or switching to a lighter task). This maximizes the user's performance while maintaining their physical and mental health.

[1244] ---

[1245] Example 2: Emotional analysis feedback

[1246] If a user feels dissatisfied or stressed during product inspection, the emotion engine detects that emotional state. The server incorporates this information into the analysis results and receives it as feedback. For example, if the system detects that the user is "tired and prone to making judgment errors," it will adjust the sensitivity of the analysis algorithm next time and strengthen settings to minimize human error.

[1247] ---

[1248] Example 3: Dynamic adjustment of operating environment

[1249] The server analyzes data from the entire production line and the user's emotional data to suggest dynamic adjustments to the operating environment. For example, by appropriately changing the temperature settings for the entire line or adjusting the operating speed of specific machines, overall efficiency can be improved. When making these suggestions based on the user's emotional data, the emotion engine also takes into account whether the working environment is comfortable.

[1250] ---

[1251] This system will enable manufacturing production lines to operate with even greater efficiency and quality. By combining an emotion engine with AI analysis, appropriate suggestions that take into account the user's condition will be made in real time, enabling more effective optimization of the production process.

[1252] The processing flow will be explained below.

[1253] Step 1:

[1254] The server collects text data, audio data, image data, and video data in real time from various sensors, including text data from temperature and pressure sensors, audio data from microphones, and image and video data from cameras.

[1255] Step 2:

[1256] The server preprocesses the collected data, removing noise from the audio data, standardizing the resolution of the image data, and aligning the timestamps of all data. During this process, the data is temporarily stored in a database.

[1257] Step 3:

[1258] The server then inputs the preprocessed data into an AI model for analysis. During the analysis stage, machine learning algorithms are used to generate anomaly detection, quality assessment, and process optimization recommendations. For example, an anomaly detection model can be used to detect abnormal sounds from audio data, and an image analysis algorithm can be used to automatically detect defects on the surface of a product.

[1259] Step 4:

[1260] The server notifies the user's device of the analysis results and optimization proposals in real time, allowing the device to allow on-site workers to immediately check the analysis results and optimization proposals.

[1261] Step 5:

[1262] The device equipped with the emotion engine analyzes the user's voice tone and facial expression data in real time to evaluate the user's emotional state (e.g., stress, fatigue). Specifically, it detects changes in tone and speed of voice from the voice data, and senses changes in facial expression from the image data.

[1263] Step 6:

[1264] The server receives the emotional state data sent from the emotion engine. Based on this information, the server adjusts the analysis algorithm and generates optimization suggestions according to the user's emotional state. For example, if the user is in a high-stress state, the server generates suggestions for taking a break or reducing workload.

[1265] Step 7:

[1266] The server then sends the newly generated optimization proposals to the user's device, where the user can view these proposals in real time. The user can then use the device to input feedback, such as accepting the proposals or requesting modifications.

[1267] Step 8:

[1268] The server receives feedback data from users and continuously improves the analysis algorithm. The feedback, including emotional data, is analyzed to improve the accuracy of the next data analysis.

[1269] ---

[1270] This system monitors and optimizes manufacturing production lines in real time, continuously improving production efficiency and quality. The introduction of an emotion engine enables flexible responses according to the user's condition, achieving both an improved work environment and increased production efficiency.

[1271] Example 2

[1272] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1273] Improving production efficiency and quality are important challenges in modern manufacturing. However, conventional systems focused on collecting and analyzing data from sensors, and did not fully consider the impact of workers' emotional states. As a result, poor performance and mistakes caused by worker stress and fatigue had a negative impact on production efficiency and quality. This made it difficult to manage worker health and optimize production processes.

[1274] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1275] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality assessment, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for analyzing the emotional state of the workers from their voice data and image data, means for generating optimization proposals based on their emotional states, and means for improving the analysis algorithm based on feedback from the workers. This improves production efficiency and quality and enables optimal proposals that take into account the health status of the workers.

[1276] A "sensor" is a device that collects different types of data in real time, such as temperature, pressure, sound, images, and video.

[1277] "Data preprocessing" refers to the process of performing preprocessing such as noise removal and resolution standardization on collected data, and then temporarily storing it.

[1278] A "database" is a storage device for temporarily storing preprocessed data.

[1279] An "AI model" is an artificial intelligence algorithm that performs analysis for anomaly detection, quality evaluation, and process optimization.

[1280] "Analysis results" are the results of anomaly detection and quality assessment generated by the AI ​​model using preprocessed data.

[1281] An "optimization proposal" is a proposal for improving the production process that is generated based on the analysis results.

[1282] "Field workers" are workers who work on the manufacturing floor and act based on data collected from sensors.

[1283] "Feedback" refers to opinions and reactions to the system analysis and suggestions provided by field workers.

[1284] "Emotional state" refers to the emotional state of a worker, such as stress or fatigue, analyzed from the worker's voice data and image data.

[1285] The "emotion engine" is a system that analyzes the emotional state of workers from collected voice and image data.

[1286] "Notification" is a communication method by which the server conveys analysis results and optimization suggestions to field workers in real time.

[1287] An "analysis algorithm" is a calculation procedure for detecting anomalies and evaluating quality based on collected data.

[1288] This invention is an AI-driven robot support system for improving production efficiency and quality in the manufacturing industry, incorporating an emotion engine that recognizes the user's emotional state in real time and reflects it in optimization proposals. This system is composed of a server, various sensors, user terminals, and the emotion engine.

[1289] The server is the central part that collects data in real time from sensors placed on the production line and performs preprocessing and analysis. The main hardware used is a high-performance server computer, and the software includes Python, TensorFlow, PyTorch, OpenCV, Librosa, etc.

[1290] Data collection

[1291] The server uses various sensors to collect temperature, pressure, audio, image, and video data. For example, the temperature sensor measures the temperature data of the line in real time, the microphone collects user voices and machine sounds, and the image sensor takes photos of the product and records video of the production line operation.

[1292] Data Preprocessing

[1293] The collected data is sent to a server where preprocessing such as noise removal and resolution unification is performed. For example, the Librosa library is used to remove environmental noise from audio data, and the resolution of image data is unified using OpenCV. This organizes the data into a form optimal for analysis.

[1294] Data analysis

[1295] The preprocessed data is then analyzed by the server using AI models. Anomaly detection models are run using TensorFlow and PyTorch. Quality assessments and recommendations for optimizing the production process are also generated at this stage. If an anomaly is detected, its details are recorded in a log.

[1296] Emotion analysis

[1297] Furthermore, the system incorporates an emotion engine that recognizes the user's emotional state in real time. The emotion engine uses natural language processing models (e.g., BERT) and facial recognition models (e.g., DeepFace) to analyze the user's emotions from voice and image data, thereby determining whether the user is in a stressful state.

[1298] Proposal generation and notification

[1299] Optimization suggestions are generated based on the analysis results and the user's emotional state. The server notifies the user of these suggestions in real time. For example, if a user is in a high-stress state, the device may be notified of suggestions such as taking a break or redistributing tasks. The device application uses React Native or Android Studio, allowing users to view these notifications on their tablets or smartphones.

[1300] Feedback collection and system improvement

[1301] Users can provide feedback on the suggestions provided to them through their devices. The server collects this feedback and uses it to improve the analysis algorithm. This allows the system to learn from the feedback and improve the accuracy of future suggestions.

[1302] Example: Detecting user stress

[1303] The server uses audio and image sensors to detect the user's stress level from their tone of voice and facial expressions. If the emotion engine analyzes the user's emotional state as "high stress," it sends the result to the server. Based on this information, the server notifies the device with suggestions for reducing the user's workload (e.g., taking a break or switching to a lighter task). This maximizes the user's performance while maintaining their physical and mental health.

[1304] Example prompt sentence:

[1305] "Analyze the user's stress level based on their voice and image data, and make optimization suggestions based on the results."

[1306] This system will enable manufacturing production lines to operate with even greater efficiency and quality. By combining an emotion engine with AI analysis, appropriate suggestions that take into account the user's condition will be made in real time, enabling more effective optimization of the production process.

[1307] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1308] Step 1: Data collection

[1309] Inputs: temperature, pressure, audio, image, and video data.

[1310] Processing: The server collects data in real time from various sensors placed on the production line.

[1311] Output: Collected multimodal data.

[1312] How it works: The server collects real-time temperature data every 10 seconds from the temperature sensor, and collects user voice and machine sounds from the microphone. At the same time, the image sensor captures the product's appearance, and the video sensor records the operation of the production line.

[1313] Step 2: Data Preprocessing

[1314] Input: Collected multimodal data.

[1315] Processing: The server preprocesses the collected data, removing noise and unifying the resolution for temporary storage.

[1316] Output: Preprocessed data.

[1317] Specific operation: Noise reduction is performed on the audio data using the Librosa library, and the resolution of the image data is unified using OpenCV. The preprocessed data is temporarily stored in a database.

[1318] Step 3: Data analysis

[1319] Input: Preprocessed data.

[1320] Processing: The server uses the preprocessed data to perform analysis using an AI model.

[1321] Output: Analysis results (anomaly detection, quality assessment, process optimization suggestions).

[1322] What it does: Runs anomaly detection models using TensorFlow or PyTorch to generate quality assessments and process optimization suggestions. If an anomaly is detected, details of the anomaly are logged.

[1323] Step 4: Sentiment Analysis

[1324] Input: Preprocessed audio and image data.

[1325] Processing: The server analyzes the user's emotional state using an emotion engine.

[1326] Output: Emotion analysis results (e.g. high stress, fatigue).

[1327] Specific operation: Emotions are analyzed from voice data using BERT, and facial expressions are analyzed from image data using DeepFace. The analysis results of the emotional state are stored on the server.

[1328] Step 5: Proposal Generation

[1329] Input: Data analysis results and sentiment analysis results.

[1330] Processing: The server generates optimization suggestions based on the data analysis results and sentiment analysis results.

[1331] Output: Optimization suggestions (e.g. break suggestions, task redistribution).

[1332] Specific operation: Based on the analysis results, suggestions to reduce the user's workload are automatically generated and prepared for notification in the next step.

[1333] Step 6: Notification of results

[1334] Input: Optimization proposal.

[1335] Processing: The server notifies the user terminal of the generated optimization proposal.

[1336] Output: A notification message on the terminal.

[1337] Specific operation: The server pushes JSON format data to the device, and the application on the device interprets the data and displays it to the user.

[1338] Step 7: Gather feedback and improve the system

[1339] Input: User feedback.

[1340] Processing: The user enters feedback on the suggestions provided, and the server collects the feedback and uses it to improve the system.

[1341] Output: Feedback data, improved AI model.

[1342] How it works: Users input feedback from their devices, and the server collects that information and reflects it in the next analysis, thereby improving the accuracy of the analysis algorithm.

[1343] These steps not only improve the efficiency and quality of the production process, but also optimize the working environment, maintaining worker health and performance through suggestions that take the user's emotional state into account.

[1344] (Application example 2)

[1345] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1346] Conventional production systems focus on detecting machine anomalies and evaluating quality, but rarely consider the emotional state of human workers. This increases the risk of workers feeling stressed or fatigued, leading to reduced production efficiency. Furthermore, the lack of work instructions or improvement suggestions based on the emotional state of workers poses a challenge, making it difficult to fully optimize the entire production process.

[1347] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1348] In this invention, the server includes means for collecting text data, voice data, image data, and video data in real time from various sensors, means for preprocessing the collected data and temporarily storing it in a database, means for analyzing the preprocessed data and performing anomaly detection, quality assessment, and process optimization proposals, means for notifying on-site workers of the analysis results and optimization proposals in real time and receiving feedback from the workers, means for improving the analysis algorithm based on feedback from the workers, and means for analyzing the emotions of the workers in real time and providing work instructions based on their emotional states. This makes it possible to provide optimal work instructions and propose process improvements that take the emotional states of the workers into consideration, thereby achieving both improved production efficiency and maintaining the health of the workers.

[1349] "Multiple sensors" are various sensors used to collect text data, audio data, image data, and video data in real time.

[1350] "Preprocessing" refers to the process of converting collected data into a form suitable for analysis, and specifically includes noise removal and resolution unification.

[1351] A "database" is a system for temporarily storing preprocessed data.

[1352] "Analysis algorithms" are techniques for analyzing pre-processed data and generating suggestions for anomaly detection, quality assessment, and process optimization.

[1353] "Emotional state" refers to the emotional state of the worker analyzed from voice and image data, and specifically refers to stress and fatigue.

[1354] "Work instructions" refers to specific work tasks and break suggestions provided to workers based on their emotional state.

[1355] "Feedback" refers to responses and reactions from workers, and is information collected to help improve the system.

[1356] An "emotion analysis device" is a piece of equipment used to analyze the emotional state of a worker in real time.

[1357] "Production process improvement proposals" refer to specific proposals for improving production efficiency based on the analysis results.

[1358] "Stress state" refers to the degree of tension and strain felt by the worker, and is analyzed by the emotion engine.

[1359] To realize the present invention, the following detailed description of the system and its components is required. The system of the present invention collects and analyzes data in real time and provides appropriate feedback and work instructions. Specific components include a server, various sensors, a user's device, and an emotion analysis engine.

[1360] server

[1361] The server is the core processing unit of the system and performs the following functions:

[1362] Data collection: Collect text, audio, image, and video data in real time from a variety of sensors.

[1363] Data preprocessing: The collected data is preprocessed by noise removal and resolution standardization, and then temporarily stored in a database.

[1364] Data Analysis: The pre-processed data is analyzed using generative AI models to generate recommendations for anomaly detection, quality assessment, and process optimization.

[1365] Emotion analysis: An emotion analysis engine is used to analyze the user's emotional state from voice and image data.

[1366] Various sensors

[1367] The system uses the following sensors:

[1368] Audio sensor: Collects the user's voice tone and mechanical sounds to detect stress levels and abnormal sounds.

[1369] Image sensor: Collects footage of the user's facial expressions and production line operations to inspect the emotional state and appearance of the product.

[1370] User's device

[1371] The user's device (smartphone, tablet, etc.) performs the following functions:

[1372] Notifications: Analysis results and optimization suggestions from the server are displayed in real time.

[1373] Feedback: Collects feedback from users and sends it to the server.

[1374] Sentiment Analysis Engine

[1375] The emotion analysis engine analyzes the user's emotional state from voice data and image data and provides the data to the server.

[1376] This is implemented using the following software and libraries:

[1377] OpenCV: Face Recognition and Landmark Detection

[1378] dlib: Facial landmark analysis

[1379] Keras: Sentiment Analysis Model

[1380] requests: Send notification

[1381] Specific examples

[1382] For example, when developing an application to detect the stress level of workers at a logistics center, the smartphone's camera and microphone can be used to analyze the worker's facial expressions and tone of voice. If the worker's emotional state is determined to be "high stress," the server will send a notification to take a break. In this way, optimal work instructions that take the worker's emotions into consideration can be provided, thereby improving production efficiency and maintaining the worker's health at the same time.

[1383] Prompt Sentence Examples

[1384] Develop an application that evaluates workers' emotions through real-time analysis of camera footage and suggests breaks when they are stressed. The technologies used are OpenCV, dlib, and Keras. The goal is to analyze emotions from workers' facial expressions and suggest breaks when stress or fatigue is detected.

[1385] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1386] Step 1:

[1387] The server collects text data, voice data, image data, and video data in real time from various sensors (audio sensors, image sensors). The input is data from the sensors, and the output is raw data for preprocessing. Specifically, the audio sensor collects the tone of the user's voice, and the image sensor captures the user's facial expressions and footage of the production line operation.

[1388] Step 2:

[1389] The server preprocesses the collected data. Specifically, it removes noise from the audio data and standardizes the resolution of the image data. At this point, the input is raw data and the output is preprocessed data. The preprocessed data is temporarily stored in a database.

[1390] Step 3:

[1391] The server inputs the preprocessed data into the generative AI model for analysis. The input is the preprocessed data, and the output of the analysis is anomaly detection, quality assessment, and process optimization recommendations. Specifically, the AI ​​model uses image data to perform visual inspections of products and analyzes audio data to detect abnormal machine sounds.

[1392] Step 4:

[1393] The server analyzes the user's emotional state using an emotion analysis engine. The input is audio and image data, and the output is the user's emotional state (stress or fatigue). Specifically, it uses OpenCV and dlib to detect facial landmarks, and classifies the emotional state using an emotion analysis model using Keras.

[1394] Step 5:

[1395] The server notifies the user device of the analysis results and optimization suggestions in real time. The input is the analysis results and sentiment analysis results, and the output is a notification message. Specifically, it uses the requests library to send a message to the notification API, which is then displayed on the user device. For example, if a high stress state is detected, a notification such as "Take a break" is sent.

[1396] Step 6:

[1397] The user checks the suggestions through the terminal and takes the necessary action. The input is a notification message and the output is feedback information. The suggestions are displayed on the terminal, and the user can check them and take action.

[1398] Step 7:

[1399] The terminal collects feedback from the user and sends it to the server. The input is the feedback information from the user, and the output is the feedback data sent to the server. Specifically, an interface is provided for the user to report task completion notifications and new problems.

[1400] Step 8:

[1401] The server improves the analysis algorithm based on the collected feedback. The input is the feedback data, and the output is an improved analysis algorithm. Specifically, it reevaluates the analysis results and adjusts the model parameters.

[1402] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1403] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1404] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1405] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1406] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1407] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1408] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1409] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1410] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1411] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1412] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1413] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1414] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1415] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1416] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1417] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1418] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1419] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1420] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1421] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1422] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1423] The following is further disclosed regarding the above embodiment.

[1424] (Claim 1)

[1425] A means for collecting text data, audio data, image data, and video data in real time from a variety of sensors;

[1426] A means for preprocessing the collected data and temporarily storing it in a database;

[1427] A means of analyzing the pre-processed data to perform anomaly detection, quality assessment, and process optimization suggestions;

[1428] A means to communicate analysis results and optimization suggestions to field workers in real time and receive feedback from them;

[1429] A means to improve the analysis algorithm based on feedback from workers;

[1430] A system including:

[1431] (Claim 2)

[1432] A means for inputting the preprocessed data into an artificial intelligence model for analysis;

[1433] means for generating improvement proposals for the production process based on the analysis results;

[1434] 10. The system of claim 1, further comprising:

[1435] (Claim 3)

[1436] A means for analyzing audio data from sensors to detect abnormal machine sounds;

[1437] means for visually inspecting the product using image data from the sensor;

[1438] 10. The system of claim 1, further comprising:

[1439] "Example 1"

[1440] (Claim 1)

[1441] A means for collecting text data, audio data, image data, and video data in real time from a variety of sensors;

[1442] A means for preprocessing the collected data and temporarily storing it in a database;

[1443] A means of analyzing the pre-processed data to perform anomaly detection, quality assessment, and process optimization suggestions;

[1444] A means to communicate analysis results and optimization suggestions to field workers in real time and receive feedback from them;

[1445] A means to improve the analysis algorithm based on feedback from workers;

[1446] A system including:

[1447] (Claim 2)

[1448] A means of inputting preprocessed data into a generative AI model for analysis;

[1449] means for generating improvement proposals for the production process based on the analysis results;

[1450] A means to notify the analysis results and improvement proposals to the mobile devices of field workers in real time,

[1451] 10. The system of claim 1, further comprising:

[1452] (Claim 3)

[1453] A means for analyzing audio data from sensors to detect abnormal machine sounds;

[1454] means for visually inspecting the product using image data from the sensor;

[1455] A means to notify on-site workers of abnormal sounds and visual inspection results and to inform them of the need for maintenance;

[1456] 10. The system of claim 1, further comprising:

[1457] "Application Example 1"

[1458] (Claim 1)

[1459] A means for collecting text data, audio data, image data, and video data in real time from a variety of sensors;

[1460] A means for preprocessing the collected data and temporarily storing it in a database;

[1461] A means of analyzing the pre-processed data to perform anomaly detection, quality assessment, and process optimization suggestions;

[1462] A means to communicate analysis results and optimization suggestions to field workers in real time and receive feedback from them;

[1463] A means to improve the analysis algorithm based on feedback from workers;

[1464] A means for the robots installed on the production line to collect operational data in real time and send it to a server;

[1465] A method for the robot to autonomously optimize itself based on the analysis results from the server, and

[1466] A means for notifying the analysis results to the terminals of field workers so that they can make manual adjustments;

[1467] A system including:

[1468] (Claim 2)

[1469] A means for inputting the preprocessed data into an artificial intelligence model for analysis;

[1470] means for generating improvement proposals for the production process based on the analysis results;

[1471] 10. The system of claim 1.

[1472] (Claim 3)

[1473] A means for analyzing audio data from sensors to detect abnormal machine sounds;

[1474] means for visually inspecting the product using image data from the sensor;

[1475] A means for robots installed on production lines to detect abnormalities and respond autonomously;

[1476] 10. The system of claim 1.

[1477] "Example 2: Combining Emotion Engines"

[1478] (Claim 1)

[1479] A means for collecting text data, audio data, image data, and video data in real time from a variety of sensors;

[1480] A means for preprocessing the collected data and temporarily storing it in a database;

[1481] A means of analyzing the pre-processed data to perform anomaly detection, quality assessment, and process optimization suggestions;

[1482] A means to communicate analysis results and optimization suggestions to field workers in real time and receive feedback from them;

[1483] means for analyzing the emotional state of a worker from voice data and image data;

[1484] means for generating optimization suggestions based on the emotional state;

[1485] A means to improve the analysis algorithm based on feedback from workers;

[1486] A system including:

[1487] (Claim 2)

[1488] A means for inputting the preprocessed data into an artificial intelligence model for analysis;

[1489] means for generating improvement proposals for the production process based on the analysis results;

[1490] 10. The system of claim 1.

[1491] (Claim 3)

[1492] A means for analyzing audio data from sensors to detect abnormal machine sounds;

[1493] means for visually inspecting the product using image data from the sensor;

[1494] 10. The system of claim 1.

[1495] "Application example 2 when combining emotion engines"

[1496] (Claim 1)

[1497] A means for collecting text data, audio data, image data, and video data in real time from a variety of sensors;

[1498] A means for preprocessing the collected data and temporarily storing it in a database;

[1499] A means of analyzing the pre-processed data to perform anomaly detection, quality assessment, and process optimization suggestions;

[1500] A means to communicate analysis results and optimization suggestions to field workers in real time and receive feedback from them;

[1501] A means to improve the analysis algorithm based on feedback from workers;

[1502] means for analyzing the emotions of the workers in real time and providing work instructions based on the emotional state;

[1503] A system including:

[1504] (Claim 2)

[1505] A means for inputting the preprocessed data into an artificial intelligence model for analysis;

[1506] means for generating improvement proposals for the production process based on the analysis results;

[1507] The system of claim 1 , further comprising: means for generating optimal work instructions based on the results of the worker's sentiment analysis.

[1508] (Claim 3)

[1509] A means for analyzing audio data from sensors to detect abnormal machine sounds;

[1510] means for visually inspecting the product using image data from the sensor;

[1511] 10. The system of claim 1, further comprising means for analyzing the worker's emotion data to detect stress or fatigue. [Explanation of symbols]

[1512] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for collecting text data, audio data, image data, and video data in real time from a variety of sensors; A means for preprocessing the collected data and temporarily storing it in a database; A means of analyzing the pre-processed data to perform anomaly detection, quality assessment, and process optimization suggestions; A means to communicate analysis results and optimization suggestions to field workers in real time and receive feedback from them; A means to improve the analysis algorithm based on feedback from workers; A system including:

2. A means for inputting the preprocessed data into an artificial intelligence model for analysis; means for generating improvement proposals for the production process based on the analysis results; The system of claim 1 further comprising:

3. A means for detecting abnormal machine sounds by analyzing audio data from sensors; means for visually inspecting the product using image data from the sensor; The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A